Long document semantic coherence processing method, device, storage medium and computer equipment
Through technical means such as sliding window converter and memory management module, the problems of poor semantic coherence and high computational complexity in traditional long document processing are solved, and efficient long document processing and semantic coherence generation are achieved.
Patent Information
- Application Number
- CN202510348649.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-24
AI Technical Summary
Traditional long document processing methods lead to poor semantic coherence and high computational complexity, making it difficult to meet real-time interaction requirements.
A sliding window converter is used to encode long documents in blocks, and through window division units, position encoding units and converter encoding units, the block position is adjusted and overlapping areas are set between adjacent blocks. At the same time, the memory management module is used to store and manage the encoded document content, deeply understand the semantics through the semantic processing module, and generate semantic coherent output content through the content generation module.
Improves semantic coherence of long document processing, reduces computational complexity, meets real-time interaction requirements, and generates coherent, smooth and fact-accurate output content.
Smart Images

Figure CN119886072B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and specifically to a method, device, storage medium and computer equipment for processing semantic coherence of a long document. The method comprises: encoding a long document in blocks; storing and managing the encoded document content; understanding the semantics of the document content; providing retrieval results for the document content; and generating output content. Background Art
[0002] Traditional long document processing methods usually adopt a fixed window size block processing strategy to divide long documents into several independent small blocks for processing. Although this method is simple and direct, it often leads to the fragmentation of contextual information, making it difficult to maintain semantic connections across blocks, especially when key information is scattered in different document blocks, the accuracy and coherence of semantic understanding are poor.
[0003] In addition, existing long document processing systems usually have high computational complexity and low processing efficiency, making it difficult to meet the needs of real-time interaction, which limits their promotion and use in large-scale application scenarios.
[0004] Therefore, it is necessary to propose a new technical solution to solve the above technical problems. Summary of the invention
[0005] The purpose of the embodiments of the present application is to provide a method, apparatus, storage medium and computer device for processing semantic coherence of a long document, aiming to improve the semantic coherence of long document processing.
[0006] The embodiment of the present application provides a method for processing the semantic coherence of a long document, comprising: encoding the long document in blocks; storing and managing the encoded document content; semantically understanding the document content; providing retrieval results of the document content; and generating output content; the method further comprises: encoding the long document in blocks by a sliding window transformer, the sliding window transformer comprising a window division unit, a position encoding unit and a transformer encoding unit, wherein the window division unit is used to adjust the block position according to the sentence and paragraph boundaries and set an overlapping area between adjacent blocks; storing and managing the encoded document content by a memory management module, the memory management module comprising a chapter-level memory Pool and gated memory unit, the chapter-level memory pool is used to store document encoding content, and the gated memory unit is used to maintain entity and event timeline; the document content is deeply semantically understood through the semantic processing module, and the semantic processing module is used to capture various dependencies between paragraphs by calculating the multi-dimensional attention weights between paragraphs; the retrieval module provides retrieval results of document content; the content generation module generates semantically coherent output content, and the content generation module includes a memory pointer network and a consistency verification component. The memory pointer network is used to enhance the coherence and accuracy of the generated content by dynamically citing historical context, and the consistency verification component is used to ensure that the generated content is consistent with the facts of the source document.
[0007] In the above method, the overlapping length of the overlapping area is O marks, where O is less than L, and the window division unit is used to create a window containing the current block and its context, which window contains the last O marks of the previous block, all marks of the current block and the first O marks of the next block, where O and L are both positive integers.
[0008] In the above method, the position encoding unit is used to assign a local position code and a global position code to each tag in the window, the local position code is used to map the serial position of the tag in the current window to a continuous encoding space, and the global position code is used to represent the absolute position of the tag in the entire document.
[0009] In the above method, the gated memory unit includes an entity recognition module, a state tracking module and a timeline maintenance module. The entity recognition module is used to identify key entities and events from the document content, the state tracking module is used to update the entity state using a gated update method, and the timeline maintenance module is used to construct and update the timeline of entity-related events.
[0010] The embodiment of the present application also provides a long document semantic coherence processing device, including: a document encoding module, which is used to encode the long document in blocks through a sliding window transformer, and the sliding window transformer includes a window division unit, a position encoding unit and a transformer encoding unit, wherein the window division unit is used to adjust the block position according to the sentence and paragraph boundaries, and set an overlapping area between adjacent blocks; a memory management module, which is used to store and manage the encoded document content, and the memory management module includes a chapter-level memory pool and a gated memory unit, the chapter-level memory pool is used to store the document encoding content, and the gated memory unit is used to maintain the entity and event timeline; a semantic processing module, which is used to perform deep semantic understanding of the document content, and the semantic processing module is used to capture multiple dependencies between paragraphs by calculating the multi-dimensional attention weights between paragraphs; a retrieval module, which is used to provide retrieval results for the document content; a content generation module, which is used to generate semantically coherent output content, and the content generation module includes a memory pointer network and a consistency verification component, the memory pointer network is used to enhance the coherence and accuracy of the generated content by dynamically citing the historical context, and the consistency verification component is used to ensure that the generated content is consistent with the facts of the source document.
[0011] In the above-mentioned device, the overlapping length of the overlapping area is O marks, where O is less than L, and the window division unit is used to create a window containing the current block and its context, which window contains the last O marks of the previous block, all marks of the current block and the first O marks of the next block, where O and L are both positive integers.
[0012] In the above device, the position encoding unit is used to assign a local position code and a global position code to each tag in the window, the local position code is used to map the serial position of the tag in the current window to a continuous encoding space, and the global position code is used to represent the absolute position of the tag in the entire document.
[0013] In the above-mentioned device, the gated memory unit includes an entity recognition module, a state tracking module and a timeline maintenance module. The entity recognition module is used to identify key entities and events from the document content, the state tracking module is used to update the entity state using a gated update method, and the timeline maintenance module is used to construct and update the timeline of entity-related events.
[0014] The embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed, the steps of the above-mentioned method for processing semantic coherence of a long document are implemented.
[0015] An embodiment of the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned method for processing semantic coherence of a long document when executing the computer program.
[0016] In the long document semantic coherence processing method, apparatus, storage medium, and computer device provided in this application, since a sliding window transformer is used to perform chunk encoding on the long document, and the sliding window transformer includes a window partitioning unit, a positional encoding unit, and a transformer encoding unit. Among them, the window partitioning unit is different from the traditional fixed-size chunking strategy. Instead, it adjusts the chunk positions according to the semantic boundaries of sentences and paragraphs, and sets an overlapping area between adjacent chunks. Therefore, it can make the chunking process respect the natural semantic units of the text, avoid the situation of splitting a semantically complete sentence or paragraph into different chunks, and thus prevent semantic fragmentation at the source. At the same time, the overlapping area between adjacent chunks provides a bridge for semantic transmission between chunks, enabling the preservation and transmission of cross-chunk semantic information, and fundamentally improving the semantic coherence of long document processing. In addition, since the encoded document content is stored and managed through a memory management module, which includes a discourse-level memory pool and a gated memory unit, the discourse-level memory pool stores the document encoded content and provides a unified information access interface for subsequent processing; the gated memory unit maintains the entity and event timeline, tracks the state changes of key entities and the development of events in the document. Therefore, the system can maintain a long-term semantic understanding of the document content, not only solving the semantic incoherence problem caused by information isolation between chunks in traditional methods, but also providing support for larger-span semantic associations, and enhancing the system's ability to grasp the overall semantics of long documents. In addition, since the semantic processing module captures various dependencies between paragraphs by calculating multi-dimensional attention weights between paragraphs, achieving a deep semantic understanding of the document content. Compared with the limitation of traditional methods that only focus on local semantics, the semantic processing module can identify and model complex semantic relationships between paragraphs, including topic coherence, logical dependencies, referential relationships, etc. In this way, the understanding of the document is no longer limited to independent text chunks, but forms a semantic network that runs through the entire text, significantly improving the accuracy and coherence of semantic understanding. In addition, since the content generation module includes a memory pointer network and a consistency verification component, jointly ensuring the semantic coherence and factual accuracy of the output content. The memory pointer network makes full use of the existing information when generating content by dynamically referring to the historical context, avoiding content incoherence caused by information loss in traditional methods; the consistency verification component ensures the factual accuracy of the output content by verifying the consistency between the generated content and the source document. The collaborative work of these two components enables the system to generate content that is both coherent and faithful to the original text, effectively solving the semantic break problem in long document processing. Description of the Drawings
[0017] Figure 1 It is a schematic diagram of the application scenario of the long document semantic coherence processing method and apparatus provided by the embodiments of this application.
[0018] Figure 2It is a block diagram of a long document semantic coherence processing device provided in an embodiment of the present application.
[0019] Figure 3 yes Figure 2 The block diagram of the document encoding module of the long document semantic coherence processing device is shown.
[0020] Figure 4 yes Figure 2 The block diagram of the memory management module of the long document semantic coherence processing device is shown.
[0021] Figure 5 yes Figure 4 Block diagram of the gated memory unit of the memory management module shown.
[0022] Figure 6 It is a flowchart of a method for processing semantic coherence of a long document provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The specific implementation methods of the present application are described in detail below with reference to the accompanying drawings.
[0024] The terms "first", "second" and similar words do not indicate any order, quantity or importance, but are only used to distinguish different technical features. The term "plurality" and similar words mean two or more, unless otherwise clearly defined.
[0025] The embodiments of the present application may be combined with each other.
[0026] The embodiments of the present application provide a method and device for processing semantic coherence of long documents, which solve the semantic coherence problem in the long document processing process and improve the efficiency of long document processing through a hierarchical memory enhancement architecture and a dual-channel retrieval strategy.
[0027] like Figure 1 , Figure 2 and Figure 6 As shown, the long document semantic coherence processing method and apparatus of the present application can be applied to a computer device 101, which can be, for example, a desktop computer, a laptop computer, a smart phone, etc.
[0028] The long document semantic coherence processing method and device provided in the present application can be implemented by hardware, and the hardware may include any combination of a processor, a memory, a communication circuit, etc., wherein the memory and the communication circuit are electrically connected to the processor. Any combination of the above processor, memory, communication circuit, etc. is used to implement the functions and steps of the long document semantic coherence processing method and device provided in the present application.
[0029] Among them, the processor may be, for example: CPU (Central Processing Unit), GPU, NPU (Neural Processing Unit), other general-purpose processors, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0030] The memory may include a read-only memory and a random access memory for providing program code and data to the processor. The memory may also include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache.
[0031] The long document semantic coherence processing method and device provided by the present application can also be implemented by software. In this case, the long document semantic coherence processing method and device provided by the present application and their respective modules can also be software modules. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product (whose carrier can be, for example, the computer-readable storage medium of the present application).
[0032] The long document semantic coherence processing method and device provided in the present application can also be implemented through a combination of software and hardware.
[0033] The computer device provided by the present application includes a processor and a memory, wherein the processor and the memory communicate via a bus. The memory is used to store program codes, and when the computer is running, the processor executes the program codes to execute the long document semantic coherence processing method provided by the present application.
[0034] The computer-readable storage medium of the present application stores program codes, and the program codes are used to enable a computer to execute the long document semantic coherence processing method provided by the present application.
[0035] The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium may be a solid state drive (SSD).
[0036] The program code instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the program code instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0037] The long document semantic coherence processing device of the present application comprises a document encoding module 201, a memory management module 202, a semantic processing module 203, a retrieval module 204 and a content generation module 205. The document encoding module 201 is used to encode the input long document in blocks (step 601); the memory management module 202 is used to store and manage the encoded document content (step 602); the semantic processing module 203 is used to achieve deep semantic understanding of the document content (step 603); the retrieval module 204 is used to provide retrieval results of the document content (step 604); and the content generation module 205 is used to generate semantically coherent output content (step 605).
[0038] The document encoding module 201 uses a sliding window transformer to perform block encoding on a long document. Figure 3 As shown, the sliding window transformer includes a window division unit 2011, a position encoding unit 2012 and a transformer encoding unit 2013.
[0039] The window division unit 2011 performs preliminary segmentation on the input long document D, with the target block length being L tokens. In actual segmentation, the window division unit 2011 implements an adaptive block strategy, adjusting the block position according to the sentence and paragraph boundaries to avoid segmenting semantically complete sentences into different blocks. This strategy makes the final block {B 1 , B 2,..., Bₙ} may have slightly different lengths but are all close to the target length L. The window partitioning unit 2011 identifies sentence boundaries such as full stops, question marks, and exclamation marks through a sentence boundary detector and preferentially splits at sentence boundaries. For extremely long sentences, the window partitioning unit 2011 performs sub-optimal splitting at grammatical pauses such as commas and semicolons. The window partitioning unit 2011 sets an overlapping region between adjacent blocks, with an overlapping length of O tokens (where O < L) to ensure context continuity. The size O of the overlapping region is automatically set according to the document type. For term-intensive documents such as professional papers, the window partitioning unit 2011 increases the overlapping region to maintain the semantic integrity of technical terms; for narrative documents, the window partitioning unit 2011 appropriately reduces the overlapping region. For each block Bᵢ, the window partitioning unit 2011 creates a window Wᵢ that includes the block and its context, which contains the last O tokens of the previous block, all the tokens of the current block, and the first O tokens of the next block. For special blocks at the beginning and end of the document, the window partitioning unit 2011 complements the context by adding special padding tokens to ensure that all window sizes are the same. Here, i, O, and L are all positive integers.
[0040] The position encoding unit 2012 assigns two types of position encodings to each token within the window: local position encoding and global position encoding. The position encoding unit 2012 maps the sequential position of the token within the current window to a continuous encoding space with a range of [0, 2L - O] through a specially set position mapping function. The position encoding unit 2012 performs a non-linear transformation on the local position, such that the tokens at the center position of the window obtain a more refined position representation, while the tokens at the edge positions obtain a coarser-grained position representation. The global position encoding represents the absolute position of the token in the entire document to maintain global consistency. The position encoding unit 2012 generates the global position encoding through a hierarchical encoding method, decomposing the document position into three levels: chapter position, paragraph position, and intra-sentence position. Each level uses an independent encoding space, and then the final global position encoding is generated through concatenation and linear projection. This hierarchical encoding method enables the model to simultaneously perceive the position relationship of the token in the local context and the global structure. The position encoding unit 2012 generates the two types of position encodings through cosine position encoding and sets a special scaling factor to control the relative strength of the two encodings, and then adds them to the word embedding vector of the token to form a position-aware token representation. The position encoding unit 2012 performs a smooth transition process for position encoding. For the tokens in the overlapping region, their global position encodings remain consistent, while the local position encodings achieve smooth transition through interpolation to avoid representation mutations caused by position discontinuity.
[0041] The transformer encoding unit 2013 uses an improved multi-head self-attention mechanism and a feedforward neural network to process the input with added position encoding. The transformer encoding unit 2013 internally constructs a multi-layer encoding structure, each layer of which contains three submodules: a multi-head self-attention submodule, a feedforward neural network submodule, and a cross-block connection submodule. The multi-head self-attention submodule projects the input to multiple attention heads, each of which independently calculates the attention weight and context representation, and then merges the multi-head results to obtain a comprehensive representation. The feedforward neural network submodule adopts a two-layer fully connected network structure, and uses the GELU (Gaussian Error Linear Unit) activation function in the middle to perform nonlinear transformation on the attention output. The cross-block connection submodule specifically handles the information interaction between blocks, stores the key information of adjacent blocks through a memory cache mechanism, and selectively introduces relevant information during the current block processing. To improve processing efficiency, the encoding unit implements a sparse attention method and sets three attention modes: local sliding attention, global sparse attention, and memory-enhanced attention. Local sliding attention limits each tag to focus only on the context within a fixed window; global sparse attention selects a small number of key tags for global attention calculation through an importance sampling strategy; memory-enhanced attention introduces an external memory vector to summarize long-distance dependency information. Through the combination of these three attention modes, the computational complexity is reduced from O(n²) to O(n×w+c), where n is the number of tags, w is the size of the attention window, and c is the number of global sampling points, significantly improving the processing speed while maintaining the ability to model long-distance dependencies.
[0042] For each tag in the window, the transformer encoding unit 2013 outputs its context representation vector, which contains two parts: local semantic information and global context information. The transformer encoding unit 2013 processes the representation vector through residual connection and layer normalization to ensure that the information will not decay during deep transmission. In order to enhance the processing robustness, the encoding unit implements an adaptive feature fusion strategy to dynamically adjust the proportion of local features and global features contained in its representation vector according to the importance and semantic integrity of the tag. For the tags in the overlapping area, the transformer encoding unit 2013 retains its representation in the central block and discards the representation in the edge block to ensure that each original tag has only one final representation. In addition, in order to handle the smooth transition of the overlapping area, the encoding unit sets a representation transition function and applies linear interpolation to the boundary tags of the overlapping area to avoid mutations in the representation vector. Finally, all tags in the document obtain a representation set H, which is organized into a multi-level structure: tag-level representation, sentence-level representation, and paragraph-level representation, forming a hierarchical document representation system, which is convenient for subsequent modules to perform multi-granular semantic processing.
[0043] like Figure 4As shown, the memory management module 202 includes two main components: a chapter-level memory pool 2021 and a gated memory unit 2022, which are synchronized through a bidirectional information channel. The chapter-level memory pool 2021 stores document encoding content and provides retrieval services, and the gated memory unit 2022 maintains entity and event timelines. The two work together to ensure consistent representation of document content and entity relationships. The memory management module 202 sets a unified memory access interface, supports three retrieval modes: exact matching, fuzzy matching, and semantic matching, and optimizes the retrieval speed of high-frequency access content through a caching mechanism. At the same time, the memory management module 202 regularly performs consistency checks to ensure information consistency between chapter-level memory and entity memory, and avoid semantic contradictions caused by memory splitting.
[0044] The chapter-level memory pool 2021 adopts a two-level storage structure: a short-term memory area and a long-term memory area. A memory bridge is set between the two to perform automatic migration and synchronization of memory entries. The short-term memory area stores the document representation of the current processing window. It adopts a hierarchical index structure to organize the memory according to three granularities: tags, sentences, and paragraphs, and supports multi-granularity retrieval. The long-term memory area stores important information of historical processing. It adopts a semantic hash index to divide the semantic space into multiple subspaces to accelerate the retrieval process. The short-term memory area is set as a fixed-size ring buffer with a size of k, and a block storage strategy is implemented internally. Adjacent tag representations are stored in continuous memory blocks to improve access locality and cache hit rate. When a new document block is encoded, the chapter-level memory pool 2021 first allocates storage space in the short-term memory area. If the short-term memory area is full, the adaptive elimination algorithm is triggered. The algorithm comprehensively considers the timestamp, access frequency, and importance score of the memory entry to select the most suitable entry for elimination. To optimize memory usage, the short-term memory area implements memory reuse technology. For repeated content fragments, the storage space is shared through a reference counting mechanism instead of creating multiple copies. For each memory entry, the chapter-level memory pool 2021 sets a structured storage format, which includes two parts: the metadata area and the content area. The metadata area stores management information such as the tag representation vector, the position index of the tag in the document, the ID of the sentence or paragraph to which it belongs, the semantic type label, and the importance score, while the content area stores the actual representation vector data.
[0045] The long-term memory area is set as a priority queue of variable size, organized according to the importance score of the content, and implements a multi-level cache architecture to store frequently accessed high-importance content in the cache layer to reduce retrieval latency. The importance score is calculated by a multi-factor scoring model that integrates three core factors: information entropy, citation frequency, and time decay factor, as well as auxiliary factors such as context relevance and keyword density. Information entropy evaluates the information richness of the content, which is quantified by calculating the entropy value of the vector distribution. The higher the entropy value, the greater the amount of information the content carries. The citation frequency records the number of times the content is accessed by subsequent processing, and is updated using exponential smoothing to give higher weights to recent citations. The time decay factor uses a piecewise decay function, using different decay rates for different types of content: factual content uses a slower decay rate, while content with strong temporality uses a faster decay rate. The long-term memory area implements a semantic clustering function, organizing semantically similar memory items together to form a topic cluster. Each topic cluster maintains a center vector and a member list, and accelerates the retrieval process through hierarchical indexing.
[0046] To prevent memory overflow, the long-term memory area implements a multi-strategy automatic compression function, and tracks memory usage in real time through the memory monitoring component. When the memory usage exceeds the preset threshold or the memory size exceeds the capacity limit, the compression process is automatically triggered. The compression process first sorts the memory items by importance score, and then performs three-level compression operations in sequence: the first-level compression discards items with importance scores lower than the dynamic threshold, and the threshold is automatically adjusted according to the current memory pressure; the second-level compression merges similar content, identifies nearly duplicate content through cosine similarity, and retains the version with higher information entropy as a representative, while recording the key difference information of the merged items; the third-level compression compresses and encodes the content that has not been accessed for a long time, and uses an autoencoder to compress the original representation into a low-dimensional representation. The compression ratio can reach 8:1, and the original representation is restored through the decoder when necessary. During the compression process, the long-term memory area maintains an operation log, records the entry ID and operation type of each compression, and supports rollback operations when necessary.
[0047] The gated memory unit 2022 tracks and maintains the status and timeline of key entities and events appearing in the document to ensure the factual consistency of the generated content. Figure 5As shown, the gated memory unit 2022 includes three core components: entity recognition module 20221, state tracking module 20222, and timeline maintenance module 20223, as well as three auxiliary components: entity linking submodule, relationship extraction submodule, and consistency verification submodule. The components transmit messages through the event bus to achieve loosely coupled inter-module collaboration. The entity recognition module 20221 combines rules and deep learning methods to identify key entities and events from the document content. Entity recognition adopts a dual-path architecture: the rule path uses domain-specific named entity recognition rules to cover common entity types; the learning path deploys a sequence annotation model based on BERT (Bidirectional Encoder Representation Transformer) to identify complex or novel entity expressions. The recognition results are merged through a confidence fusion algorithm to improve the recognition accuracy. Event recognition adopts a two-stage method of trigger word detection and argument extraction. First, the trigger word indicating the occurrence of the event is identified, and then the argument information such as event participants, time, and location related to the trigger word is extracted to construct a complete event representation. For each identified entity, the entity recognition module 20221 creates a structured entity memory entry, which contains multiple fields: a unique entity ID (for cross-document entity alignment), an entity type (to support a hierarchical classification system), an attribute set (containing static and dynamic attributes of the entity), a state vector (a continuous vector representation of the current state of the entity), a relationship network (to store relationships with other entities), and timestamp information (to record the time when the status is updated).
[0048] The state tracking module 20222 adopts a gated update method to monitor new information related to the identified entity in the document stream and update the entity state. The state tracking module 20222 sets a three-level update strategy: incremental update, replacement update, and fusion update. Incremental update is applicable to the case where the new information is compatible with the existing information, and the new information is added to the existing state; replacement update is applicable to the case where the new information clearly covers the old information, and the old state is directly replaced by the new state; fusion update is applicable to the case where there is a partial conflict between the new and old information, and a comprehensive state is generated through weighted average or conflict resolution rules. When a new event involving an entity is detected, the state tracking module 20222 first evaluates the reliability and relevance of the event, and then selects a suitable update strategy to control the update degree of the state vector through a gating network. The gating network contains two key gating units: a reset gate and an update gate. The reset gate controls which information in the old state should be ignored, and the update gate controls the degree of integration of the new information. This dual-gating setting allows the module to perform fine-grained state updates to ensure the consistency and timeliness of entity state information. In addition, the status tracking module 20222 also maintains the status update history, records the time, source and change content of each update, and supports status backtracking and change auditing.
[0049] The timeline maintenance module 20223 builds and updates the timeline of entity-related events, providing temporal support for semantic understanding and content generation. The module implements multi-granularity time representation, supports different time scales from seconds to decades, and automatically selects the appropriate time granularity according to event characteristics. The timeline maintenance module 20223 organizes the event sequence related to the entity in chronological order, and each event node contains a complete event description, timestamp, evidence source, and associated entity. In order to handle the common fuzzy time expressions in documents, the module integrates a special time normalization component, which converts relative time expressions (such as "yesterday", "last month", "several years ago") and fuzzy time expressions (such as "recently", "soon after") into standard time interval representations through contextual reasoning and external knowledge supplementation. Each time expression is mapped to a central time point and an uncertainty range. When a potential time conflict is detected, the timeline maintenance module 20223 starts the conflict resolution process: first, the severity of the conflict is evaluated. For minor conflicts (such as slight deviations in time points), the time points are adjusted using a weighted average method; for serious conflicts (such as inconsistencies in the order of events), the reliability of the information source, contextual coherence, and logical consistency are used to determine which time representation is more credible, and the timeline is updated accordingly. In addition, the module also implements event causal reasoning functions, analyzes the causal relationship between events, constructs a causal graph network, and enhances the semantic understanding of event sequences.
[0050] The semantic processing module 203 uses a paragraph-level attention algorithm to deeply understand the document content, including three main steps: paragraph representation generation, attention calculation and fusion. The semantic processing module 203 adopts a hierarchical processing architecture to abstract the document content from tags, sentences to paragraphs step by step, and capture semantic features of different granularities. At the bottom layer, the module processes the tag sequence and extracts vocabulary and phrase-level features; the middle layer processes sentences to capture syntactic structure and local semantics; the high layer processes paragraphs and chapters to understand the global theme and document structure. First, based on the natural paragraph division of the document, the semantic processing module 203 constructs an initial representation vector for each paragraph. The paragraph division adopts a multi-feature boundary detection algorithm, combining typesetting features such as line breaks, indents, paragraph spacing, and semantic features such as topic transitions and topic coherence to accurately identify paragraph boundaries. The paragraph representation is generated by the weighted average of all the tag representations in the paragraph, and the weight calculation integrates multiple features: in addition to the basic word frequency-inverse document frequency value, the syntactic role of the tag (subject, predicate, etc. have higher weights), position information (the beginning and end of the paragraph have higher weights) and semantic importance (key components such as entities and events have higher weights) are also considered. The weight generation adopts the attention mechanism to make the model adaptively focus on the key information in the paragraph and highlight the core semantics in the paragraph. In addition, the semantic processing module 203 also implements the context-aware paragraph representation enhancement function, which merges the paragraph representation with the semantic information of its context paragraph to generate a more coherent paragraph representation.
[0051] The paragraph-level attention algorithm captures multiple dependencies between paragraphs by calculating multi-dimensional attention weights between paragraphs. The algorithm sets four special attention channels: semantic similarity channel, topic coherence channel, reference relationship channel and logical relationship channel. The semantic similarity channel captures the content similarity between paragraphs and calculates the attention weight based on the cosine similarity of paragraph representations; the topic coherence channel captures the topic continuation relationship between paragraphs and detects topic coherence through the smooth transition of topic vectors; the reference relationship channel specifically handles explicit and implicit references in documents and identifies the reference relationship between paragraphs through keyword matching and entity co-reference; the logical relationship channel focuses on the logical relationships such as causality, contrast, and progression between paragraphs, and infers the logical dependencies between paragraphs by identifying logical connectives and sentence structures. The attention calculation takes into account the semantic association and position relationship between paragraphs, and applies a nonlinear position attenuation factor to paragraphs that are far away. The factor is calculated by the adaptive function d_ij = 1 / log(a·|ij|+b), where parameters a and b are automatically adjusted according to the document type and length to more accurately model the attenuation characteristics of long-distance dependencies. The multi-channel attention weights are integrated through a gated fusion network, which dynamically adjusts the weights of each channel according to the content characteristics of the current paragraph and the requirements of the processing task, so that the attention mechanism can adapt to different types of paragraph relationships. After attention calculation, each paragraph obtains a context-enhanced representation, which not only contains the semantic information of the paragraph itself, but also integrates the information of other related paragraphs to form a global context-aware paragraph representation.
[0052] In order to further enhance the understanding of paragraph semantics, the semantic processing module 203 introduces a multi-level semantic fusion network. The network adopts a tree structure, organizes paragraphs according to their hierarchical structure in the document (such as chapters, sections, paragraphs), establishes clear subordinate relationships, and fuses representations from bottom to top. Each node maintains two representations: self-representation (containing only the content information of the current unit) and synthetic representation (integrating the information of the lower-level units). For multiple paragraphs belonging to the same upper-level unit, their representations are weighted fused through a hierarchical attention mechanism, which first evaluates the contribution of each paragraph to the theme of the upper-level unit, and then allocates fusion weights accordingly to ensure that important paragraphs play a greater role in the fusion process. During the fusion process, the network also applies residual connection and layer normalization technology to avoid information attenuation or distortion in the multi-layer fusion process. In addition, the multi-level semantic fusion network realizes bidirectional information flow: in addition to bottom-up information aggregation, it also supports top-down information distribution, injecting high-level structural information and topic information into low-level units to enhance their perception of the global context. This hierarchical fusion architecture enables the semantic processing module 203 to understand both the local details and the global structure of the document, and construct a complete and coherent semantic representation of the document.
[0053] The retrieval module 204 adopts a dual-channel retrieval algorithm with a global channel and a local channel in parallel. The two channels search for different features of the document respectively, and then dynamically integrate the retrieval results through the attention fusion mechanism. The retrieval module 204 implements the query understanding pre-processing function, analyzes the type, intent and key concepts of the query, and provides guidance for the subsequent retrieval strategy selection. The global channel focuses on capturing the structural features and global semantics of the document, and performs retrieval based on the document representation stored in the chapter-level memory pool 2021. This channel adopts a multi-stage retrieval strategy: first, query expansion is performed to enrich the original query through synonym replacement, entity linking and concept generalization; then the inverted index is used for preliminary screening to quickly locate potentially related document paragraphs; finally, an exact match is performed to calculate the semantic similarity between the query and the candidate paragraph. For query q, the global channel calculates its similarity with each paragraph p in the memory pool through a deep interaction model. The model not only considers the similarity between global representations, but also analyzes the fine-grained interaction pattern between query terms and paragraph terms to capture more refined matching signals. In order to improve retrieval efficiency, the global channel implements a multi-level index structure and parallel computing optimization to significantly reduce the retrieval latency of large-scale documents.
[0054] The local channel focuses on specific semantic fragments and uses a dynamic sliding window method to perform fine-grained retrieval of documents. The channel adaptively adjusts the window size according to the query characteristics: for short and precise queries, a smaller window is used to improve accuracy; for complex descriptive queries, a larger window is used to ensure semantic integrity. The local channel implements a similarity calculation method based on kernel functions, which decomposes the query and window into multiple semantic units, calculates the similarity matrix between units, and then extracts the most significant matching patterns through kernel functions. For query q, the local channel calculates its comprehensive similarity with each window w in the document, by taking the sum of the similarities between each query token and the most similar token in the window, and applying position weighting, so that the matching patterns of adjacent words in the query receive additional rewards, thereby evaluating the overall relevance of the query and the window. In addition, the local channel also implements semantic alignment enhancement technology, which improves the matching ability of the same semantic content in different expressions through word sense disambiguation and semantic role labeling.
[0055] Finally, the results of the dual-channel retrieval are merged through adaptive fusion, which takes into account multiple factors such as query characteristics, document type, and consistency of retrieval results. The core fusion formula is:
[0056] score(q, p) = λ·sim_global(q, p) + (1-λ)·max_{w∈p} sim_local(q, w)
[0057] Where λ is the global channel weight, and the retrieval module 204 dynamically adjusts the λ value according to the query analysis results: for conceptual queries (such as "what is climate change"), the global channel weight is increased to favor macro explanations; for detailed queries (such as "what is the global average temperature in 2019"), the local channel weight is increased to focus on specific facts. In addition, the fusion process also includes result diversity optimization. Through the maximum marginal correlation algorithm, while ensuring relevance, the information coverage of the retrieval results is increased to avoid returning highly repetitive paragraphs. The retrieval module 204 finally returns the top k paragraphs with the highest scores as relevant content, and attaches metadata (such as source location, credibility score) to assist subsequent processing. At the same time, the retrieval module 204 maintains the query-result cache to speed up the response speed of repeated or similar queries and improve the overall efficiency of the system.
[0058] The content generation module 205 includes two key components: a memory pointer network and a consistency checker, which work together through a tight feedback loop to ensure that the generated content is both smooth, coherent and factual. The memory pointer network generates content, while the consistency checker supervises and corrects the generation process to prevent factual errors and logical contradictions. The memory pointer network enhances the coherence and accuracy of the generated content by dynamically referencing historical context. The network adopts an encoder-decoder architecture and adds a memory enhancement mechanism on the basis of the traditional architecture. The network includes three main parts: the encoder converts the input query into a semantic representation; the decoder generates new content based on the query representation and the generated content; and the memory pointer provides the ability to retrieve and copy content from external memory. The first layer of the encoder processes the surface representation of the query, and the second layer of the encoder integrates the retrieval results through a cross-attention mechanism to generate a knowledge-enhanced query representation. The decoder uses an autoregressive generation method to predict the next token at each time step based on the generated content and query representation. To handle long sequence generation, the decoder implements a block processing mechanism to decompose the long sequence into multiple interrelated short sequences to reduce the accumulation of errors in the generation process.
[0059] The memory pointer network dynamically decides at each generation time step whether to generate a new token or copy existing content from the memory pool. This decision is performed by the generation controller, which calculates the generation probability p_gen based on the characteristics and needs of the current context, which represents the tendency to generate new tokens from the model vocabulary. For situations where the original text needs to be accurately quoted (such as citing professional terms, numbers, or proper nouns), the controller reduces the generation probability and tends to copy content from the memory pool; for content that requires generalization or reasoning, the generation probability is increased, relying on the model's language generation ability. If it is decided to copy content, the memory pointer calculates the semantic match between the current decoder hidden state and each entry in the memory pool through the content addressing mechanism to identify the most relevant content fragments. To improve the accuracy of copying, the memory pointer implements a multi-granularity attention mechanism, calculating attention scores at the token level, phrase level, and sentence level, and then determining the final copy target through weighted combination. In addition, the memory pointer also maintains historical copy records to ensure the coherence of the copied content and avoid content fragmentation caused by repeated switching between different sources. The final token generation probability distribution is a weighted combination of the generation distribution and the replication distribution. The network samples or selects the highest probability token as output based on this distribution, gradually building the complete generated content.
[0060] The consistency check module monitors and verifies the consistency of content in real time during the generation process to prevent factual errors, logical contradictions and semantic incoherence in the generation process. The module adopts a combination of prevention and correction strategies: setting constraints before generation, real-time verification during generation, and comprehensive verification after generation. The consistency check covers three core aspects: entity consistency, logical relationship consistency and timeline consistency. The entity consistency check ensures that the entity information in the generated content is consistent with the entity state stored in the gated memory unit 2022. The check identifies entity mentions in the generated content through entity linking technology, and then queries the canonical representation and attribute state of the corresponding entity. Entity checking supports multi-level verification: the first-level verification checks whether the entity itself exists in memory; the second-level verification checks whether the entity attributes are consistent with the state in memory; the third-level verification checks whether the relationship between entities conforms to the relationship network stored in memory. At each generation step, if the current generated content involves a known entity, the entity consistency check component actively queries the latest state of the entity and ensures that the generated content is compatible with the entity state through a constraint decoding strategy. Constraint decoding adjusts the probability distribution of the next tag to increase the probability of tags that are consistent with the facts and reduce the probability of tags that are inconsistent with the facts, thereby preventing factual errors from the source.
[0061] Logical relationship verification evaluates the logical consistency between the generated content and the source document through natural language reasoning technology. This verification adopts a segmented reasoning strategy to decompose long text into multiple semantic units and verify their relationship with the source document one by one. For each complete sentence generated, the logical relationship verification component first retrieves the most relevant content fragment from the memory pool as the premise, and then applies a three-way classification model to calculate the three relationship probabilities between the premise and the generated sentence: implication (the premise supports the generated content), contradiction (the premise conflicts with the generated content), and neutrality (the premise has no direct relationship with the generated content). The logical verification model is fine-tuned based on the pre-trained language model, specifically enhancing the sensitivity to key facts such as numerical values, time, and entity relationships. If the probability of contradiction is detected to exceed the preset threshold, the logical relationship verification component triggers the content correction process: for minor contradictions, the expression is adjusted by local rewriting; for serious contradictions, it falls back to the beginning of the sentence and regenerates. In addition, logical verification also includes coherence checks to ensure that the generated content is logically consistent internally and avoids self-contradiction.
[0062] Timeline consistency verification specifically checks whether the time expression and event description in the generated content conform to the timeline maintained by the gated memory unit 2022. This verification extracts the time expression in the generated content through the time expression identifier and parses its standardized representation; identifies the events associated with the time expression through the event extractor; and then queries the records of related events in the timeline data of the gated memory unit 2022. The verification process checks consistency at three levels: consistency of event occurrence sequence (ensuring the correct sequence of events), consistency of event time points (ensuring the accuracy of specific time expressions), and consistency of event duration (ensuring the accuracy of event duration). For each time expression and associated event in the generated content, the timeline consistency verification component calculates the degree of deviation from the time record in memory. If there is an inconsistency in the time expression, but the deviation is within the allowable range, the verification component corrects the expression through fine-tuning; if the deviation exceeds the allowable range, the rewrite mechanism is triggered to replace the wrong time with the accurate time in memory. In addition, time verification also includes causal consistency checks to ensure the temporal logic of the event causal chain and avoid temporal paradoxes such as "the result precedes the cause". Through multi-level consistency checking, the content generation module 205 generates output content that is both fluent and natural and factually accurate, effectively solving the semantic coherence and factual consistency problems in long document processing.
[0063] In order to implement the adaptive window size adjustment method, the window size is automatically adjusted according to the complexity of the document content, and the sliding window transformer uses a fixed-size window to process documents with large changes in content complexity. The expansion scheme of this technology includes a complexity evaluation module and a window adjustment module, which are connected by a feedback control loop. The complexity evaluation module adopts a three-level feature extraction architecture, including a surface feature extraction layer, a linguistic feature extraction layer, and a semantic feature extraction layer. The surface feature extraction layer calculates the vocabulary diversity index, calculates the ratio of different words to the total number of words in a fixed-size text segment, and averages the results of multiple segments to eliminate the influence of text length on the evaluation. The linguistic feature extraction layer calculates the syntactic complexity index, which comprehensively considers three sub-indicators: average sentence length (in terms of the number of tokens), average depth of the syntactic tree, and density of subordinate clauses. The linguistic feature extraction layer obtains these sub-indicators through the dependency syntactic analyzer, and then applies normalization to synthesize the final syntactic complexity score. The semantic feature extraction layer calculates the semantic density index, which is evaluated by two dimensions: information entropy and keyword density. Information entropy is calculated based on the conditional probability of tags and is used to measure the uncertainty of text information. Keyword density identifies semantically important tags based on word frequency-inverse document frequency values and is used to calculate their distribution density in the text. The complexity evaluation module integrates the three types of indicators into the final complexity score through an adaptive weight combiner, and the weight parameters are dynamically adjusted according to the document type and processing task. In scientific documents and professional papers, the weight of lexical diversity is low and the weight of semantic density is high, emphasizing semantic understanding. In literary works, the weight of syntactic complexity is high, focusing on syntactic structure understanding. The complexity evaluation module converts the complexity evaluation results into window size adjustment instructions through a piecewise mapping function. The function is designed as a piecewise linear mapping, which uses gradual adjustment in the medium complexity range and step adjustment in the extremely high or extremely low complexity range to ensure a rapid response to significant complexity changes. The window adjustment module receives the adjustment instructions to achieve smooth changes in window size and reduce the risk of encoding incoherence caused by sudden changes in the window size of adjacent paragraphs. The window adjustment module weighted averages the target window size of the current paragraph with the actual window sizes of the previous paragraphs to generate the final window size used. The window adjustment module also performs key semantic unit protection to ensure that key semantic units (such as complex clauses or entity descriptions) are not improperly segmented due to window adjustment, and fine-tunes the window boundary position when necessary to align it with the semantic unit boundary. Through this adaptive adjustment process, this technical extension solution expands the context window to 1.5-2 times the baseline size when processing complex document paragraphs, and shrinks it to 0.6-0.8 times the baseline size when processing simple content, while maintaining the quality of semantic understanding and improving processing efficiency by 25%-40%.
[0064] In order to realize the hierarchical multi-granularity memory pool, store document representations of different granularities at the same time, improve retrieval flexibility, and solve the problem of single information granularity in the chapter-level memory pool 2021, the memory pool adopts a pyramid structure design, with four granularity levels from the bottom to the top, namely the tag level, sentence level, paragraph level and chapter level, and each layer maintains independent content index and representation storage. Each granularity layer uses a key-value-metadata triple structure to organize memory entries. The key is used for fast retrieval, the value stores the actual semantic representation, and the metadata contains management information such as source location, creation time, and access frequency. The relationship between layers is maintained through bidirectional links. Each high-level representation records the index of all low-level elements it contains, and each low-level element records the index of the high-level unit to which it belongs, forming a complete inclusion and subordination relationship graph. The tag level stores the original tag representation, retains the precise information at the vocabulary level, and is suitable for precise matching and detail retrieval. The sentence level stores sentence representations, captures basic semantic units, and is suitable for answer extraction and summary generation. The paragraph level stores paragraph representations, which contain topic and opinion information and are suitable for similar content retrieval. The chapter level stores chapter representations, which reflect macro structures and themes and are suitable for document classification and topic identification. The representation generation of this technical extension adopts a bidirectional aggregation process. In the bottom-up aggregation process, the token-level representation is first converted to a sentence representation through context-weighted pooling. The sentence representation is fused into a paragraph representation through a topic-aware attention mechanism. The paragraph representation is integrated into a chapter representation through a structure-aware encoder. In the aggregation process, this technical extension implements a semantic importance filter, which identifies the contribution of each low-level element to the high-level semantics through a self-attention mechanism, suppresses elements with small contributions, and ensures the information concentration of the high-level representation. In the top-down feedback process, this technical extension adopts a conditional generation network, using the high-level representation as a condition to adjust and enhance the low-level representation. For example, the topic information at the chapter level is injected into the paragraph representation through the attention gating unit, so that the paragraph representation contains both local details and global context information. The feedback mechanism adopts a residual connection design to ensure that the original information is not lost, and the enhanced low-level representation contains both the original information and the high-level semantic information. The hierarchical multi-granularity memory pool supports granular adaptive retrieval, and the granularity selection controller determines the most suitable starting retrieval granularity based on the query complexity, length, and specificity. For short and specific queries (such as "What is quantum entanglement"), the granularity selection controller starts the retrieval directly from the paragraph level. For complex descriptive queries, the granularity selection controller starts from the chapter level and goes down layer by layer. The retrieval process adopts the "coarse retrieval-fine retrieval" strategy: first, a broad search is performed at the selected starting granularity layer to obtain candidate results; then, within the coverage of these candidates, it goes down to a finer granularity layer for precise matching, and finally returns the most matching content fragment. This technical extension solution filters and sorts multi-granularity retrieval results through semantic consistency scoring to ensure the coherence and completeness of the returned content and reduce the risk of semantic breaks caused by fragment splicing.
[0065] In order to realize the graph-enhanced paragraph-level attention algorithm, a document semantic graph is constructed to comprehensively capture various implicit and explicit relationships between paragraphs, and solve the problem that the paragraph-level attention algorithm only considers the direct relationship between paragraphs and ignores the complex semantic connections in the document. The algorithm consists of a graph construction module, a graph convolution module, and an attention reading module. The graph construction module first represents the document paragraphs as nodes of the graph, each of which contains the initial semantic vector of the paragraph, structural position information, and content type label. The graph construction module identifies four types of relationships between paragraphs through a multi-channel relationship detector: similarity relationship, reference relationship, sequential relationship, and hierarchical relationship, and represents them as edges of the graph. Similarity relationship detection is achieved by calculating the cosine similarity of the paragraph representation vector, and an adaptive threshold is applied to filter, and only paragraph pairs with significant similarity are retained. Reference relationship detection combines rule recognition (based on explicit reference tags such as "as mentioned above" and "see paragraph X") and learning recognition (based on deep learning reference relationship classifiers to identify implicit references). Sequential relationship detection is based on narrative flow analysis to identify the temporal, causal, and logical inheritance relationships between paragraphs. Hierarchical relationship detection is based on document structure analysis to identify hierarchical subordinate relationships such as title and content, overview and details. Each relationship type is represented by an independent edge type, with a dedicated feature vector describing the nature and strength of the relationship. The constructed document semantic graph is a multi-relation directed graph, where nodes represent paragraphs and multiple types of edges represent different types of paragraph relationships. The graph convolution module performs message passing and information aggregation on the constructed semantic graph, allowing each paragraph to integrate other paragraph information related to it. The graph convolution module implements a relationship-aware graph convolution network, using independent parameter matrices for different types of relationships, and performs multiple rounds of information passing and updating. The graph convolution process iterates for multiple rounds (usually 3-5 rounds), and each round updates the representation of all nodes. To avoid the problem of over-smoothing, the graph convolution module applies residual connections and gating mechanisms after each round of convolution to dynamically control the proportion of original information retained. After multiple rounds of convolution, the representation of each paragraph node has integrated the information transmitted by multi-hop relationships, which can reflect the semantic role and importance in the overall structure of the document. The attention reading module uses a task-oriented attention mechanism to aggregate information in the graph in a targeted manner according to the specific task requirements currently being processed. The attention reading module trains a dedicated attention query vector for each task type (such as summary generation, question answering, and content inference), calculates the attention weight distribution based on the correlation between the query vector and the nodes in the graph, and then weighted aggregates the node representations to obtain the final graph representation. To further improve the reading efficiency, the attention reading module implements a multi-head attention mechanism, where each attention head focuses on different semantic aspects, and then integrates the multi-head results to obtain a comprehensive graph representation.Through the graph-enhanced paragraph-level attention algorithm, this technical extension enables the system to transcend the limitations of linear text order and establish a rich connection network between paragraphs. Especially for long documents with complex structures, it can effectively capture long-range semantic dependencies and implicit topic associations, and improve the ability to understand the overall structure of the document. Compared with the basic paragraph-level attention algorithm, the graph-enhanced version has a 15%-25% increase in accuracy in long document understanding tasks.
[0066] In order to integrate the external knowledge graph with the gated memory unit 2022, build a knowledge-enhanced entity state management system, and solve the problem that the gated memory unit 2022 only relies on document content to manage entity states and lacks external knowledge support, the system consists of four parts: entity linking module, knowledge acquisition module, knowledge fusion module, and consistency verification module. The entity linking module is used to accurately link the entities identified in the document to the corresponding entities in the external knowledge graph (such as Wiki database, concept network or domain-specific knowledge base). The entity linking module adopts a two-stage linking strategy: the candidate generation stage generates potential entity link candidates through character matching, alias expansion and fuzzy matching; the candidate sorting stage calculates the matching degree between the entity mention in the document and the knowledge base entity through a context-aware deep sorting model, and selects the most matching entity as the link target. The entity linking module realizes zero-sample linking capability, matching through the semantic similarity of entity descriptions, rather than relying solely on entity names, to handle uncommon entities in professional fields. The entity linking module maintains the entity co-reference resolution component during the linking process to ensure that different expressions (such as full name, abbreviation, pronoun) referring to the same entity in the document can be linked to the same knowledge graph entity. The knowledge acquisition module retrieves relevant knowledge from the knowledge graph based on the linked entity ID and builds a local knowledge subgraph. The knowledge acquisition module implements a multi-hop query engine, which not only obtains the direct attributes of the entity (such as date of birth, nationality, occupation, etc.), but also obtains multi-hop relationships (such as the founder of the organization to which the person belongs) and the relationship path between entities. The knowledge acquisition module designs a relevance filter to filter the most relevant knowledge based on the current document topic and the importance of the entity in the document, build a compact local knowledge graph, and prevent knowledge explosion. The knowledge fusion module integrates the entity information extracted from the document with the information obtained from the external knowledge graph to generate an enhanced entity representation. The knowledge fusion module implements a three-level fusion strategy: attribute-level fusion aligns the entity attributes mentioned in the document with the attributes in the knowledge graph to form a complete attribute set; relationship-level fusion integrates the relationship information between entities, including both the relationships explicitly described in the document and the implicit relationships in the knowledge graph; semantic-level fusion uses a graph neural network to fuse the structural embedding of the entity in the knowledge graph with the contextual representation in the document to generate a knowledge-enhanced entity semantic representation. The fusion process of the knowledge fusion module uses the attention mechanism to dynamically adjust the weights of information from different sources. When the document information and the knowledge graph information are highly consistent, the weights of the two are balanced; when there is an inconsistency, a more credible source tends to be selected. The consistency verification module implements multi-level knowledge consistency verification to ensure that entity status updates comply with external knowledge constraints.The consistency verification module contains three types of validators: the attribute constraint validator checks whether the entity attribute update complies with the attribute invariance constraint (such as the person's date of birth should not change), the valid value constraint (such as the age should not be negative) and the change rate constraint (such as the age increases by 1 year every year); the relationship constraint validator checks whether the entity relationship update complies with the logical constraints such as the symmetry and transitivity of the relationship; the state transition validator checks whether the entity state change follows the predetermined state transition rules (such as the person cannot directly change from "alive" to "deceased" and skip the intermediate states such as "sick"). When it is detected that the update operation violates the knowledge constraint, the consistency verification module calculates the degree of inconsistency and adjusts the value of the update gate accordingly to suppress updates that do not comply with the knowledge constraint and ensure the consistency and accuracy of the entity state. When a serious knowledge conflict is detected (the degree of inconsistency exceeds the threshold), the consistency verification module also generates a conflict report for the advanced error correction mechanism to handle. By integrating the external knowledge graph, the technical extension scheme enables the gated memory unit 2022 to obtain richer background knowledge and stricter consistency constraints, significantly improving the accuracy and robustness of entity state management, especially when processing documents involving complex entity relationships and professional domain knowledge.
[0067] In order to realize the emotion-aware dual-channel retrieval algorithm, the emotion analysis is deeply integrated with the traditional retrieval to improve the contextual adaptability in the retrieval of emotion-rich content and solve the problem that the dual-channel retrieval algorithm is mainly based on semantic similarity calculation and ignores the emotion and opinion factors. The algorithm consists of four parts: emotion representation module, multi-dimensional similarity calculation module, adaptive fusion module and emotion coherence control module. The emotion representation module is used to extract the emotion features of the query and document paragraphs and construct the emotion representation vector. The emotion representation module adopts a hierarchical emotion analysis architecture, including three levels: word-level emotion recognition, sentence-level emotion aggregation and paragraph-level emotion summarization. The word-level layer uses the emotion dictionary and the context-aware word vector model to identify the emotion vocabulary and its polarity strength. The sentence-level layer integrates the word-level emotion and identifies the complex emotion expression patterns (such as irony, assumption and transition) through the recursive neural network or transformer encoder. The paragraph-level layer comprehensively analyzes the emotion distribution of the sentence group and extracts the overall emotion tendency and emotion change pattern. The emotion representation finally generated by the emotion representation module is a multi-dimensional vector, which includes dimensions such as emotion polarity (positive / negative / neutral), emotion intensity, emotion category (such as joy, anger, sadness, etc.) and emotion object, and comprehensively describes the emotional characteristics of the content. The multi-dimensional similarity calculation module calculates the similarity between the query and the document paragraph in multiple dimensions based on semantic representation and emotion representation. The multi-dimensional similarity calculation module implements the semantic-emotion cross-attention network, which calculates both the direct matching of semantics-semantics and emotion-emotion, and the cross-matching of semantics-emotion (such as the matching degree between the semantic content of the query and the emotional expression of the document). The multi-dimensional similarity calculation module introduces emotion matching calculation for the global channel and the local channel respectively, combines the original semantic similarity with the emotion similarity, and forms an enhanced similarity measure. The emotion matching weight is dynamically set by the task type adapter according to the characteristics of the current query. For queries that explicitly contain emotional requirements (such as "recommend some inspiring stories"), the emotion matching weight will be significantly increased; for pure factual queries, the influence of emotional factors will be reduced. The adaptive fusion module is used to integrate semantic similarity and emotion similarity to generate the final retrieval score. The adaptive fusion module implements a context-aware dynamic weight allocation strategy. Through a meta-learning method, it predicts the optimal weight configuration based on the query type, query history, and dialogue state. The fusion process is not limited to simple linear combinations, but uses a deep interactive network to capture the nonlinear complementary relationship between semantic and emotional factors. The emotional coherence control module is used to maintain emotional coherence during multiple rounds of retrieval or generation, ensuring that the retrieval results are coordinated with the contextual emotional flow. The emotional coherence control module maintains a dynamic emotional state tracker, records the emotional change trajectory in historical interactions, and predicts a reasonable emotional development direction. For the task of generating multiple rounds of dialogue or continuous paragraphs, the emotional coherence control module implements an emotional smoothness evaluation function to calculate the degree of coherence between the emotions of the candidate retrieval results and the historical emotional trajectory.The emotional coherence control module performs gradual control of emotional changes to avoid sudden changes in emotions (such as jumping directly from highly positive to highly negative). When the semantic similarity of multiple candidate results is close, the emotional coherence control module will tend to select results with higher emotional smoothness, which is achieved by adjusting the final ranking score. For generation tasks that require emotional ups and downs, the emotional coherence control module also implements an emotional arc planner, which intentionally guides emotions to develop in a specific pattern (such as ups and downs, progression, and contrast) according to the needs of the narrative structure, thereby enhancing the expressiveness and appeal of the generated content. Through the emotion-aware dual-channel retrieval algorithm, this technical extension solution enables the system to provide emotionally matched retrieval results while ensuring content relevance, significantly improving the user experience, especially in emotion-driven interactive scenarios (such as psychological counseling, emotional companionship, and creative writing). In the emotion-intensive dialogue test, this technical extension solution increased user satisfaction by 35%-50%.
[0068] The long document semantic coherence processing method and device of the present application realize adaptive processing of content complexity perception, multi-granularity hierarchical memory representation, deep semantic modeling based on graph structure, knowledge-enhanced entity state management and emotionally aware intelligent retrieval capabilities through the above-mentioned technical expansion scheme. Through the adaptive window size adjustment method, the processing window size is dynamically adjusted according to the complexity of the document content to achieve efficient allocation of computing resources. Through the hierarchical multi-granularity memory pool, the document representation is stored and managed at different abstract levels to meet the retrieval needs of various types of queries. Through the graph-enhanced paragraph-level attention algorithm, the complex relationship between paragraphs in the document is captured to deepen the understanding of the overall structure of the document. By integrating the gated memory unit 2022 of the external knowledge graph, the accuracy and coherence of the entity state representation are enhanced, and the factual consistency of the generated content is improved. Through the emotionally aware dual-channel retrieval algorithm, the emotional dimension consideration is added on the basis of traditional semantic matching to make the generated content more in line with the user's emotional needs. These technical expansion schemes work together to comprehensively improve the depth of semantic understanding and the quality of content generation when processing complex documents, multi-round dialogues and knowledge-intensive tasks.
[0069] In the long document semantic coherence processing method, device, storage medium and computer equipment provided in the present application, a sliding window transformer is used to encode the long document in blocks, and the sliding window transformer includes a window division unit 2011, a position encoding unit 2012 and a transformer encoding unit 2013, wherein the window division unit 2011 is different from the traditional fixed-size block strategy, but adjusts the block position according to the semantic boundaries of sentences and paragraphs, and sets an overlapping area between adjacent blocks, so that the block scoring process can respect the natural semantic units of the text, avoid the situation of dividing semantically complete sentences or paragraphs into different blocks, and thus prevent semantic fragmentation at the source. At the same time, the overlapping area between adjacent blocks provides a bridge for semantic transmission between blocks, so that the semantic information across blocks can be retained and transmitted, fundamentally improving the semantic coherence of long document processing. In addition, since the encoded document content is stored and managed through the memory management module 202, the module includes a chapter-level memory pool 2021 and a gated memory unit 2022. The chapter-level memory pool 2021 stores the document encoding content and provides a unified information access interface for subsequent processing; the gated memory unit 2022 maintains the entity and event timeline, tracks the state changes of key entities in the document and the development of events, so that the system can maintain a semantic understanding of the document content for a long time, which not only solves the semantic incoherence problem caused by the isolation of information between blocks in traditional methods, but also provides support for semantic associations with a larger span, and enhances the system's ability to grasp the overall semantics of long documents. In addition, since the semantic processing module 203 captures multiple dependencies between paragraphs by calculating the multi-dimensional attention weights between paragraphs, it realizes a deep semantic understanding of the document content. Compared with the limitations of traditional methods that only focus on local semantics, the semantic processing module 203 can identify and model complex semantic relationships between paragraphs, including topic coherence, logical dependencies, reference relationships, etc. In this way, the understanding of the document is no longer limited to independent text blocks, but a semantic network that runs through the entire text is formed, which significantly improves the accuracy and coherence of semantic understanding. In addition, since the content generation module 205 includes a memory pointer network and a consistency check component, the semantic coherence and factual accuracy of the output content are jointly guaranteed. The memory pointer network makes full use of existing information when generating content by dynamically referencing historical contexts, avoiding content incoherence caused by information loss in traditional methods; the consistency check component ensures the factual accuracy of the output content by verifying the consistency of the generated content with the source document. The collaborative work of these two components enables the system to generate content that is both coherent and fluent and faithful to the original text, effectively solving the problem of semantic breaks in long document processing.
[0070] In terms of processing efficiency, the sliding window transformer of the present application reduces the computational complexity of global information processing by combining local and global encoding; the hierarchical storage mechanism of the memory management module 202 improves the efficiency of information retrieval; the multi-dimensional attention mechanism of the semantic processing module 203 optimizes the calculation of long-distance semantic associations; the retrieval module 204 and the content generation module 205 reduce redundant calculations by reusing information. The combined effect of these optimization measures enables the entire system to significantly improve the computational efficiency while maintaining high-quality semantic processing, meeting the needs of real-time interaction.
[0071] The embodiments of the present application are described in detail above, and the contents of this specification should not be construed as limiting the scope of protection of the present application.
Claims
1. A method for processing semantic coherence of a long document, comprising: Chunk encoding of long documents; Store and manage encoded document content; Semantic understanding of document content; Provide search results of document contents; as well as Generate output content; Characterized in that the method further comprises: The long document is encoded in blocks by a sliding window transformer, wherein the sliding window transformer comprises a window division unit, a position encoding unit and a transformer encoding unit, wherein the window division unit is used to adjust the block position according to sentence and paragraph boundaries and set an overlapping area between adjacent blocks; The encoded document content is stored and managed by a memory management module, wherein the memory management module includes a chapter-level memory pool and a gated memory unit, wherein the chapter-level memory pool is used to store the document encoding content, and the gated memory unit is used to maintain the entity and event timeline; Performing deep semantic understanding of document content through a semantic processing module, wherein the semantic processing module is used to capture multiple dependencies between paragraphs by calculating multi-dimensional attention weights between paragraphs; Providing retrieval results of document contents through retrieval module; Semantically coherent output content is generated through a content generation module, wherein the content generation module includes a memory pointer network and a consistency check component, wherein the memory pointer network is used to enhance the coherence and accuracy of the generated content by dynamically referencing historical contexts, and the consistency check component is used to ensure that the generated content is consistent with the facts of the source document.
2. The method according to claim 1, characterized in that The overlapping length of the overlapping area is O marks, where O is less than L. The window division unit is used to create a window containing the current block and its context, which includes the last O marks of the previous block, all marks of the current block and the first O marks of the next block, where O and L are both positive integers.
3. The method according to claim 1, characterized in that The position encoding unit is used to assign a local position code and a global position code to each tag in the window, the local position code is used to map the serial position of the tag in the current window to a continuous encoding space, and the global position code is used to represent the absolute position of the tag in the entire document.
4. The method according to claim 1, characterized in that: The gated memory unit includes an entity recognition module, a state tracking module and a timeline maintenance module. The entity recognition module is used to identify key entities and events from document content, the state tracking module is used to update entity status using a gated update method, and the timeline maintenance module is used to construct and update the timeline of entity-related events.
5. A device for processing semantic coherence of a long document, characterized in that: include: A document encoding module, used for encoding a long document in blocks by a sliding window transformer, wherein the sliding window transformer comprises a window division unit, a position encoding unit and a transformer encoding unit, wherein the window division unit is used for adjusting the block position according to sentence and paragraph boundaries and setting an overlapping area between adjacent blocks; A memory management module, used to store and manage the encoded document content, the memory management module includes a chapter-level memory pool and a gated memory unit, the chapter-level memory pool is used to store the document encoding content, and the gated memory unit is used to maintain the entity and event timeline; A semantic processing module, used for deep semantic understanding of document content, wherein the semantic processing module is used for capturing multiple dependencies between paragraphs by calculating multi-dimensional attention weights between paragraphs; A retrieval module, used to provide retrieval results of document content; A content generation module is used to generate semantically coherent output content. The content generation module includes a memory pointer network and a consistency check component. The memory pointer network is used to enhance the coherence and accuracy of the generated content by dynamically referencing historical contexts, and the consistency check component is used to ensure that the generated content is consistent with the facts of the source document.
6. The device according to claim 5, characterized in that The overlapping length of the overlapping area is O marks, where O is less than L. The window division unit is used to create a window containing the current block and its context, which includes the last O marks of the previous block, all marks of the current block and the first O marks of the next block, where O and L are both positive integers.
7. The device according to claim 5, characterized in that The position encoding unit is used to assign a local position code and a global position code to each tag in the window, the local position code is used to map the serial position of the tag in the current window to a continuous encoding space, and the global position code is used to represent the absolute position of the tag in the entire document.
8. The device according to claim 5, characterized in that The gated memory unit includes an entity recognition module, a state tracking module and a timeline maintenance module. The entity recognition module is used to identify key entities and events from document content, the state tracking module is used to update entity status using a gated update method, and the timeline maintenance module is used to construct and update the timeline of entity-related events.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the steps of the method for processing semantic coherence of a long document according to any one of claims 1 to 4 are implemented.
10. A computer device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor implements the steps of the method for processing semantic coherence of a long document according to any one of claims 1 to 4 when executing the computer program.
Citation Information
Patent Citations
Long text reading understanding method based on context memory
CN115293171A
Dual-prevention intelligent interaction system based on large language model
CN119579365A