Ultra-long sequence data processing method based on block annular attention mechanism
By employing a block-based circular attention mechanism, dynamic block division, sparse attention, and circular communication are used to optimize the processing of ultra-long sequence data, solving the problems of high memory complexity and communication bottlenecks, and achieving efficient processing of multimodal data.
Patent Information
- Application Number
- CN202510984778.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies suffer from high memory complexity, severe communication bottlenecks, and insufficient multimodal adaptation when processing ultra-long sequence data, making it difficult to efficiently process mixed long sequence data such as video and text.
We adopt a block-based circular attention mechanism, which optimizes computational complexity and memory usage through dynamic block division, sparse attention, circular communication and elastic memory management. We also design a multimodal adaptation strategy to achieve efficient modeling of ultra-long sequences.
It significantly reduces memory usage and communication overhead, supports efficient joint modeling of multimodal data, and improves the efficiency and performance of processing ultra-long sequences.
Smart Images

Figure CN120974398A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, and in particular relates to a method for processing ultra-long sequence data based on a block-based circular attention mechanism. Background Technology
[0002] With the development of artificial intelligence technology, large models need to process extremely long sequences of data such as videos and language to simulate the dynamic patterns of the real world. However, existing technologies face the following problems: Memory limitations: The traditional Transformer attention mechanism has a memory complexity of O(N^2). 2 This approach struggles to handle sequences exceeding tens of thousands of tokens. Communication bottlenecks exist: distributed solutions (such as RingAttention) require full transmission of the key-value matrix (KV Cache), resulting in high communication overhead. Multimodal adaptation is also insufficient: existing methods do not optimize the joint processing of mixed long sequences such as video and text. Therefore, this invention proposes an attention mechanism combining block processing, ring communication, and flexible memory management to achieve efficient modeling of ultra-long sequences. Summary of the Invention
[0003] The purpose of this invention is to provide a method for processing ultra-long sequence data based on a block-based circular attention mechanism to solve the above-mentioned technical problems.
[0004] To address the aforementioned technical problems, the specific technical solution of the ultra-long sequence data processing method based on the block-based circular attention mechanism of the present invention is as follows:
[0005] A method for processing ultra-long sequence data based on a block-based circular attention mechanism includes the following steps:
[0006] Step 1: Construct a sequence segmentation strategy;
[0007] Step 2: Calculate hierarchical attention;
[0008] Step 3: Build a memory management mechanism;
[0009] Step 4: Design a multimodal adaptation strategy;
[0010] Step 5: Optimize training and inference.
[0011] Furthermore, step 1 includes the following steps:
[0012] Step 1.1: Dynamic Block Algorithm:
[0013] Adaptive block partitioning: dynamically adjusts the size of sub-blocks based on the modal characteristics of the input data;
[0014] Video data: Divide the video stream into sub-blocks according to the time step to ensure temporal continuity;
[0015] Text data: Divide into semantic blocks to avoid truncating sentences or paragraphs;
[0016] Load balancing: A greedy algorithm is used to analyze the computational complexity of each sub-block and dynamically adjust the block size to ensure that the processing time of each device is similar.
[0017] Step 1.2: Boundary Processing:
[0018] Overlapping blocks: Adjacent sub-blocks retain a 10% overlap area, and an attention masking mechanism is used to prevent information leakage in the overlapping area;
[0019] Filling strategy: For sub-blocks that are not long enough, use zero-fill or cyclic filling to ensure that all sub-blocks have the same length.
[0020] Furthermore, step 2 includes the following steps:
[0021] Step 2.1: Local attention optimization:
[0022] Sparse attention: Using sliding window attention within each sub-block reduces computational complexity from O(N^2) to O(N^2). 2 ) decreased to O(N).
[0023] Memory compression: The local key-value cache (KV Cache) is quantized to 8-bit, reducing video memory usage by 50% and controlling the inversion error to within 0.1%;
[0024] Step 2.2: Global Attention Aggregation:
[0025] Digest generation algorithm: Mean digest: Calculate the mean of the KV cache of the sub-blocks along the sequence dimension to generate a 32-dimensional digest vector;
[0026] Low-rank projection: The KV cache is compressed to 32 dimensions through PCA, and the reconstruction error is controlled within 5%; Ring communication protocol: A bidirectional ring network is built between devices, and the communication latency is less than 1ms;
[0027] An asynchronous transmission mechanism is used: while computing the current sub-block, a background thread pre-transmits the digest of the next sub-block to the adjacent device.
[0028] Furthermore, step 3 includes the following steps:
[0029] Step 3.1: Tiered Storage Architecture
[0030] L1 memory layer: caches the digests of the currently active sub-block and adjacent sub-blocks, with a hit rate >95%;
[0031] L2 memory layer: Stores sub-blocks that may be accessed in the future, using an LRU eviction policy;
[0032] Storage layer L3: Stores large-scale inactive sub-blocks and supports direct SSD reading;
[0033] Step 3.2: Dynamic loading strategy:
[0034] Prefetch algorithm: Based on the attention weight distribution, predict the sub-blocks to be accessed in the future and load them into video memory in advance; Unloading trigger condition: When the video memory usage rate is >90%, migrate the least used sub-blocks to memory or storage layer.
[0035] Furthermore, step 4 includes the following steps:
[0036] Step 4.1: Unify the embedding space:
[0037] Video modality: Spatiotemporal features are extracted using a 3D convolutional network and mapped to the same embedding dimension as the text;
[0038] Text modality: Generate token embeddings using a standard tokenizer;
[0039] Cross-modal alignment: Constraining consistency between video and text embedding spaces through contrastive learning;
[0040] Step 4.2: Modal Mixed Attention:
[0041] Cross-attention gating: Dynamically adjusts the attention weights of video and text tokens.
[0042] Furthermore, step 5 includes the following steps:
[0043] Step 5.1: Distributed Training Strategy:
[0044] Data parallelism: Batch data is sharded and distributed to different devices, and gradients are synchronized through All-Reduce;
[0045] Model parallelism: Split the encoder and decoder onto different devices to reduce the load on a single card;
[0046] Step 5.2: Accelerating Reasoning:
[0047] Sub-block cache reuse: For recurring sub-blocks, cache their KV cache to avoid duplicate calculations;
[0048] Progressive decoding: In the generation task, sub-blocks with high attention weights are computed first, and the context window is gradually expanded.
[0049] The method for processing ultra-long sequence data based on the block-based circular attention mechanism of the present invention has the following advantages: Attached Figure Description
[0050] Appendix Figure 1This invention provides a process for processing ultra-long sequence data based on a block-based circular attention mechanism. Detailed Implementation
[0051] To better understand the purpose, structure, and function of this invention, the following detailed description of a method for processing ultra-long sequence data based on a block-based circular attention mechanism is provided in conjunction with the accompanying drawings.
[0052] The core of the ultra-long sequence data processing method based on the block-based circular attention mechanism of the present invention lies in efficiently processing ultra-long sequence data through the block-based circular attention mechanism, including the following steps:
[0053] Step 1: Construct a sequence segmentation strategy
[0054] Step 1.1: Dynamic Block Algorithm
[0055] Adaptive block partitioning:
[0056] The size of the sub-blocks is dynamically adjusted based on the modal characteristics of the input data.
[0057] Video data: Divide the video stream into sub-blocks according to time steps, for example, each sub-block contains 16 frames to ensure temporal continuity.
[0058] Text data: Divide into semantic blocks, for example, each sub-block contains 512 tokens, to avoid truncating sentences or paragraphs.
[0059] Load balancing:
[0060] A greedy algorithm is used to analyze the computational complexity (such as FLOPs) of each sub-block, and the block size is dynamically adjusted to ensure that the processing time of each device is similar.
[0061] Step 1.2: Boundary Processing
[0062] Overlapping blocks:
[0063] Adjacent sub-blocks retain a 10% overlap area (e.g., 5% overlap in video frames), and an attention masking mechanism is used to prevent information leakage in the overlapping area.
[0064] Filling strategy:
[0065] For sub-blocks that are too short, use zero padding or circular padding (such as copying the data at the end of the sequence) to ensure that all sub-blocks are of the same length.
[0066] Step 2: Calculate hierarchical attention
[0067] Step 2.1: Local Attention Optimization
[0068] Sparse attention:
[0069] Using sliding window attention (window size 64) within each sub-block reduces the computational complexity from O(N^2) to O(N^2). 2 ) decreased to O(N).
[0070] Memory compression:
[0071] By performing 8-bit quantization on the local key-value cache (KV Cache), the video memory usage is reduced by 50%, and the inquantization error is controlled within 0.1%.
[0072] Step 2.2: Global Attention Aggregation
[0073] Digest generation algorithm:
[0074] Mean digest: Calculate the mean of the KV cache of the sub-block along the sequence dimension to generate a 32-dimensional digest vector.
[0075] Low-rank projection: The KV cache is compressed to 32 dimensions using PCA, and the reconstruction error is controlled within 5%.
[0076] Ring communication protocol:
[0077] A bidirectional ring network can be built between devices (such as based on NVLink 3.0), with a communication latency of less than 1ms.
[0078] An asynchronous transmission mechanism is used: while computing the current sub-block, a background thread pre-transmits the digest of the next sub-block to the adjacent device.
[0079] Step 3: Build a memory management mechanism
[0080] Step 3.1: Tiered Storage Architecture
[0081] L1 memory layer: caches the digests of the currently active sub-block and adjacent sub-blocks, with a hit rate >95%.
[0082] Memory layer (L2): Stores sub-blocks that may be accessed in the future, using the LRU (Least Recently Used) eviction policy.
[0083] Storage layer (L3): Stores large-scale inactive sub-blocks and supports direct SSD reads (bandwidth ≥ 3GB / s).
[0084] Step 3.2: Dynamic Loading Strategy
[0085] Prefetch algorithm:
[0086] Sub-blocks to be accessed in the future are predicted based on the attention weight distribution and preloaded into video memory.
[0087] Uninstallation trigger conditions:
[0088] When the video memory usage rate is >90%, the least used sub-blocks are migrated to the memory or storage layer.
[0089] Step 4: Design a multimodal adaptation strategy
[0090] Step 4.1: Unify the embedding space
[0091] Video modality: Spatiotemporal features are extracted using a 3D convolutional network and mapped to the same embedding dimension as text (e.g., 768 dimensions).
[0092] Text modality: Generate token embeddings using standard tokenizers such as CLIP.
[0093] Cross-modal alignment: Constraining the consistency of video and text embedding spaces through contrastive learning (InfoNCE loss).
[0094] Step 4.2: Modal Mixed Attention
[0095] - Cross-attention gating:
[0096] The attention weights for video and text tokens are dynamically adjusted using the following formula:
[0097] α=σ(ω[v i ;t i ])
[0098] Where v i It's a video token, t i It is a text token.
[0099] Step 5: Optimize training and inference
[0100] Step 5.1: Distributed Training Strategy
[0101] Data parallelism: Batch data is sharded and distributed to different devices, and gradients are synchronized through All-Reduce.
[0102] Model parallelism: Split the encoder and decoder onto different devices to reduce the load on a single card.
[0103] Step 5.2: Accelerating Reasoning
[0104] Sub-block cache reuse: For recurring sub-blocks (such as video background frames), cache their KV cache to avoid duplicate calculations.
[0105] Progressive decoding: In the generation task, sub-blocks with high attention weights are computed first, and the context window is gradually expanded.
[0106] Through the above embodiments, the present invention can significantly reduce the memory usage and communication overhead of ultra-long sequence processing, while supporting efficient joint modeling of multimodal data.
[0107] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.
Claims
1. A method for processing ultra-long sequence data based on a block-wise ring-shaped attention mechanism, characterized in that, Comprising the following steps: Step 1: Construct sequence chunking strategy; Step 2: Calculate hierarchical attention; Step 3: Build memory management mechanism; Step 4: Design multi-modal adaptation strategy; Step 5: Optimize training and inference.
2. The method of claim 1, wherein, The step 1 comprises the following steps: Step 1.1: Dynamic chunking algorithm: Adaptive chunking: dynamically adjust the size of sub-chunks according to the modal characteristics of input data; Video data: divide video stream into sub-chunks by time step, ensure time continuity; Text data: chunk by semantic paragraph, avoid truncating sentences or paragraphs; Load balancing: use greedy algorithm to analyze the computational complexity of each sub-chunk, dynamically adjust the chunk size to ensure similar processing time for each device. Step 1.2: Boundary processing: Overlapping chunking: adjacent sub-chunks retain 10% overlap area, and use attention mask mechanism to avoid information leakage in overlapping area; Padding strategy: for sub-chunks with insufficient length, use zero padding or cyclic padding to ensure consistent length of all sub-chunks.
3. The method of claim 1, wherein, The step 2 comprises the following steps: Step 2.1: Local attention optimization: Sparse attention: sliding window attention is adopted within each sub-block, reducing the computational complexity from O(N 2 ) to O(N). Memory compression: 8-bit quantization of local key-value cache KV Cache, memory occupancy reduced by 50%, inverse quantization error controlled within 0.1%; Step 2.2: Global attention aggregation: Summary generation algorithm: mean summary: take the mean of KV Cache along the sequence dimension to generate a 32-dimensional summary vector; Low-rank projection: compress KV Cache to 32 dimensions through PCA, reconstruction error controlled within 5%; Ring communication protocol: build a bidirectional ring network between devices, communication delay less than 1ms; Use asynchronous transmission mechanism: while computing the current sub-chunk, the background thread pre-transmits the summary of the next sub-chunk to the adjacent device.
4. The method of claim 1, wherein, The step 3 comprises the following steps: Step 3.1: Hierarchical storage architecture L1: cache current active sub-chunks and adjacent sub-chunk summaries, hit rate > 95%; L2: store future possible access sub-chunks, use LRU eviction policy; L3: store large-scale inactive sub-chunks, support SSD direct reading; Step 3.2: Dynamic loading strategy: Prefetching algorithm: predict future access sub-chunks based on attention weight distribution, load into video memory in advance; Unloading trigger condition: when the video memory usage rate > 90%, migrate the least used sub-chunk to memory or storage layer.
5. The method of claim 1, wherein, The step 4 comprises the following steps: Step 4.1: Unified embedding space: Video modal: extract spatiotemporal features through 3D convolutional network and map to the same embedding dimension as text; Text modal: use standard tokenizer to generate token embedding; Cross-modal alignment: constrain the consistency of video and text embedding space through contrastive learning; Step 4.2: Modal mixed attention: Cross attention gating: dynamically adjust the attention weight of video and text token, formula as follows: a = s (co [v i ; t i ]) where v i is a video token, t i is a text token.
6. The method of claim 1, wherein, The step 5 comprises the following steps: Step 5.1: Distributed training strategy: Data parallelism: shard batch data to different devices, gradient synchronized through All-Reduce; Model parallelism: split encoder and decoder to different devices to reduce single card load; Step 5.2: Reasoning Acceleration: Sub-block Cache Reuse: For repeated sub-blocks, cache their KV Cache to avoid repeated computation; Progressive Decoding: In generating tasks, prefer to compute sub-blocks with high attention weights, gradually expanding the context window.