Long sequence streaming three-dimensional reconstruction method and system based on controlled memory

CN122820985APending Publication Date: 2026-09-25HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611050889.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本发明提供一种基于受控记忆的长序列流式三维重建方法及系统,其目的在于解决现有带有时序因果注意力的前馈重建模型在处理长视频流输入时,难以兼顾显存占用控制与历史上下文完整性,固定词元预算的记忆策略随序列长度增加会造成重建精度大幅衰退的问题

Benefits of technology

[0026]本发明提出了一种基于受控记忆的长序列流式三维重建方法RegVGGT,解决了带有时序因果注意力机制的前馈重建模型中显存急剧膨胀与上下文完整性受损之间的权衡问题。所述方法通过每帧最多仅保留1%最具显著性的词元,对上下文进行了激进的调控。结合兼容 FlashAttention架构的显著性评估机制,RegVGGT能够在消费级GPU上处理数千帧图像,且几乎没有对重建质量的折损。大量实验表明,本方法在长序列基准测试中全面超越了现有的基准方法,为将前馈重建模型适配于流式重建的社区研究提供了有价值的见解。本发明为免训练的在线推理阶段词元调控方法,仅修改推理阶段的记忆库管理逻辑,不修改模型网络结构与预训练权重。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820985A_ABST
    Figure CN122820985A_ABST
Patent Text Reader

Abstract

The application discloses a long-sequence streaming three-dimensional reconstruction method and system based on controlled memory, belongs to the technical field of computer vision three-dimensional reconstruction, and solves the trade-off problem between the sharp expansion of display memory and the damage of context integrity in a feedforward reconstruction model with a time sequence causal attention mechanism.The application comprises the following steps: acquiring a first frame image of an input video stream, retaining key-value tokens of each layer of the first frame image completely, constructing a global anchor memory library and permanently storing the global anchor memory library; calculating the initial saliency score of each token in a subsequent frame image; according to a preset token retention ratio, screening a preset number of tokens with high saliency score, and adding the corresponding key-value pairs to a historical context memory library; calling the tokens in the global anchor memory library and the updated historical context memory library, performing time sequence causal attention operation, and completing feedforward reconstruction reasoning of the current frame.The application is suitable for long video three-dimensional reconstruction scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision 3D reconstruction technology, specifically relating to a long-sequence streaming 3D reconstruction method and system based on controlled memory. Background Technology

[0002] When processing long video streams, feedforward reconstruction models cannot preserve the inference context of the entire video stream under limited GPU memory. Although existing feedforward reconstruction models have achieved significant results, they operate under an offline inference mechanism. This means that whenever a new frame arrives, the model is forced to reprocess all historical frames, making it impractical for processing long video streams.

[0003] By using a temporal causal approach to implement lexical attention operations instead of traditional global attention, existing techniques enable feedforward reconstruction models to be adapted to process new input images frame by frame, i.e., online inference capabilities. This significantly reduces the computational overhead of streaming reconstruction because historical lexical units can be accessed directly from the memory bank, eliminating the need to recalculate them every time a new frame arrives.

[0004] To further expand the capacity of encoding historical frame context into limited video memory, existing techniques have introduced various memory strategies to retain historical frame terms as context, either implicitly or explicitly, for reconstructing new input frames. While these methods successfully limit video memory usage, they face the problem of forgetting crucial historical context as the number of frames processed increases, leading to a decrease in reconstruction accuracy.

[0005] To mitigate the rapid increase in GPU memory consumption in feedforward reconstruction models based on temporal causal attention, recent studies have artificially employed fixed lexical budgets to limit memory usage. Existing techniques attempt to address the memory issue in long video stream inference by balancing the integrity of the inference context with GPU memory consumption. However, these methods either suffer from rapid memory expansion or compromise context integrity due to artificially set memory limits. As the number of frames processed increases, fixed-budget memory bank strategies must decide which memories to discard for new input frames, leading to distortion of historical frame context. Furthermore, it remains unclear how to ensure that currently discarded lexical terms are not significant for future frames. Summary of the Invention

[0006] This invention provides a long-sequence streaming 3D reconstruction method and system based on controlled memory. Its purpose is to solve the problem that existing feedforward reconstruction models with temporal causal attention are difficult to balance memory usage control and historical context integrity when processing long video stream inputs, and that the memory strategy with fixed lexical budget causes a significant decline in reconstruction accuracy as the sequence length increases.

[0007] Firstly, the purpose of this invention is to provide a long-sequence streaming 3D reconstruction method based on controlled memory, applied to a pre-trained feedforward reconstruction model with a temporal causal attention mechanism, comprising the following steps:

[0008] S1: Obtain the first frame image of the input video stream, completely retain the key-value terms of each Transformer layer corresponding to the first frame image, construct a global anchor memory and store it permanently;

[0009] S2: For each subsequent frame of the video stream, calculate the initial saliency score of each word in the current frame.

[0010] S3: According to the preset word retention ratio, select a preset number of words with the highest salience scores from the current frame, and add the key-value pairs corresponding to these words to the historical context memory.

[0011] S4: Call the lexical units in the global anchor memory and the updated historical context memory, and perform temporal causal attention operation through the FlashAttention architecture to complete the feedforward reconstruction inference of the current frame.

[0012] Furthermore, a preferred solution is provided: in S2, the saliency score of each word in the current frame is calculated independently for each attention head; in S3, the corresponding word is selected and retained independently for each attention head, and the historical context memory bank corresponding to each attention head is maintained respectively.

[0013] Furthermore, a preferred solution is provided: when calculating the saliency score of a key term, the query terms in the current frame are first spatially downsampled to obtain a sampled query subset; the attention scores obtained by a single key term on all queries within the sampled query subset are summed, and the sum is used as the initial saliency score of the key term.

[0014] Furthermore, a preferred solution is provided: In S4, attention scores are calculated based on logarithmic and exponential statistics to be compatible with the FlashAttention architecture. Specifically, the normalization factor cached by FlashAttention for each query during the forward propagation process is retrieved, and the results of the scaling dot product operation between the current frame query and the keyword are combined to recover the accurate attention score.

[0015] Furthermore, a preferred option is provided: the preset word retention ratio is 1% of the total number of words in a single frame.

[0016] Furthermore, a preferred solution is provided: All Transformer layers of the feedforward reconstruction model are functionally divided in advance: based on the differences in attention distribution patterns of each layer, the Transformer layers are divided into query-independent layers and sparse query layers; for the query-independent layers, only the global anchor memory is used as its historical context, and the lexical units of subsequent input frames are not stored in the memory corresponding to this type of layer; for the sparse query layers, operations S2 to S3 are performed.

[0017] Furthermore, a preferred embodiment is provided: the query-independent layer is an early layer and a late layer of the Transformer network, and the sparse query-dependent layer is an intermediate layer of the Transformer network.

[0018] Secondly, the purpose of this invention is to propose a long-sequence streaming 3D reconstruction system based on controlled memory, applied to a pre-trained feedforward reconstruction model with a temporal causal attention mechanism. The system is used to implement the long-sequence streaming 3D reconstruction method based on controlled memory described in any one or more of the above-mentioned schemes, including:

[0019] Global anchor construction module: used to acquire the first frame image of the input video stream, retain the complete key-value terms of each Transformer layer corresponding to the first frame image, build a global anchor memory and store it permanently;

[0020] Saliency calculation module: used to calculate the initial saliency score of each word in the current frame for each subsequent frame of the video stream that is input sequentially;

[0021] Lexical filtering and updating module: It is used to filter out a preset number of words with the highest salience scores from the current frame according to a preset lexical retention ratio, and add the key-value pairs corresponding to these words to the historical context memory.

[0022] Attention reasoning module: It is used to call the lexical units in the global anchor memory and the updated historical context memory, and perform temporal causal attention operation through the FlashAttention architecture to complete the feedforward reconstruction reasoning of the current frame.

[0023] Thirdly, the present invention aims to provide a computer device, the computer device including a memory and a processor, the memory storing a computer program, and when the processor runs the computer program stored in the memory, the processor executes the long sequence streaming 3D reconstruction method based on controlled memory according to any one or more of the above-described schemes.

[0024] Fourthly, the present invention aims to provide a computer-readable storage medium for storing a computer program that executes the long-sequence streaming 3D reconstruction method based on controlled memory described in any one or more of the above-described schemes.

[0025] Compared with the prior art, the advantages of the present invention are:

[0026] This invention proposes RegVGGT, a long-sequence streaming 3D reconstruction method based on controlled memory, which addresses the trade-off between rapid memory expansion and compromised contextual integrity in feedforward reconstruction models with temporal causal attention mechanisms. The method aggressively modulates the context by retaining at most 1% of the most salient words per frame. Combined with a saliency evaluation mechanism compatible with the FlashAttention architecture, RegVGGT can process thousands of frames on consumer-grade GPUs with almost no loss in reconstruction quality. Extensive experiments demonstrate that this method comprehensively outperforms existing benchmark methods in long-sequence benchmarks, providing valuable insights for community research on adapting feedforward reconstruction models to streaming reconstruction. This invention is a training-free online inference-stage word manipulation method, modifying only the memory bank management logic during the inference stage without altering the model's network structure or pre-trained weights.

[0027] This invention is applicable to long video 3D reconstruction scenarios. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0029] Figure 1 This is a framework diagram of the long-sequence streaming 3D reconstruction method based on controlled memory described in this invention;

[0030] Figure 2 A visualization of temporal causal attention;

[0031] Figure 3 A paradigm comparison between previous streaming reconstruction paradigms and the RegVGGT proposed in this invention;

[0032] Figure 4 This shows the video depth estimation results on the Bonn dataset. Detailed Implementation

[0033] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application can also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.

[0034] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0036] Implementation Method 1

[0037] This implementation uses a pre-trained StreamVGGT feedforward reconstruction model with a temporal causal attention mechanism and accelerated by the FlashAttention architecture as the carrier. The model contains a total of 24 Transformer encoders. The input is a continuous RGB-D video stream, and the output is the reconstruction results such as camera pose, depth map, and 3D point cloud.

[0038] The long-sequence streaming 3D reconstruction method (RegVGGT) based on controlled memory described in this embodiment is as follows: Figure 1 As shown, RegVGGT does not enforce a fixed memory budget, but rather uses a training-free lexical modulation strategy to regulate the information allowed to enter the memory for each frame. This invention decouples layers into two types. Completely discard intermediate historical information, and It strictly preserves the top 1% of salient words in each frame. Combined with an LSE-based saliency evaluation mechanism, RegVGGT achieves long-term geometric consistency and true scene-independent scalability.

[0039] The method flow is as follows:

[0040] S1. Initialize the global anchor memory: Obtain the first frame of the input video stream, completely store the key-value terms corresponding to each Transformer layer of the first frame, construct a global anchor memory and store it permanently, represented as:

[0041] , (1)

[0042] in, and Indicates the first frame. Layer key / value terms, This represents the corresponding anchor memory.

[0043] The process includes: acquiring the first frame of the input video stream, inputting it into a pre-trained feedforward reconstruction model, encoding it through each Transformer layer, extracting all key (K) and value (V) terms for each layer, and storing them completely as a global anchor memory for that layer. This global anchor memory is permanently retained throughout the entire video stream inference process and is not cropped, replaced, or overwritten. All subsequent temporal causal attention calculations require access to the terms in this anchor memory.

[0044] S2. Functionally partition all Transformer layers and match differentiated memory strategies: For each subsequent frame of the video stream, calculate the initial saliency score of each word in the current frame. In this step, the saliency score of each word in the current frame is calculated independently for each attention head. When calculating the word saliency score, the query words in the current frame are first spatially downsampled to obtain a sampled query subset; the attention scores obtained by a single key word on all queries in the sampled query subset are summed, and the sum is used as the initial saliency score of that key word.

[0045] This implementation visualizes the attention map to reveal the interaction between the query (row) and the key (column). For example... Figure 2 As shown, this implementation sampled frames from the ScanNet dataset with a step size of 100 to characterize the long-term behavior of temporal causal attention. (a) shows the attention map averaged layer by layer, visualizing the cross-frame index. Below (b) shows the non-increasing response characteristics of keyword saliency over time by summing each column, i.e., calculating the sum of attention to the same key for different queries. (c) is the attention map of attention heads in layer 13.

[0046] Based on the differences in temporal causal attention distribution patterns across different Transformer layers, all Transformer layers are divided into two functional groups, each matched with a different memory update strategy. Through visualization and analysis of the attention maps of each layer, the Transformer layers are categorized into two differentiated attention patterns:

[0047] Query irrelevant layer ( : These are the early and late layers of the network (e.g., layer 4 and layer 21), and the attention graph along the column direction exhibits an attention pattern independent of the query.

[0048] Sparse query layer ( : For intermediate layers of the network (e.g., layers 13 and 15), the attention graph presents a query-specific, sparse, and diagonally structured attention pattern.

[0049] for Only global anchor terms from the first frame are retained, while key-value pairs from all subsequent frames are prevented from being cached in memory.

[0050] , (2)

[0051] in, Indicates in time series index A historical memory bank maintained by the department.

[0052] In contrast, by utilizing attention patterns that present diagonal structures, for The control allows frame-by-frame information to enter the memory. Besides retaining global anchor terms, for each new input frame, only the preceding terminology is retained. The most prominent word units, among which :

[0053] , (3)

[0054] in, This represents the retention rate of each unit of memory in each frame. This differs from the enforced rigid capacity constraint (…). The fixed budget baseline method of this invention results in a linear increase in video memory usage at an extremely low rate, i.e. By retention rate By setting it to only 0.01, this invention ensures that the growth of video memory usage remains highly sustainable.

[0055] For layers exhibiting query-independent attention patterns, retaining only the key-value terms of the first frame has almost no impact on reconstruction accuracy. In other words, no updates to the memory are needed beyond the first frame. This suggests that these highly responsive key-value terms may be consistent anchors in the physical world, used to align subsequent frames with the first, thus retaining the first frame's terms is sufficient for these layers to function well. For other layers, their sparse attention patterns suggest the possibility of applying aggressive term pruning to suppress updates to the historical term memory, minimizing memory usage.

[0056] like Figure 2 As shown in (c), the attention map reveals the unique behavior of each attention head, indicating that there are differences between the attention heads. Therefore, in this embodiment, the salience of the word is evaluated independently for each attention head to achieve better model performance.

[0057] S3. According to the preset lexical retention ratio, select a preset number of lexical units with the highest saliency scores from the current frame, and append the key-value pairs corresponding to these lexical units to the historical context memory. In this step, lexical units are selected and retained independently for each attention head, and the historical context memory corresponding to each attention head is maintained separately.

[0058] This invention introduces a lexical retention strategy specific to attention heads.

[0059] First, to prevent disruptive interference caused by simply averaging attention scores among different attention heads, this implementation method... Internally, lexical saliency is evaluated independently for different attention heads, and a memory is maintained. Secondly, to minimize the computational overhead introduced in saliency evaluation, this implementation utilizes the inherent spatial redundancy between adjacent queries. Specifically, this implementation uses a spatially downsampled subset of queries (denoted as...). The saliency of lexical terms is evaluated by dividing the query in the current frame into... Non-overlapping windows are used, and the query obtained is sampled from the top left corner of each window. Then, for the header... middle The each word element Its significance score It is defined as the attention score it accumulates from the sampled subset of queries, i.e., in The Middle The sum of columns, such as Figure 2 (b) shows a diagonal view slice. Finally, for each attention head, the top [heads] with the highest significance scores are... Each lexical unit is selected and retained independently, thus effectively preserving sparse, fine-grained geometric cues.

[0060] S4. Call the lexical units in the global anchor memory and the updated historical context memory, and perform temporal causal attention operation through the FlashAttention architecture to complete the feedforward reconstruction inference of the current frame.

[0061] Attention-based lexical retention strategies, including those described in this invention, rely on precise attention scores, which FlashAttention does not explicitly generate. Constructing a complete attention matrix will produce... This incurs significant memory overhead, contradicting the design principles of high memory efficiency and low latency. To address this issue, this invention introduces LSE-Saliency, a saliency evaluation mechanism compatible with the FlashAttention architecture. It dynamically recovers accurate attention scores using log-sum-exponential (LSE) statistics without generating a complete attention map. During its forward propagation, FlashAttention computes and caches a normalization factor for each query:

[0062] , (4)

[0063] Using this cached statistics, it is possible to efficiently calculate the precise attention score of any lexical unit relative to the entire historical context.

[0064] Given a scaling attention value between the query and the current frame key, defined as The precise attention score can be recovered as:

[0065] , (5)

[0066] in, express The Middle line, number The values ​​of the columns. This allows the present invention to perform fine-grained lexical retention while fully preserving the hardware efficiency of FlashAttention.

[0067] Depend on Figure 3As can be seen from the paradigm comparison between previous streaming reconstruction paradigms and the RegVGGT proposed in this invention: (a) Unconstrained explicit memory retains all historical key-value pairs, leading to a sharp increase in video memory usage. (b) Bounded implicit memory compresses historical information into a compact hidden state, leading to catastrophic forgetting. (c) Bounded explicit memory enforces a fixed video memory budget, relying on scene-scale implicit priors to reasonably measure this budget. (d) The controlled memory of this invention allows at most 1% of the tokens to enter and update the context memory bank per frame, greatly suppressing video memory expansion during video streaming.

[0068] By incorporating a FlashAttention-compatible lexical saliency evaluation scheme, RegVGGT can process thousands of frames of images on consumer-grade GPUs with almost no loss in reconstruction quality. Extensive experiments demonstrate that RegVGGT achieves state-of-the-art performance on long-sequence benchmarks for various feedforward reconstruction model prediction tasks, significantly outperforming previous streaming reconstruction benchmark methods based on feedforward reconstruction models.

[0069] Implementation Method 2

[0070] This embodiment is a further experimental illustration of the long-sequence streaming 3D reconstruction method based on controlled memory described in Embodiment 1.

[0071] This implementation was validated on a single NVIDIA RTX 3090 (24GB VRAM) graphics card and tested on multiple benchmark datasets such as ScanNet, TUM Dynamics, 7-Scenes, NRGBD, and Bonn, covering three core tasks: camera pose estimation, 3D point cloud reconstruction, and video depth estimation.

[0072] Following the experimental setup of TTT3R, this implementation evaluated the accuracy of camera pose on the TUM Dynamics and ScanNet datasets, using Absolute Trajectory Error (ATE), Relative Translation Error (RPE trans), and Relative Rotation Error (RPE rot) as evaluation metrics. For TUM Dynamics, the initial frame of each sequence was extracted. For ScanNet, sampling was performed using a temporal step size of 3. Evaluation sequences ranging from 50 to 1000 frames were generated for both datasets. Quantitative results for the extreme sequence lengths of 800, 900, and 1000 frames are shown in Table 1. As the results show, RegVGGT consistently outperforms all comparison methods across all metrics and sequence lengths. Notably, the fixed-budget methods experienced severe performance degradation with increasing input length, while RegVGGT maintained robust tracking accuracy.

[0073] Table 1. Camera pose estimation results on the TUM Dynamics and ScanNet datasets.

[0074]

[0075] Following the standard settings of CUT3R, this implementation evaluates the quality of 3D point cloud reconstruction on 7-Scenes and NRGBD datasets. To systematically evaluate performance over long sequences, this implementation follows the TTT3R settings, sampling each sequence using a temporal step size of 2 and utilizing the first 50 to 300 frames as input. This implementation provides performance metrics for three indicators: accuracy (Accuracy, Acc), completeness (Comp), and normal consistency (NC). Quantitative results for the 200 and 300-frame limit sequence lengths are shown in Table 2. As the results show, RegVGGT establishes new state-of-the-art performance on the vast majority of metrics. Crucially, when the sequence length is extended from 200 to 300 frames, the fixed-budget baseline method suffers a sharp performance degradation in accuracy and consistency (e.g., the mean accuracy error of Evict3R spikes by more than 50% on 7-Scenes). In contrast, RegVGGT remains highly stable with minimal performance degradation, validating the effectiveness of the proposed method.

[0076] Table 2. 3D point cloud reconstruction results on the 7-Scenes and NRGBD datasets.

[0077] Since most standard benchmarks contain an extremely limited number of consecutive frames, this implementation follows the TTT3R setup to evaluate long sequence performance on the Bonn dataset. Specifically, after discarding the initial 30 frames of each sequence, consecutive subsequences with lengths ranging from 50 to 500 frames are sampled. Quantitative results are presented in... Figure 4 In this study, absolute relative error (Abs Rel) and threshold accuracy were used. The percentage of predicted depths within 1.25 times the true value is used as a metric. While Evict3R demonstrates competitiveness comparable to RegVGGT on ultra-short sequences (50 and 100 frames), the proposed model rapidly reaches state-of-the-art performance as sequence length increases. Crucially, the performance gap between the proposed method and fixed-budget baseline methods widens significantly with increasing input length. This trend directly highlights the scalability and robustness of the proposed controlled memory.

[0078] This implementation underwent ablation experiments on the ScanNet dataset to systematically evaluate the contributions of each key component in the architecture of this invention. Under the same total lexical budget, this implementation ablated global anchor preservation, layer-specific memory allocation, and attention head-specific lexical preservation to evaluate their respective contributions. As shown in Tables 3 and 4, removing either the global anchor or the layer-specific memory allocation individually leads to a significant decrease in tracking accuracy. Furthermore, while both mechanisms ensure robustness in tracking, Table 5 reveals the crucial efficiency contribution of the attention head-specific lexical preservation strategy. The strategy of this invention reduces inference latency by more than half, with a negligible impact on the main accuracy. Ultimately, these results confirm that all three components are indispensable to the architecture of this invention.

[0079] Table 3 Ablation experiments with global anchor point preservation

[0080]

[0081] Table 4. Ablation experiments on layer-specific memory allocation

[0082]

[0083] Table 5 Ablation experiments on attention head-specific lexical retention

[0084]

[0085] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A long-sequence streaming 3D reconstruction method based on controlled memory, applied to a pre-trained feedforward reconstruction model with a temporal causal attention mechanism, characterized in that... Includes the following steps: S1: Obtain the first frame image of the input video stream, completely retain the key-value terms of each Transformer layer corresponding to the first frame image, construct a global anchor memory and store it permanently; S2: For each subsequent frame of the video stream, calculate the initial saliency score of each word in the current frame. S3: According to the preset word retention ratio, select a preset number of words with the highest salience scores from the current frame, and add the key-value pairs corresponding to these words to the historical context memory. S4: Call the lexical units in the global anchor memory and the updated historical context memory, and perform temporal causal attention operation through the FlashAttention architecture to complete the feedforward reconstruction inference of the current frame.

2. The long-sequence streaming 3D reconstruction method based on controlled memory according to claim 1, characterized in that, In S2, the saliency score of each word in the current frame is calculated independently for each attention head; in S3, the corresponding word is selected and retained independently for each attention head, and the historical context memory bank corresponding to each attention head is maintained respectively.

3. The long-sequence streaming 3D reconstruction method based on controlled memory according to claim 2, characterized in that, When calculating the saliency score of a key term, the query terms in the current frame are first spatially downsampled to obtain a sampled query subset. The attention scores obtained by a single key term on all queries within the sampled query subset are summed, and the sum is used as the initial saliency score of that key term.

4. The long-sequence streaming 3D reconstruction method based on controlled memory according to claim 1, characterized in that, In step S4, attention scores are calculated based on logarithmic and exponential statistics to be compatible with the FlashAttention architecture. Specifically, the normalization factor cached by FlashAttention for each query during the forward propagation process is retrieved, and the results of the scaling dot product operation between the current frame query and the keyword are combined to recover the accurate attention score.

5. The long-sequence streaming 3D reconstruction method based on controlled memory according to claim 1, characterized in that, The preset lexical retention ratio is 1% of the total number of lexical units in a single frame.

6. The long-sequence streaming 3D reconstruction method based on controlled memory according to claim 1, characterized in that, The functions of all Transformer layers in the feedforward reconstruction model are pre-defined: based on the differences in attention distribution patterns of each layer, the Transformer layers are divided into query-independent layers and sparse query layers; for the query-independent layers, only the global anchor memory is used as its historical context, and the lexical units of subsequent input frames are not stored in the memory corresponding to this type of layer; for the sparse query layers, operations S2 to S3 are performed.

7. The long-sequence streaming 3D reconstruction method based on controlled memory according to claim 6, characterized in that, The query-independent layers are the early and late layers of the Transformer network, and the sparse query-dependent layers are the middle layers of the Transformer network.

8. A long-sequence streaming 3D reconstruction system based on controlled memory, applied to a pre-trained feedforward reconstruction model with a temporal causal attention mechanism, characterized in that... To implement the long-sequence streaming 3D reconstruction method based on controlled memory as described in any one of claims 1-7, comprising: Global anchor construction module: used to acquire the first frame image of the input video stream, retain the complete key-value terms of each Transformer layer corresponding to the first frame image, build a global anchor memory and store it permanently; Saliency calculation module: used to calculate the initial saliency score of each word in the current frame for each subsequent frame of the video stream that is input sequentially; Lexical filtering and updating module: It is used to filter out a preset number of words with the highest salience scores from the current frame according to a preset lexical retention ratio, and add the key-value pairs corresponding to these words to the historical context memory. Attention reasoning module: It is used to call the lexical units in the global anchor memory and the updated historical context memory, and perform temporal causal attention operation through the FlashAttention architecture to complete the feedforward reconstruction reasoning of the current frame.

9. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program. When the processor runs the computer program stored in the memory, the processor executes the long-sequence streaming 3D reconstruction method based on controlled memory according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that executes the long-sequence streaming 3D reconstruction method based on controlled memory as described in any one of claims 1-7.