This invention discloses a sparse attention modeling method and
system for long texts based on four-dimensional spatiotemporal coordinates, belonging to the field of large-
scale model long text optimization and sparse attention technology. Existing full-scale attention mechanisms for large models suffer from
high memory consumption, large
inference latency in long texts, attention
diffusion, and failure of long-range dependencies. Various sparse attention algorithms are mostly based on window truncation and fixed-interval sampling, which are empirical simplification strategies that easily lose key long-range information and disrupt the text's logical temporal structure. This invention utilizes four-dimensional topological constraints to construct an explicit sparse attention
mask, abandoning the fixed-window sampling mode to achieve spatiotemporal topological adaptive sparse attention modeling. This invention preserves long-range temporal correlations through temporal causal constraints, isolates invalid cross-cycle interference through contextual
branch constraints, preserves key dependencies in the
inference chain through logical hierarchy constraints, isolates mixed information from multiple subjects through identity coordinates, and dynamically weights key token attention weights based on cognitive memory strength. This invention significantly reduces the computational power and memory overhead of long texts while fully preserving long-range
causality, logic, and contextual correlations, completely solving the core pain points of degradation and
information loss in traditional sparse attention for long texts, and significantly improving the stability and logical consistency of
inference for millions of ultra-long texts.