Method and system for inference task, and computing device cluster and program product
By calculating attention scores of text data lexical units and filtering out important lexical units during the pre-filling stage of LLM, the problem of text data exceeding the context window limit in long document question answering applications is solved, achieving efficient text data compression and accurate inference results.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-12-01
- Publication Date
- 2026-07-23
AI Technical Summary
Existing technologies face challenges in long document question answering applications, such as text data exceeding the LLM context window limit and high inference costs, leading to decreased inference performance and increased resource consumption. They are particularly ineffective in complex tasks such as logical reasoning, multi-hop question answering, and code generation.
By calculating the attention score of each word in the text data during the pre-filling stage of the first LLM, words that meet the compression conditions are selected and input into the second LLM for inference, avoiding the execution of subsequent stages, shortening the compression time, and improving efficiency.
While ensuring the accuracy of the inference results, it significantly shortens the text data compression time, improves compression efficiency, and reduces resource consumption and latency.
Smart Images

Figure CN2025139034_23072026_PF_FP_ABST