Method and system for inference task, and computing device cluster and program product

By calculating attention scores of text data lexical units and filtering out important lexical units during the pre-filling stage of LLM, the problem of text data exceeding the context window limit in long document question answering applications is solved, achieving efficient text data compression and accurate inference results.

WO2026152907A1PCT designated stage Publication Date: 2026-07-23HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-12-01
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing technologies face challenges in long document question answering applications, such as text data exceeding the LLM context window limit and high inference costs, leading to decreased inference performance and increased resource consumption. They are particularly ineffective in complex tasks such as logical reasoning, multi-hop question answering, and code generation.

Method used

By calculating the attention score of each word in the text data during the pre-filling stage of the first LLM, words that meet the compression conditions are selected and input into the second LLM for inference, avoiding the execution of subsequent stages, shortening the compression time, and improving efficiency.

Benefits of technology

While ensuring the accuracy of the inference results, it significantly shortens the text data compression time, improves compression efficiency, and reduces resource consumption and latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025139034_23072026_PF_FP_ABST
    Figure CN2025139034_23072026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a method and system for an inference task, and a computing device cluster and a related product, which belong to the technical field of AI. The method comprises: inputting tokens of an inference task and tokens of first text data into a first LLM, acquiring an attention score of each token in the first text data during a prefill phase of the first LLM, and on the basis of the attention score of each token, acquiring tokens, the attention scores of which meet a compression condition, so as to obtain tokens of compressed first text data; and inputting the tokens of the compressed first text data and the tokens of the inference task into a second LLM for inference, so as to obtain an inference result of the inference task. Since it is only necessary to calculate an attention score of each token in first text data on the basis of a prefill phase of a first LLM, during the compression of the first text data, there is no need for the first LLM to execute phases following the prefill phase, thereby shortening the duration for compressing the first text data, and thus improving the compression efficiency of the first text data.
Need to check novelty before this filing date? Find Prior Art