Large language model reasoning optimization method, system and device for resource-constrained equipment, and medium

By dynamically sensing resource status before and during inference, optimizing tree-structured attention masks and model configuration, the memory problem of speculative decoding techniques on resource-constrained devices is solved, achieving efficient and stable inference for large language models.

CN120996186APending Publication Date: 2025-11-21SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511034965.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing speculative decoding techniques incur huge memory overhead on resource-constrained devices, making stable deployment difficult and hindering the widespread adoption of large language models on edge devices.

Method used

By dynamically sensing the device resource status before and during inference, a series of collaborative optimization strategies are adopted, including: establishing a resource budget, optimizing tree-like attention masks, pruning, reducing the number of decoders and lowering model accuracy, and balancing inference speed, memory usage and generation quality.

Benefits of technology

It achieves efficient and stable large-scale speculative decoding inference on devices with limited memory, maintaining speed advantage while controlling memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996186A_ABST
    Figure CN120996186A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and relates to a large language model reasoning optimization method, system and device for resource-constrained equipment, and a medium. The method comprises the following steps: establishing a resource budget, acquiring video memory parameters of equipment, including a total available video memory and a current available video memory, and setting a memory total budget of a reasoning session; evaluating memory requirements, including static calculation and dynamic calculation; calculating a total memory demand, and if the total memory demand is greater than the total memory budget, starting an iterative optimization sub-process; and loading the configuration meeting the budget, and executing speculation decoding through the optimized attention mask, which comprises the following steps of: generating a candidate lexical element sequence in parallel based on a mask structure, verifying the candidate sequence in batches and accepting or rejecting lexical elements, and updating a confirmed sequence and key value cache. Through a collaborative optimization strategy, the reasoning speed, memory occupation and generation quality are intelligently balanced, so that efficient and stable large-model speculation decoding reasoning is realized on equipment with limited memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and to a method, system, device, and medium for optimizing large language model reasoning for resource-constrained devices. Background Technology

[0002] In recent years, large language models (LLMs) have made revolutionary progress in fields such as natural language processing, content generation, and complex decision-making. However, the inference process of LLMs, especially the autoregressive generation method, suffers from computational intensity and high latency. Each generation of a token requires a complete model forward propagation, which limits its real-time interactive applications.

[0003] To accelerate the reasoning process, academia and industry have proposed speculative decoding techniques. The core idea of ​​this technique is to use a small draft model or multiple decoding heads to generate a candidate word sequence in parallel and "speculatively," and then have the entire sequence validated at once by the original large language model (base model). Since most candidate words can be validated, this is equivalent to exchanging the computation of a large model for the generation of multiple words, thus significantly improving the reasoning speed.

[0004] However, existing speculative decoding techniques, such as Medusa, while improving speed, rely on additional memory overhead. To generate and validate candidate lexical units, the system needs to allocate significant memory to store parameters for multiple decoding heads, manage candidate lexical trees, and the corresponding key-value cache (KVcache). This high memory cost makes speculative decoding techniques difficult to deploy on devices with limited memory resources, such as GPUs in consumer laptops, smartphones, and edge computing devices. Forcing their execution on these devices can easily lead to program crashes due to exceeding memory limits, severely hindering the widespread adoption and application of high-performance large language models on edge devices.

[0005] Therefore, how to effectively control memory usage while maintaining the speed advantage of speculative decoding, so that it can run stably and efficiently in a resource-constrained environment, is a technical problem that urgently needs to be solved in the field of large language model deployment.

[0006] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0007] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0008] This disclosure provides a method, system, apparatus, and medium for optimizing large language model inference on resource-constrained devices, aiming to address the technical pain points of existing speculative decoding techniques, such as high memory overhead and difficulty in deployment on resource-constrained devices. Specifically, this invention provides a method and system for optimizing large language model inference, which can dynamically perceive the device resource status before and during inference, and intelligently balance inference speed, memory usage, and generation quality through a series of collaborative optimization strategies, thereby achieving efficient and stable large model speculative decoding inference on memory-limited devices.

[0009] In some embodiments, the method includes:

[0010] Establish a resource budget, obtain the device's video memory parameters, including total available video memory and currently available video memory, and set the total memory budget for the inference session;

[0011] The memory requirements are assessed, including static and dynamic calculations. The static calculation calculates the memory required to load the base model and the memory required to load one decoder head at a preset precision, and accurately calculates the key-value cache memory requirements based on the total sequence length of the task. The dynamic calculation obtains the full tree-structured attention mask and its total number of nodes, and calculates the corresponding dynamic buffer memory.

[0012] Budget check and optimization: Calculate the total memory requirement. If the total memory requirement is greater than the total memory budget, start the iterative optimization sub-process and perform the following operations in priority order until the budget is met: tree attention mask pruning, reduce the number of decoders, and reduce model accuracy.

[0013] Perform inference, load configurations that meet the budget, and perform speculative decoding using an optimized attention mask, including: generating candidate lexical sequences in parallel based on the mask structure, batch validating candidate sequences and accepting or rejecting lexical sequences, and updating confirmed sequences and key-value caches.

[0014] Preferably, the formula for calculating the key-value cache memory requirement Mem_KV is as follows:

[0015] Mem_KV = 2 * l * b * d_kv * s * p,

[0016] Where l is the number of attention layers in the model, b is the batch size during inference, d_kv is the total dimension of all key-value headers in the model, s is the total sequence length that needs to be cached, and p is the number of bytes corresponding to the numerical precision of the model parameters.

[0017] Preferably, the total memory requirement is equal to the sum of the memory required to calculate and load the base model at a preset precision, the memory required to load one decoder head, the key-value cache memory requirement, and the dynamic buffer memory requirement.

[0018] Preferably, tree-like attention mask pruning includes:

[0019] The template pruning method based on the pruning rate function is suitable for standard full attention mask template scenarios.

[0020] The direct construction method based on feature constraints is suitable for scenarios with extremely limited memory, where the mask size needs to be precisely controlled from scratch.

[0021] Template pruning methods based on pruning rate functions include:

[0022] The maximum memory allowed for the dynamic buffer is calculated and divided by the memory required by a single node at runtime to determine the maximum number of nodes the attention mask can have. The formula for calculating the maximum memory (Budget_buffer) allowed for the dynamic buffer is as follows:

[0023] Budget_buffer=Budget_total-(Mem_model+Mem_heads+Mem_KV),

[0024] Among them, Budget_total represents the total memory budget, Mem_model represents the total memory requirement, Mem_heads represents the memory required to load one decoder head, and Mem_KV represents the memory requirement for the key-value cache.

[0025] The logistic function is used as the pruning rate function to prune layer by layer from shallow to deep, resulting in the pruned attention mask.

[0026] Preferably, the logistic function formula is as follows:

[0027] PruningRate(l)=clip(1-exp(-a*l*(1-b*l)),0,1),

[0028] Where PruningRate(l) is the function value, l is the tree level, a and b are hyperparameters that control the shape of the function, and clip represents the boundary constraint.

[0029] Preferably, the specific method for pruning layer by layer from shallow to deep is as follows:

[0030] Calculate the number of nodes that need to be removed in this layer, NumToRemove(l):

[0031] NumToRemove(l)=floor(NumNodesAtLevel_full(l)*PruningRate(l)),

[0032] Where floor represents rounding down, and NumNodesAtLevel_full(l) represents the total number of complete nodes at level l;

[0033] Remove NumToRemove(l) nodes sequentially from the rightmost node of this layer to the left.

[0034] Preferred, feature-constraint-based direct construction methods are suitable for scenarios with extremely limited memory, requiring precise control of the mask size from scratch, including:

[0035] Initialize an empty tree containing only the root node, and add child nodes level by level according to the breadth-first search principle;

[0036] During the addition process, it is checked in real time that the total number of nodes is less than the maximum total number of nodes, and the number of leaf nodes is less than or equal to the maximum number of leaf nodes.

[0037] Preferably, the speculative decoding includes:

[0038] Candidate words are sampled based on the optimized attention mask, and only candidates with node positions in the mask are generated;

[0039] The candidate sequence from the root node to the leaf node is concatenated with the input sequence and then input into the base model for verification in batches.

[0040] Validate tokens in branch order; if a token is rejected, its subsequent child nodes are automatically rejected.

[0041] The longest accepted sequence is selected and appended to the confirmed text, and the key-value cache is updated.

[0042] In some embodiments, the system includes:

[0043] The pre-inference budget and optimization module is configured to: obtain the resource budget, assess the static memory requirements, and adaptively optimize the tree-like attention mask based on the remaining memory budget; if the overall memory requirements still exceed the resource budget after optimizing the attention mask, further operations such as reducing the number of decoding heads or reducing the numerical accuracy of the base model loading are performed.

[0044] The inference execution module is configured to: load the optimized configuration determined by the previous module, and generate and verify candidate lexical sequences based on the optimized attention mask to complete speculative decoding inference.

[0045] In some embodiments, the apparatus includes a processor and a memory storing program instructions, the processor being configured to execute the large language model inference optimization method for resource-constrained devices when running the program instructions.

[0046] In some embodiments, the storage medium stores a computer program that, when executed by a processor, implements the large language model inference optimization method for resource-constrained devices.

[0047] This disclosure provides a method, system, apparatus, and medium for optimizing large language model inference for resource-constrained devices, which can achieve the following technical effects:

[0048] It can dynamically sense the status of device resources before and during inference, and intelligently balance inference speed, memory usage and generation quality through a series of collaborative optimization strategies, thereby achieving efficient and stable large-scale speculative decoding inference on devices with limited memory.

[0049] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description

[0050] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein:

[0051] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0052] Figure 2 This is a flowchart illustrating the template pruning method based on the pruning rate function;

[0053] Figure 3 This is a flowchart illustrating the direct construction method based on feature constraints;

[0054] Figure 4 This is a schematic diagram of the speculative decoding process.

[0055] Figure 5 This is a schematic diagram of a device structure according to an embodiment of the present disclosure. Detailed Implementation

[0056] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0057] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0058] Unless otherwise stated, the term "multiple" means two or more.

[0059] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0060] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0061] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.

[0062] Example 1

[0063] like Figure 1 - Figure 4 As shown, a method for optimizing large language model inference for resource-constrained devices includes a pre-inference budget and optimization process and an inference execution process.

[0064] Specifically, including:

[0065] Pre-inference budgeting and optimization process

[0066] Step 1: Establish a resource budget, obtain the device's video memory parameters, including total available video memory and currently available video memory, and set the total memory budget for the inference session;

[0067] Step 2: Assess memory requirements, including static and dynamic calculations. The static calculation calculates the memory required to load the base model and the memory required to load one decoder head at a preset precision, and accurately calculates the key-value cache memory requirements based on the total sequence length of the task. The dynamic calculation obtains the full tree-structured attention mask and its total number of nodes, and calculates the corresponding dynamic buffer memory.

[0068] Step 3: Budget check and optimization. Calculate the total memory requirement. If the total memory requirement is greater than the total memory budget, start the iterative optimization sub-process and perform the following operations in priority order until the budget is met: tree attention mask pruning, reduce the number of decoding heads, and reduce model accuracy.

[0069] Inference Execution Process

[0070] Step 4: Perform inference, load the configuration that meets the budget, and perform speculative decoding through the optimized attention mask, including: generating candidate lexical sequences in parallel based on the mask structure, batch verifying candidate sequences and accepting or rejecting lexical sequences, and updating the confirmed sequences and key-value cache.

[0071] As a refinement of the above embodiments, step 1, establishing the resource budget, specifically includes:

[0072] Step 1.1: Call the device interface to obtain the core hardware parameters of the resource-constrained device, including at least: total available video memory VRAM_total and currently available video memory VRAM_free.

[0073] Step 1.2: Establish a total memory budget of Budget_total = VRAM_free for an inference session.

[0074] As a refinement of the above embodiments, step 2 involves a preliminary assessment of static and dynamic memory requirements:

[0075] Step 2.1 (Static): Calculate the memory Mem_model required to load the base model at a preset precision p (e.g., FP16), and the memory Mem_heads required to load l decoder heads.

[0076] Step 2.2 (Static): Calculate the precise memory requirement Mem_KV for the key-value cache (KVcache). Unlike the wasteful practice in existing technologies that reserve space for the maximum context length of the model, this invention calculates the required memory precisely based on the preset total length of the inference task s (e.g., obtained by multiplying the number of session rounds n and the maximum single-round length m, i.e., s = n * m) using the following formula:

[0077] Mem_KV=2*l*b*d_kv*s*p,

[0078] Among them, 2 represents two caches for storing "Key" and "Value" respectively. l is the number of attention layers in the model, b is the batch size during inference, d_kv is the total dimension size of all key-value heads in the model, which is usually equal to the number of KV heads multiplied by the dimension of each head, s is the total sequence length to be cached, and p is the number of bytes corresponding to the numerical precision for storing model parameters (for example, p = 2 bytes under FP16 precision).

[0079] Step 2.3 (Dynamic): Obtain a default, unoptimized full tree-shaped attention mask Mask_full, which has the total number of nodes N_full. Calculate the dynamic buffer memory Mem_buffer_full required to support this full mask, and the size of this memory is positively correlated with N_full and the vocabulary size w of the model, that is, Mem_buffer_full ∝ N_full * w.

[0080] As a refinement of the above embodiment, the said step 3 budget check and core optimization loop:

[0081] Step 3.1 (Preliminary Check): Calculate the total memory required when using the default configuration (including the full attention mask Mask_full) without any optimization:

[0082] Mem_demand = Mem_model + Mem_heads + Mem_KV + Mem_buffer_full.

[0083] Step 3.2 (Decision Branch): Compare Mem_demand with the total memory budget Budget_total.

[0084] Case 1 (Sufficient Budget): If Mem_demand < Budget_total, it is determined that no optimization is required, and the system will skip all optimization sub-steps and directly enter step 4: Load Configuration and Execute Inference.

[0085] Case 2 (Insufficient Budget): If Mem_demand > Budget_total, it is determined that optimization is required, and the system will start an iterative, prioritized optimization sub-process, as described below.

[0086] Step 3.3 (Iterative Optimization Sub-Process):

[0087] This sub-process is executed in the preset priority order. After each operation is completed, the total memory requirement is recalculated, and it is checked whether the budget has been met. Once it is met, the optimization process ends and enters step 4. If the budget is still not met after traversing all optimization items, a configuration failure is reported.

[0088] Step 3.3.1 (Priority 1): Adaptive pruning using tree-like attention masking:

[0089] like Figure 2 As shown, this step aims to reduce the dynamic buffer memory requirement (Mem_buffer) by decreasing the size of the attention mask. The system includes two alternative pruning implementations:

[0090] Method A: Template pruning method based on pruning rate function

[0091] This method is suitable for scenarios where a standard full attention mask template (Mask_full) exists.

[0092] Step 3.3.1.A.1: Determine the pruning target. Calculate the maximum memory allowed for the dynamic buffer:

[0093] Budget_buffer=Budget_total-(Mem_model+Mem_heads+Mem_KV),

[0094] Among them, Budget_total represents the total memory budget, Mem_model represents the total memory requirement, Mem_heads represents the memory required to load one decoder head, and Mem_KV represents the memory requirement for the key-value cache.

[0095] Based on this budget, the maximum number of nodes N_target that the attention mask can have is calculated in reverse by dividing the budget size Budget_buffer by the memory required by a single node at runtime (which is mainly determined by the vocabulary size and numerical precision).

[0096] Step 3.3.1.A.2: Define and apply the hierarchical pruning rate function. This invention employs a pruning rate function capable of differentiating between different tree levels, and a scaling logistic function:

[0097] PruningRate(l)=clip(1-exp(-a*l*(1-b*l)),0,1),

[0098] Wherein, PruningRate(l) is the function value, l is the tree level (l=1 corresponds to the first layer of the decoder), and a and b are hyperparameters that control the shape of the function. The characteristics of this function are: when l is small (shallow layer), the value of PruningRate(l) is close to 0, that is, light pruning or no pruning, in order to preserve the diversity of candidate words; when l increases (deep layer), PruningRate(l) quickly approaches 1, that is, heavy pruning, in order to quickly reduce the size of the tree.

[0099] Step 3.3.1.A.3: Perform pruning operations layer by layer. From l=1 to the maximum level, for each level of Mask_full:

[0100] Calculate the number of nodes that need to be removed in this layer:

[0101] NumToRemove(l)=floor(NumNodesAtLevel_full(l)*PruningRate(l)),

[0102] Perform targeted removal: Candidate terms from speculative decoding are typically sorted from left to right by generation probability within the same level of the tree. To preserve high-probability sequences, the pruning operation starts from the rightmost node of that level (representing the candidate term with the lowest probability) and removes NumToRemove(l) nodes to the left in sequence.

[0103] Step 3.3.1.A.4: Generate the final mask. After all levels of processing are completed, a pruned attention mask Mask_pruned is obtained with approximately N_target nodes.

[0104] Method B: Direct construction method based on feature constraints

[0105] like Figure 3 As shown, this method is suitable for scenarios where memory is extremely limited and the mask size needs to be precisely controlled from scratch.

[0106] Step 3.3.1.B.1: Calculate feature constraints. Similar to step 3.3.1.A.1, calculate the maximum allowed total number of nodes N_max and the maximum number of leaf nodes L_max as hard constraints for construction.

[0107] Step 3.3.1.B.2: Initialize the construction. Create an empty tree containing only the root node.

[0108] Step 3.3.1.B.3: Greedy filling layer by layer.

[0109] Processing the first layer: Add k child nodes to the root node (k is the preset tree width or number of samples), as long as 1+k<=N_max. A preferred strategy of this invention is to always ensure the integrity of the first layer to maximize the initial parallelism.

[0110] Processing subsequent layers: For each layer where l > 1, traverse all parent nodes of the previous layer. For each parent node, attempt to add child nodes to it. The addition process follows the breadth-first principle, that is, first add one child node to each parent node, then add a second child node to each parent node, and so on. After adding a new node, check whether the current total number of nodes exceeds N_max and whether the number of leaf nodes exceeds L_max.

[0111] Terminate the build. The build process terminates when it is impossible to add a new node on any branch without violating the N_max or L_max constraints, generating a custom mask Mask_custom that fully meets the budget.

[0112] Step 3.3.2 (Priority 2): Reduce the number of decoder heads:

[0113] If mask pruning alone is insufficient to meet the budget, the system will perform the following step: reduce the number of decoded heads l by 1, then return to the beginning of the optimization subprocess, reassess the memory, and try optimizing again starting with mask pruning of priority 1.

[0114] Step 3.3.3 (Priority 3): Reduce model accuracy:

[0115] If reducing the number of decoders still fails to meet the budget, the system will perform the following step: reduce the loading precision p of the base model from FP16 to INT8 or lower, then return to the beginning of the optimization sub-process to perform a complete evaluation and optimization attempt again.

[0116] As a refinement of the above embodiments, step 4 involves loading the configuration and executing inference:

[0117] Step 4.1: When the optimization process ends (or if the budget is sufficient, skip directly to this step), the system obtains a final configuration combination that meets the memory budget (including the base model used, decoder head, KV cache, and an optimized or default attention mask).

[0118] Step 4.2: The inference execution module loads this final configuration and begins executing speculative decoding inference. Its internal workflow includes:

[0119] Step 4.2.1: Generation of candidate lexical sequences:

[0120] Step 4.2.1.1: In each decoding step, the currently confirmed text sequence is input into the loaded multiple decoding heads.

[0121] Step 4.2.1.2: Each decoder head outputs the probability distribution (logits) of its predicted successor words in parallel.

[0122] Step 4.2.1.3: Based on these probability distributions and strictly following the structure of the optimized attention mask Mask_optimized, sampling is performed. Specifically, the system only generates candidate lexical units for node positions that exist in Mask_optimized. Since Mask_optimized has been pruned, the total number of candidate lexical units generated in this step is much smaller than when using the full mask, thus directly utilizing the results of the previous optimization.

[0123] Step 4.2.2: Batch verification of the base model:

[0124] Step 4.2.2.1: Concatenate the multiple candidate sequences formed from the root node of Mask_optimized to all leaf nodes with the original input sequence.

[0125] Step 4.2.2.2: Input these spliced ​​complete sequences into the loaded base model in one batch at a time for single forward propagation verification. This is the core of speculative decoding acceleration.

[0126] Step 4.2.2.3: The pedestal model outputs a confirmatory probability distribution for each position of each candidate sequence.

[0127] Step 4.2.3: Acceptance and rejection of candidate terms:

[0128] Step 4.2.3.1: Traverse each candidate lexical node in Mask_optimized.

[0129] Step 4.2.3.2: Compare the raw probability generated by the decoder (Step 4.2.1.2) with the verification probability generated by the pedestal model (Step 4.2.2.3).

[0130] Step 4.2.3.3: Using a pre-defined acceptance criterion (e.g., random sampling or greedy comparison), determine whether each candidate lexical unit is "accepted" by the base model. Validation starts from the root node of the tree and proceeds down each branch.

[0131] Step 4.2.3.4: Once a word on a branch is rejected, all subsequent child nodes of that branch will be automatically rejected as well.

[0132] Step 4.2.4: Determining the optimal sequence and updating the state:

[0133] Step 4.2.4.1: Among all the verified branches, select the longest accepted candidate sequence as the final output of this decoding step.

[0134] Step 4.2.4.2: Append all the tokens in this accepted sequence to the confirmed text sequence.

[0135] Step 4.2.4.3: Update the corresponding key-value cache (KVcache) to store the attention information of the newly received lexical sequence, in preparation for the next decoding step.

[0136] Step 4.2.4.4: Move the decoding position pointer to the end of the new sequence and return to step 4.2.1 to begin the next round of speculative decoding.

[0137] Example 2

[0138] The method of the present invention will be described below through a specific application scenario.

[0139] Scene setting:

[0140] Equipment: A consumer-grade laptop equipped with an Nvidia GeForce GPU.

[0141] Available video memory (VRAM_free): According to the query, the device currently has 7.42GB of available video memory.

[0142] Model: Qwen2.5-7B-Instruct-GPTQ-INT8.

[0143] Default configuration: precision p = W8A16, each parameter occupies 1 byte, decoding header l = 4, batch size b = 1, the default full attention mask Mask_full has 64 nodes and 42 candidate sequences (leaf nodes).

[0144] Execution process:

[0145] Step 1: Resource Budget Establishment

[0146] The system detected that the available video memory VRAM_free is 7.42GB, so Budget_total is set to 7.42GB.

[0147] Step 2: Preliminary assessment of static and dynamic memory requirements

[0148] Step 2.1 (Static Memory - Model and Decoder Head):

[0149] Mem_model:Qwen2.5-7B-Instruct-GPTQ-INT8 Memory usage is 7GB.

[0150] Mem_heads: 4 decoding heads (usually small Transformer layers), the test requires about 0.3GB in total.

[0151] Step 2.2 (Static Memory - KV Cache):

[0152] Set task length: Assuming a multi-turn dialogue scenario, the preset total cache sequence length s = 2048 words.

[0153] Checking the model configuration: From config.json, we know that the number of layers l = 28, the number of key-value (KV) heads is 4, and the dimension of each head d_head = hidden_size / num_attention_heads = 3584 / 28 = 128. Therefore, the total dimension of all KV heads d_kv = 4 * 128 = 512.

[0154] Calculate Mem_KV: Use the formula Mem_KV=2*l*b*d_kv*s*p;

[0155] Mem_KV = 2 * 28 * 1 * 512 * 2048 * 2 bytes ≈ 0.11 GB.

[0156] Step 2.3 (Dynamic Memory - Default Buffer):

[0157] Mem_buffer_full: The default 64-node mask requires storing a complete logits vector for each node. From config.json, we know that the vocabulary size vocab_size = 152064.

[0158] Calculate the buffer size: 64 (number of nodes) * 152064 (vocabulary size) * 2 (BF16 bytes) ≈ 18.6MB ≈ 0.018GB.

[0159] Step 3: Budget Review and Optimization Decisions

[0160] Step 3.1 (Preliminary Inspection):

[0161] Calculate total memory requirements:

[0162] Mem_demand=Mem_model+Mem_heads+Mem_KV+Mem_buffer_full, Mem_demand=7+0.3+0.11+0.018=7.428GB.

[0163] Step 3.2 (Decision Branch):

[0164] Compare Mem_demand (7.428GB) with Budget_total (7.42GB). 7.428GB > 7.42GB, budget insufficient. Initiate step 3.3: Iteratively optimize the sub-process.

[0165] Step 3.3 Iterative optimization sub-process:

[0166] Execution priority 1 – Tree-based attention mask adaptive pruning (Method A):

[0167] Step 3.3.1.A.1 (Determine the pruning target):

[0168] Calculate static memory usage: Mem_static = Mem_model + Mem_heads + Mem_KV = 7 + 0.3 + 0.11 = 7.41 GB;

[0169] The maximum memory allowed for the dynamic buffer is Budget_buffer = 7.42 - 7.41 = 0.01 GB (approximately 10 MB);

[0170] Reverse calculation of the maximum number of nodes N_target:

[0171] N_target=Budget_buffer / (vocab_size*p);

[0172] N_target=10*1024*1024 / (152064*2)≈10,485,760 / 304,128≈34.4.

[0173] The system sets the target to N_target = 34 nodes.

[0174] Steps 3.3.1.A.2 & A.3 (Applying functions and pruning):

[0175] The system applies a hierarchical pruning function with the goal of reducing the number of nodes from 64 to 34. It prioritizes preserving shallow nodes and significantly reduces the number of deep nodes, especially starting with the low-probability nodes on the far right of each layer.

[0176] Step 3.3.1.A.4 (Generate the final mask):

[0177] Generate a new, pruned attention mask Mask_pruned with a total of 34 nodes.

[0178] Recalculate total memory requirements:

[0179] The new buffer memory Mem_buffer' = 34 * 152064 * 2 bytes ≈ 9.9 MB ≈ 0.009 GB.

[0180] New total demand Mem_demand' = 7.41 + 0.009 = 7.419 GB

[0181] Budget checked again: 7.419GB < 7.42GB. Budget is sufficient.

[0182] The optimization process has been successfully completed.

[0183] Step 4: Load configuration and execute inference

[0184] Step 4.1 (Determine the final configuration): After optimization, the final configuration is: Qwen2.5-7B-Instruct-GPTQ-INT8, 4 decoding heads, and an optimized attention mask pruned to 34 nodes.

[0185] Step 4.2 (Loading and Execution): The inference execution module loads this configuration. In each subsequent inference step, this small 34-node mask is used to generate and validate candidate terms, enabling successful operation on memory-critical devices with almost no sacrifice in model accuracy and parallelism (still 4 heads).

[0186] Example 3

[0187] A large language model inference optimization system for resource-constrained devices includes:

[0188] The pre-inference budget and optimization module is configured to: obtain the resource budget, assess the static memory requirements, and adaptively optimize the tree-like attention mask based on the remaining memory budget; if the overall memory requirements still exceed the resource budget after optimizing the attention mask, further operations such as reducing the number of decoding heads or reducing the numerical accuracy of the base model loading are performed.

[0189] The inference execution module is configured to: load the optimized configuration determined by the previous module, and generate and verify candidate lexical sequences based on the optimized attention mask to complete speculative decoding inference.

[0190] Combination Figure 5 As shown, this disclosure provides a large language model inference optimization apparatus 300 for resource-constrained devices, including a processor 304 and a memory 301. Optionally, the apparatus may further include a communication interface 302 and a bus 303. The processor 304, communication interface 302, and memory 301 can communicate with each other via the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call logical instructions in the memory 301 to execute a large language model inference optimization method for resource-constrained devices according to the above embodiment.

[0191] Furthermore, the logic instructions in the aforementioned memory 301 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0192] The memory 301, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 304 executes functional applications and data processing by running the program instructions / modules stored in the memory 301, thereby implementing the large language model inference optimization method for resource-constrained devices described in the above embodiments.

[0193] The memory 301 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 301 may include high-speed random access memory and may also include non-volatile memory.

[0194] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to execute the above-described method for optimizing large language model inference for resource-constrained devices.

[0195] The aforementioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

[0196] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code; it can also be a transient storage medium.

[0197] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0198] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0199] The methods and products (including but not limited to devices and equipment) disclosed in the embodiments herein can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection shown or discussed between each other may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0200] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A method for optimizing large language model inference for resource-constrained devices, characterized in that, Includes the following steps: Establish a resource budget, obtain the device's video memory parameters, including total available video memory and currently available video memory, and set the total memory budget for the inference session; The memory requirements are assessed, including static and dynamic calculations. The static calculations calculate the memory required to load the base model and the memory required to load one decoder head at a preset precision, and accurately calculate the key-value cache memory requirements based on the total task sequence length. The dynamic calculation obtains the full tree-structured attention mask and its total number of nodes, and calculates the corresponding dynamic buffer memory. Budget check and optimization: Calculate the total memory requirement. If the total memory requirement is greater than the total memory budget, start the iterative optimization sub-process and perform the following operations in priority order until the budget is met: tree attention mask pruning, reduce the number of decoders, and reduce model accuracy. Perform inference, load configurations that meet the budget, and perform speculative decoding using an optimized attention mask, including: generating candidate lexical sequences in parallel based on the mask structure, batch validating candidate sequences and accepting or rejecting lexical sequences, and updating confirmed sequences and key-value caches.

2. The method for optimizing large language model inference for resource-constrained devices according to claim 1, characterized in that, The formula for calculating the key-value cache memory requirement Mem_KV is as follows: Mem_KV = 2 * l * b * d_kv * s * p, Where l is the number of attention layers in the model, b is the batch size during inference, d_kv is the total dimension of all key-value headers in the model, s is the total sequence length that needs to be cached, and p is the number of bytes corresponding to the numerical precision of the model parameters.

3. The method for optimizing large language model inference for resource-constrained devices according to claim 1, characterized in that, The total memory requirement is equal to the sum of the memory required to calculate and load the base model at a preset precision, the memory required to load one decoder head, the memory requirement for the key-value cache, and the memory requirement for the dynamic buffer.

4. The method for optimizing large language model inference for resource-constrained devices according to claim 1, characterized in that, Tree-like attention mask pruning includes: The template pruning method based on the pruning rate function is suitable for standard full attention mask template scenarios. The direct construction method based on feature constraints is suitable for scenarios with extremely limited memory, where the mask size needs to be precisely controlled from scratch. Template pruning methods based on pruning rate functions include: The maximum memory allowed for the dynamic buffer is calculated and divided by the memory required by a single node at runtime to determine the maximum number of nodes the attention mask can have. The formula for calculating the maximum memory (Budget_buffer) allowed for the dynamic buffer is as follows: Budget_buffer=Budget_total-(Mem_model+Mem_heads+Mem_KV), Among them, Budget_total represents the total memory budget, Mem_model represents the total memory requirement, Mem_heads represents the memory required to load one decoder head, and Mem_KV represents the memory requirement for the key-value cache. The logistic function is used as the pruning rate function to prune layer by layer from shallow to deep, resulting in the pruned attention mask.

5. The method for optimizing large language model inference for resource-constrained devices according to claim 4, characterized in that, The formula for the logistic function is as follows: PruningRate(l)=clip(1-exp(-a*l*(1-b*l)),0,1), Where PruningRate(l) is the function value, l is the tree level, a and b are hyperparameters that control the shape of the function, and clip represents the boundary constraint, which truncates the calculation result to the interval [0,1]. The specific method for pruning layer by layer from shallow to deep is as follows: Calculate the number of nodes that need to be removed in this layer, NumToRemove(l): NumToRemove(l)=floor(NumNodesAtLevel_full(l)*PruningRate(l)), Where floor represents rounding down, and NumNodesAtLevel_full(l) represents the total number of complete nodes at level l. Remove NumToRemove(l) nodes sequentially from the rightmost node of this layer to the left.

6. The method for optimizing large language model inference for resource-constrained devices according to claim 4, characterized in that, The direct construction method based on feature constraints is suitable for scenarios with extremely limited memory, where the mask size needs to be precisely controlled from scratch, including: Initialize an empty tree containing only the root node, and add child nodes level by level according to the breadth-first search principle; During the addition process, it is checked in real time that the total number of nodes is less than the maximum total number of nodes, and the number of leaf nodes is less than or equal to the maximum number of leaf nodes.

7. The method for optimizing large language model inference for resource-constrained devices according to claim 4, characterized in that, The speculative decoding includes: Candidate words are sampled based on the optimized attention mask, and only candidates with node positions in the mask are generated; The candidate sequence from the root node to the leaf node is concatenated with the input sequence and then input into the base model for verification in batches. Validate tokens in branch order; if a token is rejected, its subsequent child nodes are automatically rejected. The longest accepted sequence is selected and appended to the confirmed text, and the key-value cache is updated.

8. A large language model inference optimization system for resource-constrained devices, characterized in that, Performing the large language model inference optimization method for resource-constrained devices as described in any one of claims 1-7, comprising: The pre-inference budget and optimization module is configured to: obtain the resource budget, assess the static memory requirements, and adaptively optimize the tree-like attention mask based on the remaining memory budget; if the overall memory requirements still exceed the resource budget after optimizing the attention mask, further operations such as reducing the number of decoding heads or reducing the numerical accuracy of the base model loading are performed. The inference execution module is configured to: load the optimized configuration determined by the previous module, and generate and verify candidate lexical sequences based on the optimized attention mask to complete speculative decoding inference.

9. A large language model inference optimization device for resource-constrained devices, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute, when running the program instructions, the large language model inference optimization method for resource-constrained devices as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the large language model inference optimization method for resource-constrained devices as described in any one of claims 1-7 above.

Citation Information

Cited By

  • Text acceleration generation method and system based on dynamic mask and parallel decoding

    CN121881994A

  • Text acceleration generation method and system based on dynamic masking and parallel decoding

    CN121881994B