Inference request processing method, electronic equipment, storage medium and computer program product
By distinguishing between low-precision and high-precision requests in large-model inference scenarios, using quantitative computing and sparse attention mechanisms to process low-precision requests, generating lightweight cache collections, and ensuring full calculation of high-precision requests, it solves the problem of balancing efficiency and accuracy in existing technologies and realizes efficient and accurate inference services.
Patent Information
- Application Number
- CN202511250826.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing technologies find it difficult to strike a balance between processing efficiency and result accuracy in large-model inference scenarios. Low-precision inference requests lead to waste of computing resources and response delays. The key-value cache reuse mechanism for high-precision requests may reduce accuracy and lacks differentiated design.
By receiving inference requests and determining their types, the system uses the linear layer of quantized calculations and the sparse attention mechanism to process low-precision requests and generate a lightweight key-value cache set. It performs full calculations on high-precision requests and generates a comprehensive key-value cache set. In the decoding stage, it filters them according to the attention scoring function or uses the full cache for processing.
It supports multiple request levels under the same model instance, reduces overall computing resource consumption, improves resource utilization efficiency and response speed, and ensures service quality and user experience.
Smart Images

Figure CN120745847A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of computer technology, and in particular relates to an inference request processing method, electronic device, storage medium, and computer program product. Background Art
[0002] In large-scale model inference scenarios, as user demands for real-time performance and precision diverge, existing technologies face the challenge of balancing processing efficiency and result accuracy. Low-precision inference requests (such as those in real-time interaction and rapid response scenarios) are sensitive to latency, but general-purpose inference frameworks often employ a unified, full-scale computation model that is not optimized for these characteristics, resulting in wasted computing resources and response delays. High-precision inference requests (such as those in medical diagnosis and financial analysis) require extremely high result accuracy, but traditional key-value cache reuse mechanisms can reduce accuracy due to accumulated errors (such as feature drift and cache aging), making them unable to meet these stringent requirements. Furthermore, existing key-value cache technologies often rely on indiscriminate reuse and lack differentiated designs for requests of varying precision. This makes it difficult to adapt to the efficiency requirements of low-precision scenarios while also failing to guarantee the accuracy integrity of high-precision scenarios. Summary of the Invention
[0003] The present disclosure provides an inference request processing method, an electronic device, a storage medium, and a computer program product.
[0004] According to one aspect of the present disclosure, a method for processing an inference request is provided, wherein the inference request is processed through a model, the method comprising: receiving an inference request, determining the request type of the inference request, and determining the inference request as a first type inference request or a second type inference request, wherein the first type inference request is a low-precision inference request, and the second type inference request is a high-precision inference request. In the pre-filling stage of the model, the first type inference request is processed using a linear layer of quantized calculation and a sparse attention mechanism to determine the first key-value cache set corresponding to the first type inference request, or the second type inference request is fully calculated to determine the second key-value cache set corresponding to the second type inference request. And in the decoding stage of the model, the inference result of the first type inference request is determined based on the target key-value cache in the first key-value cache set determined by the attention scoring function, or the inference result of the second type inference request is determined based on all key-value caches in the second key-value cache set.
[0005] According to the inference request processing method of this embodiment, it is possible to intelligently differentiate low-precision first-type inference requests and high-precision second-type inference requests based on the request type, significantly improving resource utilization efficiency while ensuring service quality and user experience. In the pre-filling phase, a linear layer and sparse attention mechanism with quantized calculations are used for the first-type inference requests to reduce the amount of computation and memory usage, and determine a lightweight first key-value cache set. For the second-type inference requests, full calculations are performed to ensure the highest accuracy, and a comprehensive second key-value cache set is determined. After entering the decoding phase, the target key-value cache is filtered from the first key-value cache set based on the attention scoring function, and the results of the first-type inference request are quickly generated. The second-type inference request is directly processed based on the complete second key-value cache set to ensure accuracy. This flexible scheduling and optimization strategy not only enables a single large language model instance to support multiple request levels, but also reduces overall computing resource consumption, simplifies the system deployment and maintenance process, and ultimately provides a more efficient, accurate, and responsive service experience.
[0006] According to at least one embodiment of the present disclosure, the inference request processing method uses a quantized linear layer and a sparse attention mechanism to process the first type of inference request during the pre-population phase of the model. The method includes converting the weights and activation values of the linear layer into low-precision values based on a target quantization accuracy. Forward propagation of the first type of inference request through the quantized linear layer is performed. Matrix operations are performed on the first type of inference request using the sparse attention mechanism.
[0007] According to the technical solution of this embodiment, by converting the weights and activation values of the linear layer to low-precision numerical values and using these quantized parameters for calculations during the forward propagation process, the computational complexity and memory usage of the model are significantly reduced. In addition, the sparse attention mechanism is used to perform sparse matrix operations on the input sequence, further reducing redundant calculations.
[0008] According to at least one embodiment of the present disclosure, an inference request processing method is provided, wherein an inference result of the first type of inference request is determined based on a target key-value cache in the first key-value cache set determined by an attention scoring function, including: determining an attention score of a first key-value cache in the first key-value cache set based on an attention scoring function during a decoding phase of the model. The first key-value cache set having an attention score greater than or equal to the target score is used as a target key-value cache. The inference result of the first type of inference request is determined based on the target key-value cache.
[0009] According to the technical solution of this embodiment, efficient identification and utilization of key context information in low-precision reasoning tasks are achieved, which can significantly reduce the amount of calculation and memory access while retaining the most critical information for the current reasoning task, thereby improving decoding efficiency and reasoning speed.
[0010] According to the inference request processing method of at least one embodiment of the present disclosure, in the decoding stage of the model, the attention score of the first key-value cache in the first key-value cache set is determined based on the attention scoring function, including: determining the query vector based on the context information of the first type of inference request. The first key-value cache in the first key-value cache set is processed based on the linear layer and self-attention mechanism of the decoding stage to determine the key-value cache entry of the intermediate state information of the first key-value cache; the key-value cache entry includes a key vector and a corresponding value vector. The attention score of the first key-value cache in the first key-value cache set is determined by comparing the similarity between the query vector and the key vector based on the attention scoring function.
[0011] According to the technical solution of this embodiment, the key-value cache most relevant to the current decoding task can be efficiently screened out.
[0012] According to at least one embodiment of the present disclosure, the inference request processing method performs full computation on the second-type inference request to determine a second key-value cache set corresponding to the second-type inference request, including: performing a linear transformation on the second-type inference request using a target weight; performing attention computation on the second-type inference request after the linear transformation to determine the second key-value cache set for the second-type inference request.
[0013] According to the technical solution of this embodiment, the semantic information and association relationship of the input context can be completely retained without introducing any precision loss, ensuring higher accuracy and stability when processing complex tasks.
[0014] According to at least one embodiment of the present disclosure, the inference request processing method further includes: processing the first type of inference requests or the second type of inference requests in a single batch during a pre-population phase to determine a corresponding first key-value cache set or second key-value cache set; and processing the first key-value cache set or the second key-value cache set in multiple batches during a decoding phase to determine a corresponding inference result.
[0015] According to the technical solution of this embodiment, the initial computing overhead is effectively controlled in the pre-filling stage, and the system throughput and response flexibility are improved through multi-batch processing in the decoding stage, thereby improving the efficiency and scalability of the overall inference system while ensuring the quality of service.
[0016] According to at least one embodiment of the present disclosure, an inference request processing method monitors GPU memory occupancy and CPU queue depth in real time, and adjusts the batch size through a sliding window mechanism, including: when the GPU memory occupancy is less than a first target threshold and the CPU queue depth is greater than or equal to a second target threshold, increasing the batch size based on the sliding window mechanism; when the GPU memory occupancy is greater than or equal to a third target threshold, or the CPU queue depth is less than a fourth target threshold, reducing the batch size based on the sliding window mechanism.
[0017] According to the technical solution of this embodiment, when GPU resources are sufficient and the CPU task backlog is large, the batch size is automatically increased to improve throughput and GPU utilization. When GPU memory is nearing saturation or the CPU task backlog is low, the batch size is reduced to avoid resource overload and reduce task latency. This effectively balances system throughput and responsiveness, improving the stability of inference services in high-concurrency and load-fluctuating scenarios.
[0018] According to at least one embodiment of the present disclosure, the inference request processing method further includes: performing static analysis on operators in the model to determine control-intensive operators and parallel-intensive operators, allocating the control-intensive operators to the CPU for execution and allocating the parallel-intensive operators to the GPU for execution.
[0019] According to the technical solution of this embodiment, the advantages of the CPU in processing logical control tasks and the high performance characteristics of the GPU in large-scale parallel computing are brought into play, thereby effectively improving the overall reasoning efficiency and reducing the reasoning delay.
[0020] According to another aspect of the present disclosure, an electronic device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, so that the processor executes the inference request processing method of any embodiment of the present disclosure.
[0021] According to another aspect of the present disclosure, a storage medium is provided, in which execution instructions are stored. When the execution instructions are executed by a processor, they are used to implement the inference request processing method of any embodiment of the present disclosure.
[0022] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method for processing an inference request according to any embodiment of the present disclosure is implemented.
[0023] The beneficial effects of the technical solution of the present application are at least: it can intelligently differentiate low-precision first-type inference requests and high-precision second-type inference requests based on the request type, significantly improving resource utilization efficiency while ensuring service quality and user experience. In the pre-filling phase, a linear layer and sparse attention mechanism with quantized calculations are used for first-type inference requests to reduce computational complexity and memory usage, and determine a lightweight first key-value cache set. For second-type inference requests, full calculations are performed to ensure the highest accuracy, and a comprehensive second key-value cache set is determined. After entering the decoding phase, the target key-value cache is filtered from the first key-value cache set based on the attention scoring function, and the results of the first-type inference request are quickly generated. The second-type inference request is directly processed based on the complete second key-value cache set to ensure accuracy. This flexible scheduling and optimization strategy not only enables a single large language model instance to support multiple request levels, but also reduces overall computing resource consumption, simplifies the system deployment and maintenance process, and ultimately provides a more efficient, accurate, and responsive service experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings illustrate exemplary embodiments of the present disclosure and together with the description serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0025] Figure 1 It is a schematic diagram of the overall process of an inference request processing method according to one embodiment of the present disclosure.
[0026] Figure 2 This is a flowchart of processing a first type of inference request in a pre-filling phase in an inference request processing method according to an embodiment of the present disclosure.
[0027] Figure 3 It is a flowchart of determining a second key-value cache set in an inference request processing method according to an embodiment of the present disclosure.
[0028] Figure 4 This is a flowchart of determining an inference result of a first type of inference request in an inference request processing method according to an embodiment of the present disclosure.
[0029] Figure 5 This is a flowchart of determining the attention score of the first key-value cache in the inference request processing method according to an embodiment of the present disclosure.
[0030] Figure 6 It is a schematic block diagram of the structure of an inference request processing device according to an embodiment of the present disclosure.
[0031] Figure 7is a schematic block diagram of the structure of an electronic device according to one embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] The present disclosure is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are intended only to illustrate the relevant content and are not intended to limit the present disclosure. It should also be noted that, for ease of description, only the portions relevant to the present disclosure are shown in the accompanying drawings.
[0033] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure can be combined with each other. The technical solution of the present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0034] In real-world applications, users may need to simultaneously handle simple tasks with extremely high response time requirements (such as instant chatbots) and complex analytical tasks with strict accuracy requirements (such as professional document translation or deep data analysis). Current technical solutions often require deploying multiple model instances with different configurations to address these requirements. This not only increases computing resource consumption and management costs, but also limits the flexibility and scalability of the system.
[0035] To this end, the present disclosure proposes the following technical solutions, including: receiving an inference request, determining the request type of the inference request, and determining the inference request as a first type of inference request or a second type of inference request, wherein the first type of inference request is a low-precision inference request, and the second type of inference request is a high-precision inference request. In the pre-filling stage of the model, the first type of inference request is processed using a linear layer of quantized calculation and a sparse attention mechanism to determine the first key-value cache set corresponding to the first type of inference request, or the second type of inference request is fully calculated to determine the second key-value cache set corresponding to the second type of inference request. And in the decoding stage of the model, the inference result of the first type of inference request is determined based on the target key-value cache in the first key-value cache set determined by the attention scoring function, or the inference result of the second type of inference request is determined based on all the key-value caches in the second key-value cache set.
[0036] In the technical solution of the present application, low-precision first-type inference requests and high-precision second-type inference requests can be intelligently differentiated according to the request type, significantly improving resource utilization efficiency while ensuring service quality and user experience. In the pre-filling stage, a linear layer of quantized calculation and a sparse attention mechanism are used for the first-type inference request to reduce the amount of calculation and memory usage, and determine a lightweight first key-value cache set; while for the second-type inference request, a full calculation is performed to ensure the highest accuracy, and a comprehensive second key-value cache set is determined. After entering the decoding stage, the target key-value cache is filtered out from the first key-value cache set based on the attention scoring function, and the result of the first-type inference request is quickly generated, while the second-type inference request is directly processed based on the complete second key-value cache set to ensure accuracy. This flexible scheduling and optimization strategy not only enables a single large language model instance to support multiple request levels, but also reduces overall computing resource consumption, simplifies the system deployment and maintenance process, and ultimately provides a more efficient, accurate and responsive service experience.
[0037] For the convenience of description and to make the technical solution of the present disclosure easier to understand, the terms of the present disclosure are first explained before describing the technical solution of the present disclosure.
[0038] The prefill phase occurs after the AI model receives a complete input request but before it begins generating the first output token. The main tasks of this phase are to process the input request, calculate the context representation of all input tokens, and initialize the key-value cache required for the subsequent decoding phase.
[0039] The decoding phase begins after the pre-population phase. In this phase, the model generates output tokens one by one in an autoregressive manner. Each generated token is appended to the generated sequence and serves as input for the next step.
[0040] The linear layer, also known as the fully connected layer or dense layer, is one of the most basic layers in a neural network. Its main function is to perform a linear transformation on the input data through matrix multiplication and the addition of a bias vector.
[0041] The attention mechanism is a data processing method in machine learning, widely used in various machine learning tasks, such as natural language processing, image recognition, and speech recognition. The attention mechanism uses weights to reflect the degree of attention (importance) given to different pieces of information. The attention mechanism can be viewed as a multilayer perceptron (MLP) composed of a query matrix (Query), a key (Key), and a weighted average.
[0042] The attention scoring function measures the correlation or matching between a query vector (Query) and a key vector (Key). Its core function is to determine which parts of the input sequence are most relevant to the current output when the model is processing the current task. By calculating the similarity score between the query vector and each key vector, the attention scoring function provides the basis for subsequent Softmax normalization and value-weighted summation, thereby focusing on and selectively utilizing key information. Common attention scoring methods include dot product, scaled dot product, and additive attention, which affect the model's focus on contextual information and computational efficiency in different scenarios.
[0043] In a server environment, by dynamically adjusting computing strategies and resource allocation, simple user queries that require quick responses or complex analysis tasks that require high precision can be efficiently processed. For example, in a cloud-based intelligent customer service system, the appropriate model configuration is automatically selected based on the complexity of the user's question to ensure immediate response while maintaining a high level of service quality. In addition, for resource-constrained terminal devices such as smartphones or IoT devices, a single model can be optimized through quantization and sparsification techniques so that it can run locally without sacrificing too much performance, providing personalized recommendations, real-time speech recognition and other services, reducing dependence on network connections and protecting user privacy. At the same time, this flexible scheduling mechanism can also be applied to edge computing scenarios, enabling edge devices close to the data source to intelligently decide when and how to perform inference tasks, thereby reducing latency and improving overall inference efficiency.
[0044] Figure 1 FIG. 1 shows a schematic diagram of the overall process of the inference request processing method according to an embodiment of the present disclosure. Figure 1 In the illustrated method M100 , the inference request is processed by the model, and the method M100 includes steps S110 to S130 , wherein the method may be executed by a server.
[0045] In step S110, an inference request is received, and the request type of the inference request is determined to be a first type inference request or a second type inference request. The first type inference request is a low-precision inference request, and the second type inference request is a high-precision inference request.
[0046] When an inference request is received, the inference request type is determined based on the request type tag or by analyzing the information contained in the request. Tasks that require high response time but relatively low accuracy are marked as type 1 inference requests; complex tasks requiring high accuracy are marked as type 2 inference requests. Dynamically adjusting the processing strategy based on the request type maximizes resource utilization while ensuring quality of service and reducing overall computing costs.
[0047] The above reasoning requests include multiple modal data such as text sequences, images, audio, video, or tables.
[0048] The above request type can be determined based on the category selected by the user, or can be determined by analyzing the information carried in the inference request through the recognition module.
[0049] The above models include the trained GPT model and BERT model.
[0050] In a specific embodiment, an inference request is received from a client or server, the inference request containing data such as input text, task description, user identity, and / or device information. The information contained in the inference request is parsed, and key information is extracted based on semantic features, such as the length of the input text (short or long sentences), the semantics of the request content (simple question and answer, complex reasoning, or professional consultation), user identity, device information (mobile, edge device, or server), and the tag or priority carried in the request (high real-time requirements or high accuracy). The extracted key information is converted into structured features that can be used for classification. The structured features are analyzed and judged based on target rules or classification models to obtain the final classification results.
[0051] In step S120, during the pre-filling stage of the model, the first type of inference request is processed using the linear layer of quantized calculation and the sparse attention mechanism to determine the first key-value cache set corresponding to the first type of inference request, or the second type of inference request is fully calculated to determine the second key-value cache set corresponding to the second type of inference request.
[0052] During the pre-population phase of the AI model, different computation strategies are selected based on the request type. To accelerate processing and reduce resource consumption, a linear layer with quantized computation and a sparse attention mechanism are applied to the first type of inference requests. The precision of the model parameters is reduced, and the most representative attention connections are selected to determine the first key-value cache set. To ensure the highest accuracy, no compression or simplification strategies are employed, and instead a full computation method is used to determine the second key-value cache set for the second type of inference requests. This ensures that all necessary intermediate state information is fully preserved, providing the most accurate results during the decoding phase.
[0053] Optionally, in the pre-filling stage of the model, the first type of inference request is processed using a linear layer of quantized calculation and a sparse attention mechanism. In some embodiments of the present disclosure, the following may be included: Figure 2 Steps S210 to S230 are shown.
[0054] In step S210, the weights and activation values of the linear layer are converted into low-precision values based on the target quantization accuracy.
[0055] During AI model inference, the linear layer typically carries a large number of matrix operations and is a major source of computing resource consumption. To improve inference efficiency and reduce resource consumption, the weights and activations of the linear layer are quantized. This involves mapping the original high-precision floating-point numbers to low-precision values based on a preset target quantization accuracy. This process can be completed during model deployment or pre-inference processing, supporting both static quantization (e.g., post-training quantization) and dynamic quantization (e.g., dynamic precision adjustment based on input at runtime). It seamlessly integrates with subsequent inference processes, providing foundational support for lightweight inference. The target quantization accuracy is determined by the model configuration, preferably 8 bits.
[0056] In step S220 , the first type of inference request is forward propagated through the quantized linear layer.
[0057] In the AI model inference process, the first type of inference request is forward propagated through a quantized linear layer. The weights and activations of this linear layer have been converted from high-precision to low-precision to support efficient execution of low-precision computation instructions. This significantly improves inference speed, reduces memory bandwidth requirements, and optimizes overall resource utilization, while maintaining basic semantic understanding of the task.
[0058] In step S230 , matrix operations are performed on the first type of inference request using a sparse attention mechanism.
[0059] When processing first-type inference requests, a sparse attention mechanism is used to optimize the attention matrix. Specifically, after calculating the attention score for a first-type inference request, only the key-value pairs with the highest scores are retained, while the remaining redundant information with lower scores is ignored. This reduces the computational complexity and memory access overhead of subsequent matrix operations. This significantly improves inference efficiency without significantly affecting the ability to understand the semantics of the task.
[0060] Optionally, performing full calculation on the second type of reasoning request to determine the second key-value cache set corresponding to the second type of reasoning request may include: Figure 3 Steps S310 to S320 are shown.
[0061] In step S310 , a linear transformation is performed on the second type inference request using the target weight.
[0062] When processing the second type of inference request, the target weights are used to linearly transform the input. This involves applying an affine transformation to the input data using a high-precision weight matrix. This transformation does not quantize or compress the weights or activation values, ensuring that the model fully utilizes the original model's expressive power during inference, resulting in high-quality output.
[0063] In step S320, attention calculation is performed on the second type of reasoning request after the linear transformation to determine a second key-value cache set of the second type of reasoning request.
[0064] After the linear transformation of the second-type inference request is completed, attention calculation is further performed to extract the correlation between each position in the input sequence of the second-type inference request. By calculating the attention score between the query vector and the key vector, key context information is filtered out, and the corresponding key vector and value vector are cached to form a second key-value cache set. This second key-value cache set is reused in the subsequent decoding stage to improve generation efficiency and semantic coherence.
[0065] In step S130, during the decoding phase of the model, the inference result of the first type of inference request is determined based on the target key-value cache in the first key-value cache set determined by the attention scoring function, or the inference result of the second type of inference request is determined based on all key-value caches in the second key-value cache set.
[0066] During the decoding phase of the AI model, for the first type of inference requests, an attention scoring function is used to evaluate the relevance of each key-value cache entry in the first key-value cache set, screening out target key-value caches with higher attention scores. This reduces the amount of computation while determining inference results that meet the requirements for fast response. For the second type of inference requests, the complete second key-value cache set is directly used for inference to ensure the high accuracy and completeness of the output results. This demonstrates the intelligent scheduling capability of resource utilization, meeting the efficiency requirements of low-latency scenarios while ensuring the accuracy of high-precision tasks.
[0067] Optionally, determining the inference result of the first type of inference request according to the target key-value cache in the first key-value cache set determined by the attention scoring function may include: Figure 4 Steps S410 to S430 are shown.
[0068] In step S410, during the decoding phase of the model, an attention score of a first key-value cache in a first key-value cache set is determined based on an attention scoring function.
[0069] The attention scoring function evaluates the relevance of the key-value cache entries in the first key-value cache set to determine the context information that should be focused on at the current decoding moment. This lays the foundation for the attention mechanism to model dynamic context during the decoding process.
[0070] Optionally, in the decoding phase of the model, the attention score of the first key-value cache in the first key-value cache set is determined based on the attention scoring function. In some embodiments of the present disclosure, the following steps may be included: Figure 5 Steps S510 to S530 are shown.
[0071] In step S510 , a query vector is determined based on context information of the first type of reasoning request.
[0072] When processing the first type of inference request, a query vector for the current time step or task is generated based on the input context. This query vector is typically generated by linearly transforming the input representation. Its role is to represent "what to look for" in the current task in the attention mechanism and is used for relevance matching with the key vector.
[0073] In step S520, the first key-value cache in the first key-value cache set is processed based on the linear layer and self-attention mechanism of the decoding stage to determine the key-value cache entry of the intermediate state information of the first key-value cache; the key-value cache entry includes a key vector and a corresponding value vector.
[0074] To improve inference efficiency and resource utilization during the decoding process of an AI model, the key-value cache of the input sequence is typically pre-calculated and cached for reuse in subsequent decoding stages. In this step, based on the linear layer output of the current decoding stage and the self-attention mechanism, the state of the key-value cache entries in the first key-value cache set is updated to reflect the model's contextual understanding state at the current decoding time step.
[0075] In step S530 , the similarity between the query vector and the key vector is compared based on the attention scoring function to determine the attention score of the first key-value cache in the first key-value cache set.
[0076] During the decoding phase of the AI model, a query vector is generated based on the current decoding state, and the cached key vectors are read from the first key-value cache. An attention scoring function is used to calculate the similarity score between the query vector and each key vector, i.e., the attention score. This attention score reflects the semantic relevance between the current decoding position and each position in the input sequence.
[0077] In step S420, the first key-value cache set with an attention score greater than or equal to the target score is used as the target key-value cache. During the model decoding process, the key-value cache entries in the first key-value cache set are filtered based on the attention score between the query vector and the key vector. By setting a target score, the key-value cache entries with an attention score greater than or equal to the target score are retained to form the target key-value cache. The target key-value cache is considered to be the most relevant and semantically contributing part to the current decoding state.
[0078] Preferably, the above target score is 0.9.
[0079] In step S430 , an inference result of the first type of inference request is determined based on the target key-value cache.
[0080] When processing the first type of inference request, an attention-weighted sum is performed based on the target key-value cache to obtain the context vector for the current decoding time step. This context vector is then processed through modules such as the decoding layer, activation function, and output layer to ultimately generate the inference result. This significantly reduces computing resource consumption and inference latency without sacrificing the core semantic understanding capabilities of the task.
[0081] Optionally, during the pre-population phase, model inference is typically computationally intensive. Therefore, during the pre-population phase, a single batch of first-type inference requests or second-type inference requests is processed to determine the corresponding first key-value cache set or second key-value cache set. During the decoding phase, model inference typically incurs high communication overhead. Therefore, during the decoding phase, the first key-value cache set or second key-value cache set is processed in multiple batches to determine the corresponding inference results.
[0082] Optionally, when the GPU memory occupancy is less than a first target threshold and the CPU queue depth is greater than or equal to a second target threshold, the batch size is increased based on a sliding window mechanism. When the GPU memory occupancy is greater than or equal to a third target threshold, or the CPU queue depth is less than a fourth target threshold, the batch size is decreased based on a sliding window mechanism.
[0083] Specifically, the GPU memory occupancy rate and CPU queue depth are monitored in real time, and the average or current value of these indicators is calculated within a sliding window (for example, data from the past few seconds) to evaluate the current resource usage status of the system. If the GPU memory occupancy rate is lower than the first target threshold and the CPU queue depth is greater than or equal to the second target threshold, the batch size is increased based on the data trend within the sliding window to improve throughput; conversely, if the GPU memory occupancy rate is greater than or equal to the third target threshold, or the CPU queue depth is less than the fourth target threshold, the batch size is reduced to reduce resource consumption and prevent overload. Throughout the process, the changes in batch size are smoothly adjusted through the sliding window mechanism to avoid frequent adjustments due to instantaneous fluctuations, thereby ensuring system stability and response efficiency.
[0084] Optionally, during the AI model inference deployment process, different types of operators may have significant differences in computing characteristics, data flow control, and parallelism. Perform static analysis of the operators in the model to identify control-intensive and parallel-intensive operators. Assign control-intensive operators to the CPU for execution, assign parallel-intensive operators to the GPU for execution, or assign both control-intensive and parallel-intensive operators to the GPU for execution.
[0085] Specifically, control-intensive operators are characterized by complex control flow, low parallelism, and reliance on CPU instruction flow control. These include conditional judgment, branch control, or dynamic routing. Parallel-intensive operators are highly parallelizable and suitable for GPU parallel computing. These include matrix multiplication, convolution, or attention calculations.
[0086] Optionally, a mathematical optimization method is introduced in the attention calculation in the pre-filling stage and the decoding stage of the artificial intelligence model, which reduces the number of accesses to the high-bandwidth memory (HBM) by rearranging the calculation order.
[0087] The aforementioned mathematical optimization methods include the FlashAttention method. Specifically, the input query vector, key vector, and value vector are divided into multiple small blocks of fixed size, which are then processed in parallel on the GPU. The computation order is optimized to avoid unnecessary repeated calculations and reduce computational overhead. Furthermore, memory access patterns are adjusted to reduce the number of HBM accesses and communication overhead. Furthermore, knowledge distillation can be used to transfer knowledge from trained large models (such as GPT and BERT) to smaller models for inference, reducing computational resource consumption while maintaining high accuracy.
[0088] The inference request processing method provided in this embodiment can distinguish between first-type and second-type inference requests based on the type of inference request. In the pre-filling phase, it uses quantized computing, a sparse attention mechanism, or full computing to generate corresponding key-value caches. In the decoding phase, it further filters key information based on the attention scoring function or directly uses the full cache for inference. This enables adaptive processing of tasks of different precision levels within the same model architecture, improving the inference speed and resource utilization of low-precision requests while ensuring the output quality of high-precision requests. This overall optimizes inference efficiency, reduces computing resource consumption, and enhances the system's flexibility and adaptability in diverse application scenarios.
[0089] The following uses the intelligent customer service application scenario as an example to further illustrate the technical solution of the present disclosure. Those skilled in the art should understand that in addition to the intelligent customer service application scenario, it can also be applied to other application scenarios.
[0090] When a user initiates an inference request to the customer service system via mobile or web, the system first determines the request type based on the complexity of the request and the user's identity tag (e.g., regular user or VIP user). This classifies the request as either a Type 1 or Type 2 inference request. For example, a user asking about "business hours" is a simple question-and-answer task and is identified as a Type 1 inference request. However, a user requesting "contract terms interpretation" is a more specialized and complex task and is identified as a Type 2 inference request.
[0091] Then, during the pre-population phase of the AI model, different processing strategies are selected based on the request type. For the first type of inference request, the linear layer and sparse attention mechanism of quantized computing are enabled to perform lightweight processing on the input content, significantly reducing computing resource consumption while ensuring basic semantic understanding, and generating the first key-value cache set. For the second type of inference request, a full computation method is used without any compression or sparsification operations, ensuring that all context information is fully preserved, and generating the second key-value cache set to ensure high accuracy of the output results.
[0092] During the decoding phase of the AI model, different decoding strategies are selected based on the request type. For the first type of inference requests, an attention scoring function is used to evaluate the entries in the first key-value cache set, filtering out those with attention scores greater than the target score to form the target key-value cache. Responses are then quickly generated based on these key entries, shortening response times. For the second type of inference requests, the entire second key-value cache set is directly used for decoding, ensuring the comprehensiveness and accuracy of the output.
[0093] Finally, the generated inference results are returned to the user, completing the entire inference process. This inference process enables flexible support of inference tasks at different precision levels within the same model architecture, improving resource utilization efficiency while ensuring service quality.
[0094] Based on any of the above embodiments, the present disclosure also provides an inference request processing device.
[0095] Figure 6 It is a schematic block diagram of the structure of an inference request processing device according to an embodiment of the present disclosure.
[0096] like Figure 6 As shown, the inference request processing device includes: The inference request receiving module 610 receives the inference request information input by the user; The reasoning request identification module 620 performs intent recognition or identification recognition on the received reasoning request information, obtains a recognition result, and determines the request type of the reasoning request information based on the recognition result, which is a low-precision type reasoning request or a high-precision type reasoning request; Artificial Intelligence Model 630: If the request type is a low-precision inference request, Artificial Intelligence Model 630 uses a linear layer of quantized computation and a sparse attention mechanism for processing during the pre-fill phase of the inference process. During the decoding phase, the model determines the target key-value cache based on the attention scoring function to determine the corresponding inference result. If the request type is a high-precision inference request, the model uses full computation during the pre-fill phase of the inference process and uses the full key-value cache to determine the corresponding inference result during the decoding phase.
[0097] The above-mentioned inference request processing device can be in the form of computer software, and each module of the above-mentioned inference request processing device can be implemented by a computer software module.
[0098] The implementation process of the functions and effects of each module in the above-mentioned inference request processing device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0099] The present disclosure also provides an electronic device 1000 (corresponding to the inference request processing method). Figure 7 A schematic diagram showing a hardware implementation using a processing system is shown.
[0100] The hardware structure of an electronic device can be implemented using a bus architecture. The bus architecture can include any number of interconnecting buses and bridges, depending on the specific application and overall design constraints of the hardware. Bus 1100 connects various circuits including one or more processors 1200, memory 1300, and / or hardware modules. Bus 1100 can also connect various other circuits 1400, such as peripheral devices, voltage regulators, power management circuits, external antennas, etc. Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of illustration, the figure shows only one connecting line, but this does not mean that there is only one bus or only one type of bus.
[0101] For ease of explanation, some steps of the above method are described as corresponding to modules. It should be understood that the corresponding modules for performing one or more steps of the above method can be one or more hardware modules specifically configured to perform the corresponding steps, or implemented by a processor configured to perform the corresponding steps, or stored in a computer-readable medium for implementation by a processor, or implemented by some combination thereof.
[0102] The present disclosure also provides a storage medium having a computer program stored therein, which is used to implement the above-mentioned method when the computer program is executed by a processor. "Storage medium" can be any device that can contain, store, communicate, propagate or transmit a program for use with an instruction execution system, device or equipment or in conjunction with these instruction execution systems, devices or equipment. More specific examples of storage media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0103] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the process or function of the present disclosure is performed in whole or in part.
[0104] A computer program or instruction can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, the computer program or instruction can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The readable storage medium can be any accessible medium or a data storage device such as a server or data center that integrates one or more accessible media. The accessible medium can be a magnetic medium such as a floppy disk, hard disk, or magnetic tape; an optical medium such as a digital video disk; or a semiconductor medium such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0105] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0106] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, electronic devices, and computer program products according to the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0107] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0109] In the description of this specification, the description with reference to the terms "one embodiment / method", "some embodiments / methods", "example", "specific example", or "some examples" means that the specific features, structures, or characteristics described in conjunction with the embodiment / method or example are included in at least one embodiment / method or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / method or example. Moreover, the specific features, structures, or characteristics described may be combined in a suitable manner in any one or more embodiments / methods or examples. In addition, those skilled in the art may combine and combine different embodiments / methods or examples described in this specification and the features of different embodiments / methods or examples, unless they are contradictory.
[0110] Those skilled in the art will appreciate that the above embodiments are merely intended to clearly illustrate the present disclosure and are not intended to limit the scope of the present disclosure. Other changes or modifications may be made based on the above disclosure, and such changes or modifications are still within the scope of the present disclosure.
Claims
1. A method for processing an inference request, wherein the inference request is processed by a model, characterized in that: include: receiving an inference request, determining a request type for the inference request, and determining the inference request as a first type inference request or a second type inference request, wherein the first type inference request is a low-precision inference request, and the second type inference request is a high-precision inference request; In a pre-population phase of the model, the first type of inference request is processed using a linear layer of quantized computation and a sparse attention mechanism to determine a first key-value cache set corresponding to the first type of inference request, or the second type of inference request is fully computed to determine a second key-value cache set corresponding to the second type of inference request; and During the decoding phase of the model, an inference result of the first type of inference request is determined based on a target key-value cache in the first key-value cache set determined by an attention scoring function, or an inference result of the second type of inference request is determined based on all key-value caches in the second key-value cache set.
2. The inference request processing method according to claim 1, wherein: During the pre-population phase of the model, the first type of inference request is processed using a linear layer of quantized computation and a sparse attention mechanism, including: Convert the weights and activation values of the linear layer to low-precision values based on the target quantization accuracy; Perform forward propagation calculation on the first type of inference request through the quantized linear layer; Matrix operations are performed on the first type of inference request through a sparse attention mechanism.
3. The inference request processing method according to claim 1, wherein: Determining an inference result of the first type of inference request according to a target key-value cache in the first key-value cache set determined by an attention scoring function includes: During a decoding phase of the model, determining an attention score for a first key-value cache in the first set of key-value caches based on an attention scoring function; The first key-value cache set whose attention score is greater than or equal to the target score is used as the target key-value cache; An inference result of the first type of inference request is determined based on the target key-value cache.
4. The inference request processing method according to claim 3, wherein: During a decoding phase of the model, determining an attention score for a first key-value cache in the first set of key-value caches based on an attention scoring function includes: determining a query vector based on context information of the first type of reasoning request; Processing a first key-value cache in the first key-value cache set based on a linear layer and a self-attention mechanism in a decoding phase to determine a key-value cache entry of intermediate state information of the first key-value cache; the key-value cache entry includes a key vector and a corresponding value vector; The similarity between the query vector and the key vector is compared based on an attention scoring function to determine an attention score of a first key-value cache in the first key-value cache set.
5. The inference request processing method according to claim 1, wherein: Performing full computation on the second-type inference request to determine a second key-value cache set corresponding to the second-type inference request includes: performing a linear transformation on the second type of inference request by a target weight; Perform attention calculation on the second type of reasoning request after the linear transformation to determine a second key-value cache set of the second type of reasoning request.
6. The inference request processing method according to claim 1, wherein: Also includes: Processing the first type of inference request or the second type of inference request in a single batch in the pre-filling phase to determine the corresponding first key-value cache set or second key-value cache set; In the decoding stage, the first key-value cache set or the second key-value cache set is processed in multiple batches to determine a corresponding inference result.
7. The inference request processing method according to claim 6, wherein: Also includes: When the GPU memory occupancy is less than the first target threshold and the CPU queue depth is greater than or equal to the second target threshold, the batch size is increased based on the sliding window mechanism; When the GPU memory occupancy is greater than or equal to a third target threshold, or the CPU queue depth is less than a fourth target threshold, the batch size is reduced based on a sliding window mechanism.
8. The inference request processing method according to claim 1, wherein: Also includes: Performing static analysis on operators in the model to identify control-intensive operators and parallel-intensive operators; Assign control-intensive operators to the CPU for execution, and assign parallel-intensive operators to the GPU for execution.
9. An electronic device, characterized in that: include: a memory storing execution instructions; as well as A processor, wherein the processor executes the execution instruction stored in the memory, so that the processor executes the inference request processing method according to any one of claims 1 to 8.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the inference request processing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Model reasoning method and device based on key value matrix cache and medium
CN118036754A
Big language model-based reasoning method and device, electronic equipment and storage medium
CN119168054A
Memory pooling method and system for model reasoning acceleration and computer program product
CN120525063A
Method and apparatus for inference in large language model
WO2025152397A1