Model quantitative reasoning acceleration method and device, equipment and medium
By dividing the input text into multiple processing blocks and calculating the token importance score based on the self-attention matrix for accuracy allocation, the problem of video memory overhead in long text inference is solved, and efficient and accurate inference effect is achieved.
Patent Information
- Application Number
- CN202510525474.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing technology has memory overhead problems in the long text inference process of large language models, especially in terms of the efficiency, importance recognition mechanism and block-level reasoning resource scheduling of quantitative configuration strategies, which is difficult to meet the needs of efficient and accurate reasoning in the fields of financial technology and medical health.
By dividing the input text into multiple processing blocks, fixing the calculation accuracy format of the first processing block into a high-precision format and disabling quantization processing, generating a self-attention matrix through the language model for other processing blocks, calculating the importance score of each token, and allocating the accuracy format based on the threshold. Then, count the number of tokens in each processing block, select the most accurate format as the unified quantization configuration, and reuse the configuration within the network module group, perform block-level batch quantization and complete model inference.
While ensuring inference accuracy, it greatly reduces the overhead of video memory and configuration time, effectively improving the execution efficiency and memory utilization rate in long text inference tasks.
Smart Images

Figure CN120086355A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and storage medium for accelerating model quantization inference. Background Art
[0002] In recent years, large language models (LLMs) have performed excellently in natural language processing tasks and are widely used in scenarios such as dialogue systems, machine translation, and question answering systems. However, with the continuous expansion of model scale, they face significant video memory occupancy problems when performing long text inference tasks. Most mainstream LLMs have billions or even hundreds of billions of parameters, and during the inference process, the model and related intermediate states need to be fully loaded into the GPU video memory, which poses extremely high requirements on hardware resources.
[0003] When applied to long text inference tasks, such as long document generation, complex literature abstract extraction, or multi-round dialogue processing, the input sequence often contains thousands to tens of thousands of tokens. Such extremely long sequences significantly increase the computational complexity and video memory requirements of the attention mechanism. Specifically, the dimension of the attention matrix grows with the square of the input sequence length, and a large number of intermediate activation values also need to be temporarily stored in the video memory, which is extremely likely to cause memory overflow or inference failure. Existing video memory optimization means, such as gradient checkpointing, activation value offloading, etc., are mainly for the training stage and are difficult to be directly applied to the inference process. In addition, although distributed inference can alleviate the resource bottleneck of a single device, it has high requirements for infrastructure, complex configuration in actual deployment, and is difficult to operate stably in edge or medium-low resource environments.
[0004] In the field of medical and health business, LLMs are used in scenarios such as medical record abstract generation, medical question answering systems, and diagnosis and treatment record analysis. These tasks often involve a large number of medical terms and context-related information, and have a stronger dependence on the model's inference ability for long texts. Although the current general quantization compression method can reduce the video memory occupancy, it often ignores the retention of some key information in medical texts, resulting in a decrease in model accuracy and even a deviation in the understanding of medical information, affecting the credibility of the results.
[0005] In the field of fintech business, large language models are widely used in text-intensive tasks such as contract parsing, financial summary generation, and customer risk assessment. These tasks usually involve compliance reports or historical transaction records with complex structures and extremely long texts, which pose extremely high requirements on the video memory efficiency during the inference process. However, the current model is prone to failure due to insufficient video memory when processing long text inputs such as financial documents, seriously affecting the stability and scalability of the model.
[0006] To alleviate the video memory pressure, quantization technology has become one of the mainstream compression means, especially widely adopted in the inference scenario. Existing research attempts to use the mixed-precision quantization method, that is, allocating higher bit widths to high-importance tokens and lower bit widths to low-importance tokens, so as to reduce the overall video memory usage while maintaining the model accuracy. However, the mixed-precision quantization faces the problem of low configuration selection efficiency during the inference process. Since the current method needs to dynamically analyze the importance of each token and determine the bit width strategy during each round of inference, this process itself is time-consuming and weakens the speed advantage brought by quantization, especially more significantly in long text tasks.
[0007] In summary, there are still key deficiencies in the existing technology in dealing with the video memory overhead problem during the long text inference process of large language models, especially in aspects such as the efficiency of the quantization configuration strategy, the importance recognition mechanism across task domains, and the block-level inference resource scheduling. Further optimization is still needed to meet the actual requirements of efficient and accurate inference in fields such as fintech and healthcare. Summary of the Invention
[0008] The main objective of the present invention is to provide a model quantization inference acceleration method, device, equipment, and storage medium, aiming to solve the technical problem that the existing technology needs to determine the quantization bit width for each token during the inference process, resulting in low quantization configuration efficiency, especially seriously affecting the inference speed and video memory optimization effect in long text inference tasks.
[0009] To achieve the above objective, the present invention provides a model quantization inference acceleration method, including: Dividing the input text into multiple processing blocks, fixing the processing precision format of the first processing block as the high-precision format, and disabling the quantization processing of the first processing block; For the other processing blocks except the first processing block among the multiple processing blocks, generating the self-attention matrix of each other processing block through a language model, determining the sum of all element values in the column corresponding to each token position in the self-attention matrix, and using the sum of all element values as the importance score of each token position; Allocating the token positions with importance scores greater than the first threshold to the high-precision format, allocating the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and allocating the token positions with importance scores less than the second threshold to the low-precision format; Counting the number of token positions allocated to the high-precision format, medium-precision format, and low-precision format in each processing block, and selecting the precision format with the largest number as the unified quantization configuration of the corresponding processing block; Divide the network modules of the language model into multiple configuration sharing groups, where each configuration sharing group contains at least two network modules; Within each configuration sharing group, share the unified quantization configuration of the processing block corresponding to the first network module with other network modules within the same configuration sharing group; According to the unified quantization configuration corresponding to each processing block, perform block-level batch quantization on all processing blocks and complete model inference to generate a model inference result.
[0010] Furthermore, to achieve the above object, the present invention provides a model quantization inference acceleration device, including: An input text preprocessing module, configured to divide the input text into multiple processing blocks, fix the processing precision format of the first processing block to a high-precision format, and disable the quantization processing of the first processing block; A self-attention analysis module, configured to generate a self-attention matrix for each of the other processing blocks except the first processing block among the multiple processing blocks through a language model, determine the sum of all element values in the column corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score for each token position; An accuracy allocation module, configured to allocate token positions with importance scores greater than a first threshold to a high-precision format, allocate token positions with importance scores above a second threshold and below the first threshold to a medium-precision format, and allocate token positions with importance scores less than the second threshold to a low-precision format; A quantization configuration decision module, configured to count the number of token positions allocated to the high-precision format, medium-precision format, and low-precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration for the corresponding processing block; A network module grouping control module, configured to divide the network modules of the language model into multiple configuration sharing groups, where each configuration sharing group contains at least two network modules; A configuration sharing management module, configured to share the unified quantization configuration of the processing block corresponding to the first network module with other network modules within the same configuration sharing group; An inference execution module, configured to perform block-level batch quantization on all processing blocks according to the unified quantization configuration corresponding to each processing block and complete model inference to generate a model inference result.
[0011] Furthermore, to achieve the above object, the present invention also provides a computer device, which includes a memory, a processor, and a model quantization inference acceleration program stored in the memory and executable on the processor. When the model quantization inference acceleration program is executed by the processor, it implements the steps of the model quantization inference acceleration method as described above.
[0012] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium with a model quantization inference acceleration program stored thereon. When the model quantization inference acceleration program is executed by a processor, it implements the steps of the model quantization inference acceleration method as described above.
[0013] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health and fintech. It discloses a model quantization inference acceleration method, which includes: dividing the input text into multiple processing blocks, fixing the calculation precision format of the first processing block as a high-precision format and disabling quantization operations; for other processing blocks except the first one, generating a self-attention matrix through a language model and calculating importance scores based on the sum of the values in the corresponding columns of each token position in the self-attention matrix; allocating each token position to a high-precision, medium-precision, or low-precision format based on two preset thresholds; counting the number of tokens in each precision format within each processing block and selecting the precision format with the largest number as the unified quantization configuration; dividing the network module into multiple configuration sharing groups and sharing the unified quantization configuration of the processing blocks within the group; performing block-level batch quantization according to the unified quantization configuration of the processing blocks and completing model inference to generate inference results. The present invention uniformly determines the quantization configuration of each processing block based on the token importance score and reuses this configuration within the network module group, achieving block-level precision allocation and parallel quantization inference, significantly reducing the video memory overhead and configuration time overhead while ensuring the inference precision, and effectively improving the execution efficiency and video memory utilization rate in long text inference tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The following will further illustrate the present invention in conjunction with the drawings, where: Figure 1 is a schematic diagram of an application environment of the model quantization inference acceleration method in an embodiment of the present invention; Figure 2 is a schematic flowchart of an embodiment of the model quantization inference acceleration method of the present invention; Figure 3 is a schematic diagram of the functional modules of a preferred embodiment of the model quantization inference acceleration device of the present invention; Figure 4 is a schematic diagram of the structure of a computer device in an embodiment of the present invention; Figure 5Another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners
[0015] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0016] The model quantization inference acceleration method provided by the embodiments of the present invention can be applied to an application environment such as Figure 1 , where the client communicates with the server through a network. The server can divide the input text into multiple processing blocks through the client, fix the calculation precision format of the first processing block as a high-precision format and disable the quantization operation; for other processing blocks except the first processing block, generate a self-attention matrix through a language model, and calculate the importance score based on the sum of the values in the corresponding column of each token position in the self-attention matrix; allocate each token position to a high-precision, medium-precision or low-precision format based on two preset thresholds; count the number of tokens in each precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration; divide the network module into multiple configuration sharing groups, and share the unified quantization configuration of the processing block within the group; perform block-level batch quantization according to the unified quantization configuration of the processing block and complete model inference to generate an inference result. The present invention realizes precision allocation and parallel quantization inference at the block level by uniformly determining the quantization configuration of each processing block based on the token importance score and reusing this configuration within the network module group, greatly reducing the video memory overhead and configuration time overhead while ensuring the inference precision, and effectively improving the execution efficiency and video memory utilization rate in the long text inference task. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail through specific embodiments below.
[0017] Please refer to Figure 2 , Figure 2 which is a flowchart of an embodiment of the model quantization inference acceleration method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described herein may be executed in a different order.
[0018] As Figure 2 shown, the model quantization inference acceleration method proposed by the present invention includes the following steps: S10, divide the input text into multiple processing blocks, fix the processing precision format of the first processing block as a high-precision format, and disable the quantization processing of the first processing block; In this embodiment, dividing the input text into multiple processing blocks is to control the distribution of computing resources during model inference for long text inputs, and to perform local management of the activation values, attention matrices, and video memory usage of the model. In practical applications, text processing blocks can be generated through a sliding window of a fixed length. For example, every 512 tokens are divided into one block, and this length is determined according to the maximum processing length of the training window or inference window of the model. This division method comes from the window sharding mechanism widely used in sequence modeling, and its purpose is to transform long text into local context blocks, which not only ensures local semantic coherence but also avoids the computational bottleneck caused by loading all tokens at once.
[0019] Fixing the processing precision format of the first processing block to a high-precision format is to provide a stable computational starting point for the model in the initial stage of inference. The high-precision format generally refers to the floating-point format, such as FP32 or FP16, which has stronger robustness and fault tolerance in terms of numerical expression range and expression precision compared to low-bitwidth fixed-point formats (such as INT8 or INT4).
[0020] Disabling the quantization process for the first processing block is to further ensure that this segment of the input runs in full-precision representation, avoiding the cumulative interference of quantization errors on the propagation paths of subsequent blocks in the initial stage of model propagation. Disabling quantization not only includes not applying quantization operations to the weights and activation values of this processing block but also skipping the quantization configuration generation process for this processing block. Its core purpose is to provide a reference benchmark in quantization decision-making, making the first processing block a relative anchor point for the importance scoring and precision configuration judgment of subsequent blocks.
[0021] In practical applications, the quantization process usually involves discretizing the weight parameters of the model and compressing the activation values into a finite bitwidth representation domain. If this operation is still performed on the first processing block, it will cause the initial attention result to be distorted, thereby reducing the accuracy of precision allocation for subsequent token positions. Therefore, disabling the quantization process for the first processing block is not only a consideration of numerical stability but also a fundamental premise for the dynamic precision adjustment strategy in the entire inference path.
[0022] Since the model usually adopts residual connections and normalization mechanisms, the high-precision first block helps to stabilize the statistical behavior of the first few layers, thereby improving the discrimination accuracy of subsequent processing blocks in quantization strategy decision-making. This feature also has good generalization and does not depend on the specific model structure, so it can be applied to large language models with different architectures such as GPT, T5, and BERT.
[0023] In a specific implementation, after loading the input text as a token sequence, it can be segmented according to a preset processing block length, and each segment forms a processing block. The token range of the first processing block can be the first 512 tokens of the token sequence. After the division is completed, configure the computing precision format of this processing block as FP16 or FP32. During model inference, skip the quantization configuration generation process of this block, that is, do not participate in importance scoring and do not perform bit-width decision-making. To disable quantization processing, when calling the model for execution, for the first processing block, forcefully bypass the quantization operator deployed in the model. This can be achieved by configuring the quantization mask parameter of the inference engine, or during the model graph conversion process, setting the data path of this block to a non-quantized path. Further, when this processing block passes through the embedding layer, multi-head attention layer, and feed-forward layer of the model, all can be executed with floating-point precision, without calling the low-precision computing kernel function. In an actual model running environment, if using an inference framework such as TensorRT, the first block can be marked as a resident high-precision area when constructing the engine, or the first block precision locking instruction can be added during compilation. In the multi-block execution process, the output of the first processing block can also be used as a reference input for the importance analysis and precision allocation strategy generation of subsequent processing blocks.
[0024] Example illustration: In the field of medical and health business, when processing text containing long medical diagnosis records, it is often the case that key disease descriptions appear at the beginning of the text. If the first processing block loses information due to quantization, it may cause the model to fail to correctly extract disease symptoms and corresponding analyses. After adopting fixed high-precision processing for the first block, the model can more accurately capture the high-value content in the first paragraph, which is beneficial for subsequent generation of accurate diagnostic summaries or disease course predictions.
[0025] In the field of fintech business, when processing a complete risk report, the beginning of the report usually contains global risk classification and summary information. If it is placed in a low-precision computing path, it is easy to cause the model to misjudge the risk level, thus affecting the subsequent judgment process. By maintaining high-precision computing and disabling quantization for the first processing block, the core points of the report can be effectively captured, ensuring the accuracy and stability of the risk control model inference, and enhancing the reliability and security of financial service decisions.
[0026] By fixing the first processing block as a high-precision format and disabling its quantization processing during the inference process, the stability and accuracy of subsequent processing blocks in precision strategy discrimination can be significantly improved. The high-precision first block provides complete context representation ability, making its output can be used as a benchmark for weight allocation and quantization level selection, while avoiding model deviation caused by quantization errors in the initial information propagation. In the long text scenario, this strategy can effectively reduce the global precision degradation problem caused by misjudgment in the first paragraph, thereby enhancing the overall inference performance and stability.
[0027] S20. For other processing blocks except the first processing block among the multiple processing blocks, generate the self-attention matrix of each other processing block through a language model, determine the sum of all element values corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score of each token position. In this embodiment, for other processing blocks except the first processing block among the multiple processing blocks, the self-attention mechanism is introduced to generate the corresponding self-attention matrix, which is mainly used to measure the information dependence relationship between different token positions within the same processing block. In natural language processing tasks, large language models (LLMs) usually adopt the self-attention mechanism based on the Transformer structure to construct the context representation between tokens. By separately constructing the self-attention matrix for each processing block, the calculation scope can be localized, thereby effectively controlling the memory occupancy and improving the calculation efficiency.
[0028] The sum of all element values corresponding to each token position is used as the importance score. This operation is essentially a column vector summation, and its technical meaning is to measure the overall attention concentration degree of a certain token in this processing block. The larger this value is, the stronger the connection between this token and other tokens, and the wider the spread range of its semantic information within the block, so it is considered more "important".
[0029] In self-attention calculation, the column vector corresponding to each token position represents the attention weights assigned by this token as the target token to all source tokens. On this basis, summing the column vectors can obtain its global attention degree. The importance score obtained in this way does not depend on specific task labels, so it has good generality and is applicable to the requirements of token selective precision control in various tasks such as machine translation, question answering systems, and information extraction.
[0030] To avoid introducing unnecessary noise and redundant calculations, padding tokens should be skipped when generating the self-attention matrix, that is, a mask strategy can be adopted for the padding token positions to force their corresponding attention values to zero, preventing them from interfering with the importance evaluation of other tokens. This design is also applicable in the multi-head attention mechanism. The local matrices calculated by each attention head can be first averaged and fused into the final self-attention matrix.
[0031] In a specific embodiment, the token sequence of each processing block is fed into a multi-head self-attention module for forward propagation to obtain the attention matrices of multiple attention heads. By performing weighted average or simple arithmetic average on these local matrices, the final self-attention matrix of the current processing block is generated. Subsequently, each column of this matrix is traversed, and the sum operation is performed on all the numerical values in the column vector as the importance score of the token position corresponding to this column. During this process, the column vector corresponding to the padding token will be replaced with a vector of all zeros, so that its importance score is always zero, ensuring that it does not participate in the subsequent selection of precision configuration.
[0032] For scenarios with high computational performance requirements, the generation of the self-attention matrix and the column vector summation operation can be fused into a GPU kernel function (CUDA Kernel) to batch process multiple processing blocks under the GPU parallel framework, improving the overall processing efficiency. In addition, for the case where the numerical precision in the attention matrix is relatively low (such as FP16), a normalization or numerical smoothing mechanism can be introduced to avoid misjudgment caused by the influence of numerical drift on the token scores.
[0033] Example illustration: In the field of medical and health, texts such as doctors' consultation records, patients' self-reports, and test reports often contain a large number of tokens, and the words that truly affect subsequent decisions are only a very small number of highly specialized and contextually highly relevant medical terms. Taking an electronic medical record as an example, it may contain information segments such as patient age, underlying diseases, chief complaints, and test results. When generating the self-attention matrix of this processing block, for example, for the phrase "abnormal liver function", the token positions corresponding to it in the matrix show strong attention connections with the token positions such as "elevated ALT", "jaundice", and "positive hepatitis B surface antigen". By calculating the sum of the elements in the column corresponding to this token, its semantic importance in the entire segment can be quantified. Furthermore, these high-score tokens will be marked as important tokens, providing higher computational precision support for subsequent diagnostic classification or automatic generation of medical record summaries. For tokens with high generality but low context semantic weight, such as "patient gender" and "this follow-up visit", the sum of the elements in the corresponding column is smaller, and naturally, they are given lower importance scores, so that they are calculated with lower precision in subsequent inferences, saving computing power resources without affecting the expression of the core semantics.
[0034] In intelligent customer service conversations, user complaint handling, or transaction log parsing in the financial field, important context information is usually distributed in tokens containing words such as "risk", "freeze", "fraud", "delayed receipt of funds", etc. Moreover, the context dependence of such words is usually relatively high. In the self-attention matrix, the token "fraud" forms strong connections with tokens such as "transfer failure", "funds frozen", "unknown contact", etc. The sum of the elements of its corresponding column vector is significantly higher than that of other background words such as "hello", "excuse me", "thank you", etc., and thus it is marked as a high-importance position in this step. In the subsequent inference of the risk judgment model, the key tokens are assigned high-precision calculations to ensure that the understanding of sensitive expressions will not be distorted due to low-bitwidth quantization, which is particularly crucial in user inputs with fuzzy semantic boundaries and non-standard expressions. Through this attention-based token-level precision allocation strategy, it is possible to effectively compress the video memory occupancy and computational load during inference without sacrificing the sensitivity and accuracy of risk control.
[0035] These example scenarios show that in long-text tasks, by capturing the semantic coupling degree between tokens through the self-attention matrix and determining importance based on the sum of column vectors, not only the technical goal of token-level precision control is achieved, but also it has general adaptability across domains and high inference efficiency guarantee.
[0036] By summing each column of the self-attention matrix of each processing block and calculating the importance score, it is possible to accurately identify the token positions that contribute the most to semantic expression based on the context dependence strength of the tokens. When further performing precision allocation on this basis, it no longer relies on external feature extraction or complex rule matching, significantly reducing the pre-processing computational overhead in the inference stage and achieving the unity of the identification of key tokens and precision control.
[0037] S30, assign the token positions with importance scores greater than the first threshold to the high-precision format, assign the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and assign the token positions with importance scores less than the second threshold to the low-precision format; In this embodiment, using the importance score of each token position to drive the configuration of subsequent calculation precision is the key path to implementing the differential quantization strategy. The "importance score" here comes from the numerical representation generated based on the self-attention matrix in the previous stage, which is used to reflect the information influence of each token in its context. The higher this score, the higher the degree to which the token is concerned by other tokens in the current processing block, that is, the stronger its role in semantic propagation and the lower the tolerance for semantic masking. Therefore, this value directly determines the calculation precision required for the token in inference.
[0038] High-precision formats usually correspond to floating-point computing formats, such as half-precision floating-point (FP16) or single-precision floating-point (FP32), which achieve a balance between computational complexity and expression accuracy. Medium-precision formats can be mid-width fixed-point quantization, such as INT4, which can save some video memory resources and computational overhead while ensuring the accuracy of the main semantic expression. Low-precision formats usually adopt fixed-point quantization formats with smaller bit widths, such as INT2 or further INT1 (binary), which are applicable to token positions with higher redundancy and strong semantic fault tolerance, significantly compressing the computational load.
[0039] Hierarchical control is achieved by setting a first threshold and a second threshold, where the first threshold is greater than the second threshold. The second threshold is used to identify the tokens with the lowest importance, while the first threshold identifies the most important batch of tokens. The tokens between the two fall into the medium-precision range. The setting principles of the first threshold and the second threshold can be dynamically set according to the processing block length, model scale, and current hardware resources, or determined through empirical model training. This dual-threshold mechanism enables the precision configuration to have both controllable policy flexibility and fine-grained scheduling ability for computational resources.
[0040] The process of comparing the scores with the two thresholds can usually be executed in parallel in the weight calculation thread during the preprocessing stage. The precision level of each token position can be determined through a single floating-point comparison operation, and then it is labeled with the corresponding precision label for subsequent configuration generation and inference allocation. The entire allocation mechanism does not depend on the modification of the model structure and only acts on the quantization strategy layer, with good system compatibility and deployment adaptability.
[0041] Hierarchical precision configuration can be achieved through static threshold setting. For example, the second threshold is set to 0.2, and the first threshold is set to 0.8. The token importance scores are uniformly distributed in the range of [0, 1]. Then, tokens greater than 0.8 are assigned FP16, the middle section (between 0.2 and 0.8, including 0.8 and 0.2) uses INT4, and those below 0.2 are INT2. Dynamic threshold calculation can also be based on the score distribution of each processing block. For example, the quantile method is used, with the top 10% as high precision, the bottom 30% as low precision, and the middle as medium precision. In addition, task-aware factors can be incorporated to preferentially label known important tokens in specific semantic annotation tasks as high precision, that is, to enhance precision under the guidance of attention.
[0042] In specific implementation, a precision label is recorded at each token position (for example, a 2-bit identifier: 00 for low precision, 01 for medium, and 10 for high precision), stored in the video memory in the form of a bitmap. After the quantization configuration module reads this identifier, it calls the quantization kernel function with the corresponding bit width to complete the binding of low-level inference. Multiple tokens can be batch-labeled and distributed concurrently with the help of the SIMD (Single Instruction Multiple Data) instruction set to improve the throughput efficiency. For a distributed scenario, the precision allocation and labeling can be locally completed on each GPU, and the configured bitmap can be uploaded to the shared control module to achieve cross-block consistent scheduling.
[0043] By allocating the precision format according to the importance score at the token level, the key inclination of computing resources can be achieved, that is, higher computing power is used for key tokens, and the precision resource investment in low-weight information is reduced. While maintaining a high inference accuracy of the overall model, the video memory overhead and inference latency are effectively controlled.
[0044] S40, count the number of token positions allocated as high-precision format, medium-precision format, and low-precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration for the corresponding processing block; In this embodiment, after the multi-level precision allocation is completed, it is necessary to count the distribution of different precision formats within each processing block to provide a decision basis for the subsequent selection of the unified quantization configuration. Each token position has been allocated as high-precision format, medium-precision format, or low-precision format after the previous importance score comparison. The goal of this step is to count the number of token positions corresponding to the three types of precision formats by traversing the valid token positions of the entire processing block.
[0045] The quantization configuration of each processing block must be unified to match the scheduling requirements of block-level parallel execution in the hardware computing unit. For example, when the Tensor Core of the GPU executes calculations such as INT8 or FP16, it is required that each operation targets a unified matrix precision. Therefore, it is necessary to pre-determine the overall computing format of the processing block. This step determines the dominant precision distribution of the current processing block through the precision statistics results after quantization.
[0046] During the traversal process, invalid token positions do not participate in the statistics to avoid the deviation of the real distribution caused by padding tokens. In actual implementation, the mask recorded during the padding stage is usually used to quickly exclude the padding area, and only the precision labels of the valid positions are counted. After the counting is completed, compare the quantities of the three precision formats, and select the precision format with the largest number as the unified quantization configuration of this processing block.
[0047] In some special cases, if the number of tokens in two or three precision formats is the same and reaches the maximum value, to ensure the stability and conservativeness of the model inference accuracy, the format with a higher precision level is preferentially selected. For example, when the number of high-precision and medium-precision tokens is equal, the high-precision format is selected as the quantization configuration for this processing block. This conservative bias strategy ensures that the model retains the information expression ability of key semantic paths as much as possible without increasing the computational burden.
[0048] This operation is not only a configuration strategy based on statistical dominant trends but also a technical mechanism to ensure the consistency of the internal quantization strategy of the processing block. Through this mechanism, each processing block is assigned a unique unified quantization configuration, so that the corresponding kernel function can be called in subsequent inferences and hardware-level block parallel acceleration can be achieved.
[0049] A bitmap structure can be used to represent the precision labels at each token position. Each bit represents the precision category of the current token. For example, two binary identifiers are used to represent high, medium, and low precision respectively. During the statistical stage, logical AND operations are used to extract the precision flag bits of each processing block, and invalid token positions are skipped according to the mask, and the bit counts of the three precision markings are performed respectively.
[0050] In a multi-core system, the statistical operations can be executed in parallel in chunks. Each thread processes the statistical tasks of one processing block and writes the results to the shared memory or the quantization configuration register table. After the three types of count values are saved in the shared memory, a comparison operation is performed to determine the maximum value, and the precision configuration identifier is generated according to the priority strategy. If a dynamic precision adjustment strategy is adopted, the dynamic quantization bitwidth setting of the current processing block can be adjusted synchronously after the statistics are completed.
[0051] To further improve the intelligence and adaptability of the unified quantization configuration of the processing block, on the basis of selecting the precision configuration based on the statistical results of the precision format, multiple semantic and historical context features can be introduced to perform weighted fine-tuning on the statistical count results to reflect the sensitivity of the language model to local semantic changes and dynamically adapt to the actual computational resource requirements of different semantic regions in the inference task.
[0052] First, when processing long text input in natural language, there is a phenomenon that the semantic density of tokens changes with the context semantic field and presents an uneven distribution. Taking highly structured documents such as legal contracts and medical test reports as examples, their text fragments may contain a large number of repeated, high-frequency and information-redundant tokens (such as terms, connectors, unit names, etc.). Although such tokens appear frequently in the literal sense, they often do not constitute the focus of information in actual semantics. If the precision format is divided only based on the original importance score calculated by attention, the number of medium or low-precision tokens may be too small in the statistical process, thereby mistakenly configuring such processing blocks as high-precision, increasing computational redundancy.
[0053] To this end, a high-frequency token importance suppression mechanism based on a sliding window can be introduced. This mechanism counts the frequency of tokens that appear in multiple adjacent processing blocks during the statistical process. If the frequency of a token in the sliding window exceeds the preset threshold (such as appearing in 70% of the blocks), the importance score of the token is appropriately reduced by a proportional factor (such as multiplying by 0.8), so that it is more inclined to be classified as a medium or low-precision format during statistics, avoiding high-frequency redundant tokens dominating the overall accuracy configuration selection. The sliding window can be a processing block sequence of length 3 to 5, and the window length can be dynamically adjusted according to the total input length to take into account both response speed and local feature capture capabilities.
[0054] Secondly, the influencing factor of the previous semantic block can be introduced. Since there is often strong semantic coherence between processing blocks, such as in financial public opinion analysis, the "credit card overdue" processing block is followed by the "collection, default, credit investigation" semantic block. When calculating the accuracy, the accuracy tendency in the previous processing block should be referred to, so as to enhance the accuracy configuration strategy of the current block and the semantic consistency of the previous block. In terms of implementation, the proportion of the number of tokens in various precision formats in the previous processing block can be recorded, and a weighting factor can be introduced based on the precision count of the current processing block for adjustment. For example, the current number of medium-precision tokens is 200. If the proportion of high-precision tokens in the previous block is high, the current high-precision count can be multiplied by a forward adjustment factor of 1.1 to increase its configuration selection possibility.
[0055] Furthermore, for the specific requirements of certain tasks, a context weight awareness model can also be introduced to dynamically generate the precision configuration weight vector of the current processing block through a lightweight forward network module (such as one-layer FFN + Softmax). The input features include the original three-category precision token count of this block, the high-frequency token ratio of the sliding window, the precision distribution of the previous block, etc. The relative weights of each category of precision are adjusted through the weights learned by the forward model, so as to more finely control the selection logic of the final precision configuration. For example, in the scenario of medical and health text generation, this awareness model can identify the points of "pathological state transition" or "change in treatment response" and tend to select a higher precision configuration.
[0056] Finally, when the numerical values of the high, medium, and low three-category precision format counts are close, to avoid the problem of discontinuous context precision caused by quantization jitter, a dynamic gating mechanism can be introduced. The gating module judges whether the precision configuration distribution of the current processing block is balanced based on the precision distribution entropy (such as Shannon entropy or Top-k ratio). When it is determined that there is "no obvious dominant precision", the medium precision configuration is forcibly selected to achieve soft regulation of the overall computing resources of the system. This mechanism is particularly applicable to large-scale model deployment scenarios and can effectively control the power consumption peak and improve the overall inference throughput of the model.
[0057] Example illustration: In the field of medical and health, when facing the task of medical record abstract, the input may include content such as diagnostic conclusions, symptom descriptions, treatment plans, etc. For example, the processing block contains "The patient has intermittent headache and blurred vision, suspected of increased intracranial pressure". In the previous steps, "headache", "blurred", and "intracranial pressure" are marked as high precision, "patient", "has", "suspected", etc. are marked as medium precision, and the rest of the conjunctions are marked as low precision. During the statistics process, it is found that the number of medium-precision tokens is slightly more than that of high-precision tokens, but the numbers of high-precision and medium-precision tokens are close. According to the conservative strategy, the system selects high precision as the unified configuration of the processing block to ensure the inference precision of the subsequent diagnostic model.
[0058] In the financial business scenario, when analyzing the user's transaction history behavior to generate risk predictions, a certain processing block includes "The user transferred money three times today, and each amount exceeded the limit". In the importance scoring stage, "transfer", "limit", and "amount" are assigned high-precision formats, and the rest are medium or low precision. After precision statistics, the number of high-precision tokens dominates, and the system configures this processing block as FP16 precision and uniformly calls the floating-point kernel function for processing. This strategy ensures that the system can accurately identify and early warn of abnormal transaction patterns.
[0059] By counting the number of tokens in each precision format and selecting configurations using the principle of the largest proportion, the standardization of the precision strategy within the processing block is effectively achieved, enabling parallel quantization calculations based on the unified block-level configuration in the subsequent inference stage, improving the system execution efficiency and simplifying the kernel function scheduling path. At the same time, the precision priority rule is introduced to ensure that the semantic expression ability of the model is not damaged first in case of configuration conflicts, thus maintaining the model precision stability while saving resources.
[0060] S50, divide the network modules of the language model into multiple configuration sharing groups, and each configuration sharing group contains at least two network modules; In this embodiment, a language model is usually composed of multiple continuously stacked network modules. Each module includes components such as an embedding layer, a feed-forward network layer, and a multi-head self-attention layer. To control the video memory overhead and improve the inference efficiency in the context of long text input, these network modules can be divided into several "configuration sharing groups" in the processing order. Modules within each group share the same precision configuration and quantization parameters to reduce unnecessary repeated calculations.
[0061] The processing order means grouping according to the natural front-back arrangement of network modules in the model execution path, rather than randomly shuffling or dynamically organizing based on task allocation. This order-based partitioning method can ensure the stability and simplicity of the precision configuration propagation logic. Each configuration sharing group contains several adjacent network modules, and the number can be determined by a preset parameter, which can be set according to dimensions such as the depth of the model, the amount of calculation per layer, and available computing resources. Common preset numbers include 4, 8, 12, etc., and the specific value can be determined through experimental optimization. Each configuration sharing group contains at least two network modules.
[0062] The core of introducing the configuration sharing group is to form an execution structure of "block-level inference - group-level reuse". By unifying the quantization schemes of all modules in the configuration sharing group, the number of quantization configurations that need to be independently managed during execution can be significantly reduced, and the video memory and scheduling overhead of frequently loading configurations by different modules in model inference can be reduced. This strategy of partitioning in order and setting a fixed number within the group can be divisible by the number of modules in terms of structure, thus achieving unified scheduling and allocation in hardware orchestration.
[0063] It should be noted that the configuration sharing group and the "processing block" generated by chunking the previous input text are two-dimensional structures: the former acts inside the model structure for optimizing the quantization configuration between network modules; the latter acts on the organizational structure of the input data for controlling the input length and calculation partitioning. A connection is established between the two through a binding mechanism (such as the first module in the group binding to the corresponding processing block).
[0064] For example, a large language model with 96 network modules can be divided into 12 configuration sharing groups, with each group containing 8 consecutive modules. The division is carried out in the order of module numbers. For example, Group_1 includes Module_1 to Module_8, Group_2 includes Module_9 to Module_16, and so on. During the configuration phase, a unique group identifier is assigned to each sharing group, and a mapping relationship between the group identifier and the bound processing block is established.
[0065] During the execution phase, when the model inference task reaches a certain configuration sharing group, the system first reads the corresponding unified quantization configuration parameters of the group from the bound processing block and maps the parameters for use by all modules within the group. No separate quantization configuration initialization is performed for all modules within the group, thus saving time and storage overhead.
[0066] In large deep models in special scenarios, due to the non-uniformity of network modules at different depths in terms of semantic modeling strength, information flow characteristics, and quantization sensitivity, using a unified number division to configure sharing groups may not achieve the best balance between accuracy and efficiency. In this case, a non-uniform division method can be introduced, and by flexibly adjusting the number of network modules included in each configuration sharing group, the division strategy can better fit the model hierarchical structure and specific task requirements.
[0067] For example, modules near the input side usually undertake underlying semantic encoding tasks (such as position embedding, lexical modeling, etc.). Their computational structures are relatively simple and are not sensitive to accuracy loss. Therefore, multiple adjacent modules can be combined into larger configuration sharing groups to improve the parameter reuse efficiency and reduce the configuration switching frequency. On the contrary, modules near the output side mainly complete high-level semantic abstraction and decision information integration. These modules are more sensitive to quantization accuracy. If a unified configuration sharing strategy is adopted, it is easy to cause a decrease in accuracy. Therefore, these modules can be divided into smaller configuration sharing groups, or even configured separately for every two modules, so as to retain more quantization flexibility at the output layer and adapt to the representation requirements of high-dimensional output features.
[0068] To further enhance adaptability, this non-uniform partitioning strategy can be combined with the Neural Architecture Search (NAS) mechanism. Specifically, different partitioning strategies (such as shared group size, shared group location distribution, bound block numbers, etc.) are regarded as part of the architecture search space, and control variables are introduced during the pre-training or distillation phase for joint performance-efficiency optimization. The architecture search results can feedback which deep regions are more sensitive to accuracy, thereby guiding the dynamic adjustment of the shared group layout, and finally forming an adaptive configuration sharing mechanism that takes into account computing power distribution and accuracy stability.
[0069] In addition, when deploying deep Transformer architectures for tasks such as medical and health dialogue systems or financial transaction behavior modeling, a semantic density-driven dynamic partitioning mechanism can also be introduced: for example, according to the gradient activation distribution caused by the input tokens in the model, the group partitioning method of each segment of the network is adjusted dynamically in real time. Regions with high semantic density (such as disease diagnosis conclusions or transaction behavior breakpoints) can be configured with a smaller granularity of group partitioning to improve configuration fidelity; while information redundancy regions (such as repeated inquiries or invalid behaviors) can use large group sharing, thus reducing unnecessary waste of computing resources.
[0070] In summary, non-uniform partitioning not only provides a more flexible configuration reuse path, but also provides a structure-aware optimization path for high-performance quantization inference. It is an important enhancement strategy for the efficient deployment of complex models in resource-constrained environments.
[0071] By dividing the network modules of the language model into multiple configuration sharing groups containing a preset number of modules, the time and video memory consumption for repeatedly calculating quantization configurations for each layer of modules during the inference phase can be significantly reduced, and the configuration scheduling logic can be simplified. This strategy effectively introduces a configuration reuse mechanism at the module structure level, realizing batch control of computing resources while ensuring quantization flexibility. By sharing the precision configuration within the group, a large number of redundant configuration loading and switching operations are avoided, resulting in an overall improvement in the execution efficiency of long text inference tasks.
[0072] S60, sharing the unified quantization configuration of the processing block corresponding to the first network module within each configuration sharing group with other network modules within the same configuration sharing group; In this embodiment, to improve the execution efficiency and accuracy consistency of block-level quantization inference, the network modules are divided into multiple configuration sharing groups, and the unified quantization configuration of the processing block associated with the first network module is reused within each group. The essence of this operation is to construct a cross-module parameter sharing mechanism, avoiding generating or loading quantization configurations separately for each network module during the inference process, thereby reducing redundant configuration overhead and improving the overall operation efficiency and memory usage.
[0073] In a specific implementation, first, the network modules of the language model need to be divided into multiple configuration sharing groups according to the processing order. Each sharing group contains a preset number of network modules. For example, each group contains 4 consecutive Transformer sub-layers, or they are equally spaced according to the model depth. Subsequently, the first network module is selected from each sharing group, and the unified quantization configuration of the processing block associated with it is used as the benchmark configuration for the group. This processing block may come from the actual inference data in the model input stage, or it may be generated by simulated data or task prior data to ensure that the selected configuration is representative.
[0074] To achieve efficient sharing, the unified quantization configuration needs to be written into a shared memory area, and the same memory address is mapped for all network modules within the same configuration sharing group, thus avoiding repeated configuration loading and conversion operations. The shared quantization configuration usually includes quantization bit width (such as INT8 or INT4), scale factor, zero point, and calculation precision format identifier (such as whether it is mixed-precision floating-point calculation, etc.). For configuration parameters generated by strategies such as weighted average, historical accumulation, or semantic importance aggregation, they can also be directly written into the shared area, so that the sharing mechanism does not affect the flexibility of the original configuration generation strategy.
[0075] This configuration sharing mechanism not only improves the running efficiency of the inference stage but also enhances the local consistency of precision. Especially in the structural design where there is semantic recursive enhancement between the deep modules of the model, using consistent quantization configurations can reduce error accumulation and quantization jitter effects. To ensure effectiveness in dynamic task scenarios, a lightweight synchronization mechanism can be further introduced into the sharing group. For example, if it is detected that the configuration of the processing block bound to the first network module is updated, the system can automatically trigger the shared memory refresh operation through the configuration mapping table to ensure that the subsequent modules read the latest configuration.
[0076] In the Transformer architecture, assume there are 48 network modules, which are divided into 12 configuration sharing groups, with 4 modules in each group. The first module in each group is bound to a processing block with a representative importance score distribution (such as a block extracted from high-frequency word segments or key areas of the conversation). The unified quantization configuration is generated according to the token precision allocation of this processing block. The generated configuration is written into the shared memory area, and the other modules in each sharing group access this configuration through the same memory pointer during the inference process to complete the unified quantization operation of weights and activation values. In a multi-card deployment environment, the quantization configuration of the group can also be synchronously stored in the shared memory of each card, reducing communication latency through pointer or IPC mapping. If it is detected that the structure of the processing block changes drastically or the input switches across scenarios, the system can dynamically update the bound processing block and configuration, and remap the shared memory address to ensure the adaptability of configuration sharing.
[0077] By constructing a configuration sharing group based on processing blocks and reusing the quantization configuration associated with the first network module within the group, the generation and switching frequency of quantization parameters during model inference can be significantly reduced, while reducing the repetitive calculation overhead while ensuring semantic fidelity. It strengthens the computational consistency within the model and improves the execution stability and scalability of large model inference under multi-scenario deployment.
[0078] S70, according to the unified quantization configuration corresponding to each processing block, perform block-level batch quantization on all processing blocks and complete model inference to generate model inference results.
[0079] In this embodiment, in order to improve the computational efficiency and video memory utilization rate of the large language model in long text inference tasks, it is necessary to perform block-level batch quantization calculations based on the unified quantization configuration corresponding to each processing block and complete the corresponding model inference process. This operation integrates the capabilities of configuration-driven quantization calculation and processing block parallel scheduling in the technical path, ensuring that each processing block completes refined inference processing according to its importance and precision requirements, and finally splicing to generate a complete inference result.
[0080] The unified quantization configuration of each processing block comes from the previous importance score analysis and precision statistics operations, and usually includes precision format identifiers (such as FP16, INT8, INT4), quantization information of weight parameters (such as scale factor scale_w and zero point zero_point_w), and quantization information of activation values (scale_a and zero_point_a). The execution of quantization calculations usually depends on the selection of efficient computational kernel functions. After the precision configuration of the processing block takes effect, the corresponding execution path needs to be allocated according to the precision format. For example, floating-point calculations use FP16 kernel functions implemented based on CUDA or ROCm, and the fixed-point quantization path calls INT kernel functions while loading relevant quantization parameters.
[0081] To improve the operating efficiency, all processing blocks are mapped to the parallel processing units (SM, Streaming Multiprocessor) on the graphics processing unit (GPU, Graphics Processing Unit) or tensor processing unit (TPU, Tensor Processing Unit), and task scheduling is performed according to the starting index or processing order of the processing blocks. During the scheduling process, each processing block loads its own bound unified quantization configuration and performs corresponding matrix multiplication, normalization, activation function, and residual connection operations. During the process, the activation value quantization parameters are used to dynamically adjust the input range, and the weight quantization parameters are used to restore the model weights from fixed-point representation to computable quantization expressions.
[0082] After the inference calculation is completed, it is necessary to perform precision verification and invalid position clearing operations on the intermediate results of each processing block. This step is usually completed based on the recorded invalid token position index table, setting the intermediate tensor results at the positions corresponding to the invalid tokens to zero or deleting them to avoid interference with the final output. All valid calculation results are then reconstructed by splicing according to the starting position index of the processing block into a complete model output sequence.
[0083] Finally, perform standard post-processing procedures on the spliced output sequence, including dimension alignment (such as padding tensors of different blocks to a unified shape), normalization (such as LayerNorm normalizing the activation output), and decoding (such as generating natural language text through GreedyDecoding or Beam Search), and finally output the inference result text or structured vector that can be directly used by the user.
[0084] For example, in a text generation task, the input text is split into a sequence of processing blocks with 512 tokens in each block. After analysis in steps 4 and 5, each processing block obtains a corresponding unified quantization configuration. For example, processing block A is in high-precision format (FP16), block B is in medium precision (INT8), and block C is in low precision (INT4). During the inference execution phase, after the GPU loads the unified quantization configuration of each processing block, it allocates processing block A to the high-precision execution channel, and blocks B and C to the fixed-point execution channel, and performs matrix multiplication and non-linear transformation calculations through different kernel functions. After the calculation of each processing block is completed, the system clears the calculation results corresponding to the padding tokens according to the token position index recorded during preprocessing. For example, if block C contains 12 padding tokens, the outputs of its last 12 positions will be set to zero. Then, the valid outputs of blocks A, B, and C are sorted and spliced according to the starting index, and post-processing (such as normalization and decoding) is performed to obtain the complete text generation result. In the scenario of batch inference with multiple inputs, the micro-batch scheduling mechanism can also be combined to schedule the processing blocks of multiple input samples on the same GPU, and share the weight cache of some high-frequency structures when loading the shared configuration to further improve the throughput.
[0085] Example illustration: In the field of healthcare business, an intelligent assisted diagnosis platform hopes to deploy a large language model for structured analysis and reasoning of extremely long electronic medical record texts. The goal is to extract the main diagnosis, the evolution process of key symptoms, and the conclusion of etiological inference from tens of thousands of words of hospitalization records. In this task, the length of the original input text exceeds the standard window limit of the model, so the processing block method needs to be adopted for segmented reasoning. First, the entire medical record text is divided into multiple processing blocks according to a preset length (for example, every 512 tokens). The system marks the first processing block as a high-precision processing area, which usually contains high-semantic-density parts such as the patient's basic information, chief complaint, and current medical history, and directly determines the credibility of subsequent diagnostic reasoning. Therefore, the quantization operation of this processing block is disabled, and only high-precision operators are used for floating-point calculation. For the remaining processing blocks, the platform automatically generates the self-attention matrix corresponding to each processing block before the model execution, and sums each column in the matrix to obtain the degree of attention of each token in the entire context, that is, the importance score. On this basis, the system compares the score with two dynamic thresholds, configures the tokens with higher importance scores as high-precision format, which generally corresponds to the core disease course nodes or etiological information, configures the tokens with medium scores as medium-precision format for processing descriptive paragraphs or disease condition observation records, and configures the tokens with lower scores such as format information and table data as low-precision format. After each processing block completes the token-level precision annotation, the system counts the number of tokens in different precision formats and selects the precision level with the largest number as the unified quantization configuration for this processing block. For example, in a certain processing block, the medium-precision format tokens account for the highest proportion, then the block is configured as a medium-precision calculation path as a whole. This strategy avoids frequent precision switching and improves execution efficiency. Then, the system divides the inference module of the model into several configuration sharing groups in sequence, and each group contains several consecutive network modules. The block processed by the first module in each group is bound and configured with precision parameters, and the remaining modules share this configuration, avoiding the repeated execution of the precision selection logic for each layer. Finally, all processing blocks call the quantization calculation kernel functions of the corresponding precision to execute inference in parallel according to their unified quantization configurations. The high-precision blocks use the FP calculation kernel function, and the medium- and low-precision blocks call the fixed-point kernel function, running concurrently in the GPU or TPU parallel units. The positions corresponding to the padding tokens will be automatically skipped during the inference to ensure the validity of the output results. After the inference is completed, the system concatenates the valid results of all processing blocks into a whole inference result text according to the original token position index, and performs unified dimension alignment and normalization, and finally decodes to generate the diagnostic conclusion, the key symptom timeline, and the list of suspected etiologies. The entire inference process ensures the precision and stability of the diagnostic core information while controlling the video memory occupancy, and realizes the deployment ability of high-performance long-text medical understanding tasks.
[0086] In the field of fintech business, a credit assessment platform for small and medium-sized enterprises needs to use large language model inference to judge the credit risk of enterprises based on long text data such as complete business information submitted by enterprises, historical credit contracts, public annual reports, and customer feedback records. Since these documents often have loose structures and complex content, with the total number of tokens easily reaching tens of thousands, directly inputting them into the model will face serious problems such as out-of-memory errors and long response delays. The system first divides all text information into multiple processing blocks according to a preset length (such as every 1024 tokens), ensuring that the length of each block is controllable and the computing resources are acceptable. The first processing block usually contains high-weight fields such as enterprise entity information, registered capital, core business, and operating income in the past three years, which are the core factors affecting the credit score. Therefore, the platform strategy is to fix this block to use a high-precision calculation path and disable quantization processing to retain the expression details of key financial features and legal statements to the greatest extent. For the remaining processing blocks, the platform introduces a multi-head self-attention mechanism to analyze the degree of attention received by each token in the context. After constructing the attention matrix, it extracts the sum of the global attention weights received by each token in each column as an importance indicator for the token. Subsequently, the system assigns different precision levels to each token based on two dynamic thresholds, ensuring that high-risk factors such as tax anomalies, contract defaults, and bank credit records are processed in high-precision format, while relatively less important information such as financial statement annotations and customer descriptions is processed in medium or low precision. Then, the system counts the number of tokens of various precision levels in each processing block and selects the precision level with the highest frequency as the unified quantization configuration for that block, thereby simplifying the execution path and reducing the scheduling overhead. If there is a boundary situation where the numbers of high, medium, and low precision are similar, a preference is given to high precision to avoid weakening the key semantics in credit judgment. The platform divides the entire large language model into multiple configuration sharing groups at different levels. Each group contains a fixed number of network modules, and the quantization configuration used for the block processed by the first module in the group is used as the shared configuration, which is recorded in the configuration mapping table that binds the group identifier and the block identifier, ensuring that all modules use the same quantization path when processing the tasks of this group. During the execution phase, the system loads the corresponding unified quantization configuration parameters for each processing block, including the scaling ratio of weights, the calibration range of activation values, and the precision format identifier of the current block. Different blocks are scheduled to the GPU parallel units that support mixed precision. Among them, high-precision blocks call floating-point processing units, and medium and low-precision blocks use INT4 / INT8 instruction sets to perform inference. At the same time, the intermediate results corresponding to the positions of invalid tokens are removed to avoid introducing redundant interference. Finally, the platform sorts and splices all valid inference results according to the original token index to form a complete output sequence, and through the post-processing module, performs structure alignment, risk label normalization, and multi-label score decoding to generate the credit rating, risk exposure range, and warning dimension recommendation list of the enterprise in the current financial scenario, assisting credit review personnel or automated credit engines to complete the decision-making.
[0087] By performing block-level batch quantization calculations based on a unified quantization configuration, not only an accuracy-aware heterogeneous execution strategy is achieved, but also the parallel computing power of multi-core processors is fully unleashed. While ensuring that different token segments perform differential calculations according to importance, it effectively reduces the overhead of repeated configuration loading, dynamic quantization generation, etc. It reduces the average inference video memory occupancy and improves the processing throughput rate, and is applicable to resource-sensitive large-scale deployment scenarios in long text tasks.
[0088] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health and fintech. It discloses a method for accelerating model quantization inference, including: dividing the input text into multiple processing blocks, fixing the calculation precision format of the first processing block as a high-precision format and disabling quantization operations; for other processing blocks except the first processing block, generating a self-attention matrix through a language model, and calculating an importance score based on the sum of the values in the corresponding column of each token position in the self-attention matrix; allocating each token position to a high-precision, medium-precision, or low-precision format based on two preset thresholds; counting the number of tokens in each precision format within each processing block, and selecting the precision format with the largest number as the unified quantization configuration; dividing the network module into multiple configuration sharing groups, and sharing the unified quantization configuration of the processing blocks within the group; performing block-level batch quantization according to the unified quantization configuration of the processing blocks and completing model inference to generate an inference result. The present invention realizes precision allocation and parallel quantization inference at the block level by uniformly determining the quantization configuration of each processing block based on the token importance score and reusing this configuration within the network module group, greatly reducing the video memory overhead and configuration time overhead while ensuring the inference accuracy, and effectively improving the execution efficiency and video memory utilization rate in long text inference tasks.
[0089] In one embodiment, the above step S10 includes: S101, dividing the input text into multiple equally long processing blocks according to a preset block length; S102, disabling quantization parameter adjustment for all token positions of the first processing block; S103, fixing the processing precision format of the embedding layer, self-attention layer, and feed-forward network layer of the first processing block in the language model as a high-precision format; S104, when there are remaining tokens at the end of the input text that are less than the preset block length, forming the text segment composed of the remaining tokens as an independent processing block; S105, padding the independent processing block with invalid tokens to the preset block length, fixing the processing precision format of the padded independent processing block as a high-precision format, and disabling quantization operations on the independent processing block; S106, recording the start position index and end position index of all processing blocks.
[0090] In this embodiment, in order to reduce the video memory pressure of the large language model during long text inference, a processing block division strategy is adopted to divide the input text into multiple sub-segments according to a preset block length. Each processing block is the smallest quantization scheduling unit in model inference, with clear start and end position indexes. This division is not only used to control the locality of video memory usage, but also provides boundary support for subsequent quantization strategy generation.
[0091] The processing precision format of the first processing block is fixed to a high-precision format, usually the FP16 or FP32 floating-point format, which means that the model uses a full-precision calculation path when processing this block, and explicitly prohibits this block from participating in any form of quantization parameter adjustment, including but not limited to bit-width compression of weight parameters and activation values, dynamic scaling factor generation, etc., thus skipping the process of this block in quantization configuration generation.
[0092] The Embedding Layer, Self-Attention Layer, and Feed-Forward Network Layer together constitute the core module for token representation transformation in the language model, playing a key role in the semantic representation of tokens. Adopting a high-precision format for these core structures of the first processing block helps to stabilize the context awareness ability of the model at the initial stage of inference and improve the nested inference performance of subsequent tokens, especially in scenarios where key topic cues are included at the initial stage of long text input.
[0093] If the text length cannot be evenly divided by the preset block length, the remaining tokens at the end will be assembled into an independent processing block. To maintain the consistency of the block-level inference structure and the alignment requirements for parallel scheduling, this independent processing block will be padded with invalid tokens (Padding tokens) to make it up to the full block length and also processed using a high-precision format. The padding tokens do not participate in the effective calculation of the model output during actual inference, but are only used to keep the input dimensions consistent.
[0094] After the system completes text chunking, it needs to record the start and end position index information of each processing block, which serves as a key auxiliary identifier for subsequent calculation scheduling, output splicing, precision allocation, etc. These index information can be constructed into a bidirectional mapping table (such as a hash table) to quickly locate the processing block corresponding to the token and its precision control strategy during runtime.
[0095] The input text can be sequentially partitioned by setting the block length as a subset of the maximum sequence length supported by the model (such as 512 or 1024 tokens). The partitioning operation can be done offline during the preprocessing stage or dynamically based on streaming input at runtime. In the GPU inference framework, a position encoding mask can be used to distinguish actual tokens from padding tokens, and a high-precision kernel function can be used to process the specified block.
[0096] In the current mainstream quantization inference hardware systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), and TPUs (Tensor Processing Units) are widely used in the deployment of high-performance large language models. For these platforms, flexible precision scheduling capabilities need to be provided during actual deployment to accommodate the high-precision processing requirements of the first processing block.
[0097] In a programmable chip architecture (such as an FPGA or an AI acceleration chip), the various operations in the model execution path are generally mapped to a series of reconfigurable logic units or operator chains. To achieve high-precision execution of the first processing block, the computing precision mode of the current block can be set through a bit-width control register or an operator scheduling instruction table in the intermediate computing path. When it is recognized that the token position being processed belongs to the first processing block, the system controller can, through a combination of hardware and software, set the precision control field of the corresponding operator chain to a high-precision flag, such as enabling the FP16 or FP32 channel, and bypass the quantization scaling unit, skipping operations such as zero-point adjustment and discrete mapping, so that the data is directly transmitted to the operator input port in its original floating-point format, implementing the so-called "precision passthrough path".
[0098] The key to this mechanism lies in binding the position information of the processing block to the register configuration strategy. The purpose of dynamically switching the precision path at the block level can be achieved by maintaining a processing block - precision strategy mapping table in the scheduling unit. When processing streaming token input, this mechanism can implement runtime configuration updates through a boundary detection module.
[0099] In highly integrated heterogeneous computing platforms such as TPU, it usually contains multiple processing cores (Cores) that support different computing precisions. Among them, some cores are high-computing-power main cores (High-Precision Cores) that support floating-point computing, while other cores are optimized for dedicated fixed-point computing paths. The TPU architecture often integrates a scheduler and a resource allocator to dynamically schedule workloads among different Cores.
[0100] When processing the first processing block, the scheduler can identify it as a high-priority computing unit and directly allocate its tasks to the main core, using the floating-point computing core to execute the calculations of the Embedding, self-attention, and feed-forward network modules, thus avoiding the introduction of the quantization process; while the inference tasks of the remaining processing blocks can be divided into other low-bitwidth Cores and executed in a low-precision high-throughput mode for quantized inference. To achieve dynamic reuse of resources, the scheduling strategy of TPU can also support time slicing and task prefetching to ensure low latency and high parallelism in the scheduling between multiple processing blocks.
[0101] If deployed on a general GPU platform, the high-precision execution flag can also be passed through the CUDA kernel launch parameter to disable the quantization process in a specific kernel path. For example, launch a group of kernels without quantization kernel functions for the first processing block, or force the scheduling of the float path instead of the int path through the precision tag injection mechanism in the general model framework.
[0102] In addition, in a higher-order system, a dynamic fine-tuning logic (precision adaptation controller) can also be introduced to automatically determine whether to enable the high-precision path according to the importance score distribution or semantic weight of the first processing block, thus achieving a flexible balance between performance and precision.
[0103] In this embodiment, by dividing the input text into multiple equal-length processing blocks and disabling the quantization process for the first processing block, more original semantics and context information can be retained in the initial stage of model inference, thereby improving the stability of subsequent token importance evaluation and precision allocation. It helps to improve the initial inference accuracy of long text input without significantly increasing the overall video memory overhead, providing a more reliable baseline representation for subsequent block quantization configuration.
[0104] In one embodiment, the above step S20 includes: S201. Input each other processing block into the multi-head self-attention layer of the language model to obtain the local self-attention matrices of multiple attention heads corresponding to each other processing block; S202. Perform weighted average or arithmetic average processing on the local self-attention matrices of multiple attention heads to generate the final self-attention matrix corresponding to each other processing block; S203. Extract the column vector corresponding to each token position from the final self-attention matrix; S204. Determine the sum of the values of all elements in each column vector, and use the sum as the importance score of the token position corresponding to each column vector; S205. If there are filled invalid token positions in other processing blocks, when determining the importance score of the filled invalid token positions, set all element values of the column vector corresponding to the invalid token position to zero.
[0105] In this embodiment, before allocating precision to the processing block, it is necessary to first obtain the importance of each token in the current semantic context. To achieve this goal, the processing flow starts from the multi-head self-attention layer and constructs an attention matrix for each processing block. After each processing block is input into the language model, it will pass through multiple attention heads. Each attention head independently calculates a local self-attention matrix. Each matrix takes all tokens in the processing block as units and describes the mutual attention relationship between them. These matrices reflect the language model's ability to model the internal structure of the sentence from different attention perspectives.
[0106] To obtain a unified and comparable attention representation, it is necessary to fuse these local matrices. The fusion method can be flexibly set according to the actual task requirements. Common methods include taking the arithmetic average of the matrices of all attention heads, that is, simply calculating the mean; or by configuring different weight coefficients for each attention head to achieve weighted average. This fusion process generates the final self-attention matrix of each processing block, which serves as a comprehensive representation of the mutual dependence relationship of the tokens within the processing block.
[0107] In this final matrix, each column vector represents the degree of attention focus that a specific token receives from other tokens. Specifically, fixing the position of a certain token and extracting the column vector corresponding to that position can be regarded as "the set of degrees to which this token is concerned by all other tokens". This is an important basis for evaluating whether a token occupies an important semantic position in the current processing block. That is: Current token: The token for which the attention output is being calculated, denoted as the token at the i-th position.
[0108] Other tokens: All valid tokens in the same processing block except the i-th position, i.e., the position numbers within the block are from 0 to N-1 (where N is the length of the processing block); excluding the i-th position itself; excluding the positions of invalid tokens (such as padding tokens).
[0109] To extract this importance signal, all elements in each column vector are numerically summed to obtain an aggregated value that measures the global attention received by the token. This aggregated value is the importance score of the token, and the higher the score, the greater the impact of the token on the overall context semantic structure. In the subsequent precision allocation process, this score will be directly used to determine whether the token should be processed using higher-precision computing resources.
[0110] Considering that the processing block may contain invalid tokens padded for the fill length, these tokens do not carry semantic content and may interfere with the importance score if they participate in the calculation of the attention matrix. To prevent such interference, after generating the final self-attention matrix, the positions corresponding to all invalid tokens need to be identified, and their column vectors in the matrix are uniformly set to zero. This can effectively prevent invalid tokens from obtaining false high attention values and ensure the accuracy and robustness of the importance calculation. This operation can be implemented through a structured mask or an explicit indexing mechanism, and the specific method is adjusted according to different model architectures and deployment platforms.
[0111] In actual deployment, when the processing block is fed into the language model for inference, each layer of the self-attention mechanism of the model will output the attention matrices corresponding to a number of attention heads. These matrices are usually organized in a three-dimensional structure, where the first dimension is the number of attention heads, and the second and third dimensions respectively represent the number of tokens in the processing block, forming a three-dimensional tensor with a structure of "number of attention heads × number of tokens × number of tokens".
[0112] During the forward inference process of the model, these attention tensors can be cached at the output of each layer's self-attention module through an intermediate layer interception mechanism. To reduce the interruption overhead during inference, the caching operation should be inserted into the model's computational graph in a non-blocking manner and its lifecycle should be uniformly managed by the scheduler. After caching, different fusion strategies can be achieved by operating on the first dimension (i.e., the attention head dimension): one way is to simply average the matrices of all attention heads to obtain a two-dimensional matrix; another way is to assign different weighting coefficients to each head and perform weighted average processing, which is applicable when the differences in the semantic perception capabilities of different attention heads are evaluated in advance and strategic adjustments are made in the weight settings to enhance the expressive power of the fusion matrix.
[0113] After completing the attention head fusion, a two-dimensional matrix of token number × token number will be obtained, representing the comprehensive attention relationship among tokens within the processing block. The global attention degree of each token is determined by the sum of the elements in the corresponding column of this matrix. In specific operations, by fixing a certain column index, column vectors can be extracted from the entire matrix, and this operation can be implemented through standard matrix slicing functions in the tensor operation framework (such as PyTorch, TensorFlow, JAX, etc. all support such slicing access). Subsequently, all the values in this column vector are accumulated to obtain the attention aggregation value of the token at this position, that is, its importance score. Since vector addition is highly optimized in modern tensor calculation libraries, the calculation of the importance score can be completed with extremely low latency on the GPU / TPU, with good engineering feasibility.
[0114] Regarding the processing of padding tokens, during forward inference, the model usually generates a padding mask, which is used to mark which tokens are invalid positions generated by padding. The mask is generally a boolean vector or a 0 / 1 tensor of the same length as the processing block, where 1 represents a valid position and 0 represents a padding position. This mask can be extended to a matrix form and aligned with the column structure of the final self-attention matrix. In implementation, this mask can be broadcast as a two-dimensional matrix and applied to each column of the attention matrix, uniformly setting all elements of the column vector corresponding to the padding position to zero at the matrix level.
[0115] To achieve optimal efficiency, the application process of this mask can be embedded in the quantization preprocessing module during the model graph construction phase and attached as a post-operation after the attention matrix calculation node. Depending on the different operating platforms, the column-level zero-clearing logic can be implemented through CUDA kernel on the GPU; on the TPU, batch column masking can be completed within a single clock cycle relying on its built-in high-throughput tensor mapping operation.
[0116] In addition, to ensure that special structures that appear in the context scenario do not cause incorrect judgments, such as in some tasks, the padding tokens may be in the middle of the sequence rather than at the end (such as multi-segment concatenated input), it is necessary to record the original input structure and the padding mask generation logic in advance to ensure the accuracy of the mask. An optional method is to introduce a position label assistance mechanism, jointly using the position semantics encoding of each token and the mask to enhance the recognition stability of invalid tokens.
[0117] In this embodiment, an importance score of a token is obtained by extracting column vectors from the fused self-attention matrix and summing the total weights, thereby constructing a quantization accuracy perception mechanism starting from the language model structure. Without introducing an external scoring module, the attention information already learned by the model itself is fully utilized to internalize the basis for accuracy allocation. This avoids explicit label dependence and has good computational graph compatibility, which helps with efficient deployment on subsequent acceleration chips.
[0118] In one embodiment, step S30 above includes: S301, determining a first threshold and a second threshold according to the length of the processing block; S302, if there is an invalid token position, allocating the processing precision format of the invalid token position to a low-precision format and excluding the invalid token position when counting the number of precision formats of each processing block; S303, comparing the importance score of each valid token position with the first threshold and the second threshold; S304, allocating the processing precision format of the valid token position with an importance score greater than the first threshold to a high-precision format; S305, allocating the processing precision format of the valid token position with an importance score above the second threshold and below the first threshold to a medium-precision format; S306, allocating the processing precision format of the valid token position with an importance score less than the second threshold to a low-precision format.
[0119] In this embodiment, allocating the token positions with importance scores greater than the first threshold to a high-precision format is an important operation for differentially configuring the processing precision after evaluating the attention information of all valid tokens in each processing block. The high-precision format refers to using a higher-precision representation and calculation method during the inference process, such as using a numerical representation with a higher bit width or floating-point operations that disable quantization operations. The purpose is to retain the expressive ability of key semantic tokens and avoid negative impacts on the model prediction results caused by a decrease in precision. The medium-precision format and the low-precision format correspond to medium and low bit-width quantization configurations in resource-constrained scenarios and are used to process tokens with lower importance to save video memory and computing resources.
[0120] The length of the processing block has a dynamic impact on the setting of the threshold. Since the number of tokens in a long processing block increases and its overall information density distribution tends to be average, in order to ensure that the screening mechanism has relatively consistent discrimination ability in processing blocks of different lengths, the values of the first threshold and the second threshold need to be correspondingly lowered to expand the determination range of high-precision tokens. Dynamic adaption can be achieved by setting an initial standard threshold and linearly or non-linearly scaling the two thresholds in combination with the ratio of the current block length to the standard length.
[0121] In the scenario where there are invalid token positions in the processing block, special judgment needs to be added to the processing precision allocation. Since invalid tokens do not participate in semantic modeling, there is no semantic contribution in their attention information itself, and their corresponding importance scores can be considered as structural null values or forced to zero. Uniformly allocating them to the low-precision format can save computing resources and actively exclude them when counting the precision ratio later to avoid interfering with the unified quantization configuration decision of the current processing block.
[0122] The importance score of each valid token will be compared with the two thresholds in turn. When its value is higher than the first threshold, it means that it plays a significant semantic connection role in the global attention aggregation and needs to be allocated the high-precision format; when its score is between the first and the second thresholds, it is considered to be in the sub-important information interval and can be processed using the medium-precision format; when it is lower than the second threshold, it indicates that the token has a weak effect on context construction and can be processed using the low-precision format for inference calculation. This multi-level precision configuration scheme ensures the balance of semantic expression, performance control, and computational efficiency.
[0123] If there are no invalid token positions in the processing block, all token positions are valid token positions.
[0124] In a specific implementation, when initializing the model for running, by reading the length of the processing block, a preset dynamic threshold calculation function can be called to generate the first threshold and the second threshold for the current block. This function can be implemented by means of table lookup or formula calculation, and the specific implementation can support mechanisms such as interval interpolation, exponential decay, and minimum threshold protection. Subsequently, all token positions of the current processing block are traversed in sequence. During the traversal process, the system calls the padding mask or the token validity label to determine whether the current token is an invalid token. If it is an invalid token, its processing precision format is directly marked as low precision, and the importance score comparison logic is skipped, and the judgment process of the next token is entered. For valid tokens, the importance score of the token is read from the cache or the current attention module. This score is then compared with the two thresholds corresponding to the current processing block. If the score is greater than the first threshold, it is marked as high precision in the processing block precision allocation vector; if the score is greater than the second threshold and less than or equal to the first threshold, it is marked as medium precision; if the score is less than or equal to the second threshold, it is marked as low precision. After the processing is completed, the precision labels of all valid tokens are saved in the precision control vector inside the processing block and are used for subsequent unified quantization configuration decisions and kernel function selection. If acceleration of processing is required, the importance score comparison logic can be vectorized, and the judgment of the allocation strategy can be completed in parallel on the GPU to improve the execution efficiency.
[0125] Through the above steps, in this embodiment, without significantly increasing the model complexity, the calculation precision of tokens can be differentially configured according to the actual semantic importance of the tokens, thereby greatly reducing the overall video memory consumption and calculation load while ensuring the inference precision of the model. The dynamic threshold mechanism further improves the adaptability of the algorithm when processing input sequences of different lengths, and avoids the problem of uneven precision allocation caused by fixed thresholds. The special processing of invalid tokens ensures the accuracy of resource allocation and simplifies the subsequent statistical logic.
[0126] In one embodiment, the above step S40 includes: S401, traversing all token positions of the current processing block to identify the token positions assigned to the high-precision format, medium-precision format, and low-precision format; S402, during the statistical process, if there are invalid token positions, all invalid token positions are excluded, and only the precision format allocation results of valid token positions are statistically counted; S403, respectively counting the counts of the high-precision format, medium-precision format, and low-precision format assigned to the valid token positions to generate a high-precision count, a medium-precision count, and a low-precision count; S404. Compare the values of the high-precision count, medium-precision count, and low-precision count; S405. Use the precision format corresponding to the count with the largest value as the unified quantization configuration for the current processing block; S406. If there are multiple precision formats with the same count and it is the maximum value, select the highest precision format among the multiple precision formats as the unified quantization configuration for the current processing block.
[0127] In this embodiment, counting the number of token positions assigned different precision formats within each processing block is a key analysis operation performed after token precision allocation to determine the overall quantization strategy for the current processing block. This process first requires traversing all token positions in the current processing block to identify the precision label corresponding to each position, which may be in high-precision, medium-precision, or low-precision format.
[0128] During the traversal, if the processing block contains invalid tokens, such as padding tokens filled at the end, they need to be excluded from the statistics. This is because invalid tokens do not participate in actual inference calculations and should not affect the overall precision distribution. Only counting the precision labels of valid tokens can ensure the rationality and accuracy of subsequent unified quantization configuration decisions.
[0129] Counting for each precision label is the core of this step. The system will generate three count results respectively: high-precision count, medium-precision count, and low-precision count. Each count result corresponds to the number of positions where valid tokens in the processing block are assigned this precision format, and is used to depict the precision distribution trend within the block.
[0130] Subsequently, the system compares the sizes of the three count results, finds the precision format with the largest number, and uses it as the unified quantization configuration for the current processing block, that is, the current processing block will perform unified quantization and compute kernel function calls in this format subsequently. The design purpose of the unified configuration mechanism is to avoid frequent switching of precision calculation paths within the same processing block, thereby reducing the context switching cost and improving the inference throughput efficiency.
[0131] If multiple precision count results are the same and are the maximum value, such as the high-precision and medium-precision counts being exactly the same, the system should use the precision priority principle and select the one with a higher precision level among the tied maximum values as the final configuration. This strategy reflects the tendency to protect semantic quality, that is, to preferentially retain higher computational precision within the scope of available resources, thereby reducing the risk of token misrecognition or inference deviation.
[0132] For example, during the model execution process, after the token precision allocation for each processing block is completed, the system will start the precision statistics module. First, it traverses all the token index positions in this processing block through loop or vectorization logic. The precision label corresponding to each token is stored in the precision marking vector of the previous step. Each element of the vector is the precision identifier of the token. For example, 0 represents low precision, 1 represents medium precision, and 2 represents high precision.
[0133] During the traversal, the system checks whether each token is a valid token. The judgment basis can be the padding mask or the token validity vector. If it is an invalid token, its corresponding position will be skipped during the precision statistics process and will not participate in any precision counting.
[0134] For valid tokens, they are added to the corresponding precision counters according to their precision identifiers. The system also maintains three independent integer variables to accumulate the numbers of tokens with high, medium, and low precision respectively. After the traversal is completed, the system compares the values of the three counters to find the precision format corresponding to the maximum value.
[0135] If the precision counts of two or three types are all the maximum values, for example, both high precision and medium precision are 100, then the system performs a priority judgment. It can compare through the precision level values and preferentially select the precision format with a higher level. The system can also introduce additional factors in this logic, such as the historical block configuration inertia, the context semantic density trend, etc., to enhance the stability and adaptability of the precision configuration.
[0136] The finally determined precision format will be written into the processing block control structure as the unified quantization configuration of the current processing block and passed to the subsequent quantization kernel function selection logic. This configuration can be applied in batch form on the GPU to improve the efficiency of the quantization preparation stage.
[0137] In this embodiment, through comprehensive statistics and priority analysis of the precision labels inside each processing block, it is possible to allocate the most suitable unified quantization configuration for each block in a data-driven manner according to its semantic density and precision requirements, significantly reducing the computational path complexity and precision switching overhead. While retaining the high-precision processing ability of key tokens, it also improves the inference efficiency and video memory compression through the block-level merging strategy. The exclusion process of invalid tokens ensures that the statistical results are not interfered by redundant tokens, further improving the configuration accuracy and execution stability of the system. The introduction of the precision priority strategy enables the system to still guarantee the processing quality of high-semantic contribution tokens in the face of fuzzy precision distribution scenarios, contributing to the stability and reliability of the model output.
[0138] In one embodiment, the above step S60 includes: S601, Assign a unique block identifier to each processing block and a unique group identifier to each configuration sharing group; S602, Establish a binding relationship between the unique group identifier and the unique block identifier in the configuration mapping table; S603, Based on the binding relationship between the unique group identifier and the unique block identifier, associate the first network module of each configuration sharing group with the unified quantization configuration parameters of the corresponding processing block; S604, Write the unified quantization configuration parameters into the shared memory area of each configuration sharing group, and assign the same memory address mapping to all network modules of each configuration sharing group; S605, When other network modules within the same configuration sharing group perform quantization processing, read the unified quantization configuration parameters from the shared memory area through the memory address mapping.
[0139] In this embodiment, in order to achieve compression optimization of video memory resources and calculation paths during the inference process of the large language model (LLM), a configuration sharing mechanism is introduced to reuse the unified quantization configuration among multiple network modules within the same group. The key to this mechanism lies in the structured organization of the mapping relationship between processing blocks and network modules, and the reuse of parameter access paths using shared memory.
[0140] First, during the processing block division stage, the system generates a globally unique block identifier for each processing block, which is used to identify the group of token sequences and their quantization configurations corresponding to the block. In different implementation platforms, this identifier can be an integer index, a hash result, or a combined encoding that combines the processing block position and length, with stability and traceability, facilitating quick positioning of the quantization configuration ownership in subsequent steps.
[0141] Meanwhile, the system also divides the network modules into several configuration sharing groups according to their processing order in the model inference graph. A configuration sharing group is a logical structure that contains multiple adjacent or semantically related network modules, usually set to a fixed number (such as each group contains 4, 8, or 16 modules). Each configuration sharing group is also assigned a unique group identifier, and its generation mechanism should maintain the same uniqueness and mapability as the processing block identifier.
[0142] Based on the establishment of these two types of identifiers, the system forms a binding relationship between the group identifier and the block identifier by constructing a configuration mapping table. This mapping table can adopt the key-value pair form, with the group identifier as the key and the corresponding block identifier as the value. This structure allows the model to quickly find the processing block and its quantization configuration bound to the current configuration sharing group during runtime, thereby determining the shared access path.
[0143] Furthermore, to enable the reuse of quantization configurations among multiple modules within a group, the system writes the unified quantization configuration parameters of the processing block into a shared memory area after binding. This shared area can be a shared cache (shared memory) in the GPU memory at the implementation level, a tensor accelerator register area supporting concurrent access, or a high-bandwidth cache block allocated by middleware in a system supporting heterogeneous architectures.
[0144] The unified quantization configuration parameters typically include weight quantization parameters, activation value quantization parameters, processing precision format information, bit-width encoding strategies, quantization range information, and auxiliary biases. These parameters are stored in a structured format, such as a key-value dictionary, a compressed tensor, or a configuration structure, and are exposed to all network modules within the configuration sharing group through the shared memory address.
[0145] In terms of address mapping, the system points the quantization configuration access paths of all modules within the configuration sharing group to a unified shared memory address. This mapping process can be statically completed during model loading or dynamically maintained during the inference process, depending on the time coupling relationship between the generation of the processing block and the invocation of modules within the group. It should be noted that to ensure the performance and consistency of concurrent reads by multiple modules, the shared memory area should be ensured to have thread safety and non-blocking characteristics, which can be achieved through lock mechanisms, read-write caches, or communication channels.
[0146] Once the mapping is completed, when the remaining network modules within the configuration sharing group perform quantization processing, they will no longer independently generate or load quantization configurations. Instead, they directly read the required parameter data from the shared area by looking up the mapped address. This approach avoids repeated calculations, memory copying, and configuration switching, greatly compressing the execution graph width of the inference path and improving the overall computational density and throughput.
[0147] Furthermore, if this mechanism is combined with quantization-aware training (QAT) or post-training quantization (PTQ) strategies, configuration sharing not only reduces the memory consumption of the hardware layer but also improves semantic consistency, enabling modules within the same group to collaboratively express local semantics under a consistent quantization configuration, effectively avoiding representation jitter or semantic drift caused by differences in quantization strategies.
[0148] In addition, the configuration sharing mechanism has good scalability and hardware compatibility. It can be implemented through a unified layer parameter injection strategy in frameworks supporting graph execution (such as TensorRT, ONNX Runtime), and through the cooperation of a memory address binding table and an interrupt synchronization mechanism at the low-level hardware layer (such as TPU, NPU) to achieve dynamic sharing.
[0149] This embodiment uses a configuration sharing mechanism to achieve efficient reuse of processing block quantization configurations within a configuration sharing group, significantly reducing video memory consumption and computational burden caused by repeated configuration loading. Compared with the traditional method of loading configurations for each module independently, the use of a unified memory address mapping can reduce configuration access latency to a constant level and reduce the overall peak memory usage of the system. The unified quantization strategy within the group also improves the consistency of multi-module collaborative execution, effectively reduces quantization jitter in the reasoning path, and enhances model output stability. In the scenario of multi-block concurrent execution, this mechanism can effectively avoid semantic drift or output fault problems caused by inconsistent quantization strategies.
[0150] In one embodiment, the above step S70 includes: S701, loading corresponding unified quantization configuration parameters for each processing block, wherein the unified quantization configuration parameters include a precision format identifier; S702, configuring independent processing kernel functions for different processing blocks according to the precision format identifier; S703, in a parallel processing unit of a graphics processor or a tensor processor, allocating processing resources according to the index order of the processing blocks, and concurrently executing quantization processing of all processing blocks; S704, performing token position verification on the processing result of each processing block to eliminate the intermediate results corresponding to invalid token positions; S705, sorting the valid intermediate results of all processing blocks according to the starting position index and splicing them into a complete model output sequence; S706, performing post-processing operations on the spliced output sequence to generate a final model inference result.
[0151] In this embodiment, loading unified quantization configuration parameters for each processing block is a preparation stage before model execution. The unified quantization configuration parameters include three key contents. The first is the weight quantization parameter, which is a parameter set used to map the model weight from floating-point representation to quantized representation, usually including scaling factors and zero points. These parameters are widely used in static quantization or dynamic quantization processes; the second is the activation value quantization parameter, which controls the compression accuracy of the input activation value and is used to maintain the stability of the input feature distribution during inference; the third is the precision format identifier, which is a mark that guides the processing block to select the corresponding calculation path, usually an enumeration type or a control register flag, which is used for subsequent dynamic kernel function scheduling. The specific implementation of loading these parameters can be achieved through configuration table lookup, memory pre-reading, or configuration cache to reduce memory access latency in the inference phase.
[0152] Configuring independent processing kernel functions according to precision format identifiers is a key step in implementing block-level heterogeneous computing. The processing kernel function is an execution path customized for different precision formats. High-precision formats usually correspond to floating-point kernel functions of FP16 or FP32, retaining higher numerical precision to process semantically dense or position-critical tokens; medium and low-precision formats correspond to fixed-point quantization kernel functions such as INT8 and INT4, using mechanisms such as tensor quantization and weight sharing to accelerate matrix calculations and memory transfers. The scheduling strategy of the computing kernel function can be implemented through a function pointer mapping table or a kernel selection mechanism supported by hardware, enabling the processing block execution path to be allocated on demand and improving the overall throughput.
[0153] In a graphics processing unit (GPU) or a tensor processing unit (TPU), allocating parallel processing resources for all processing blocks is the physical guarantee for implementing block-level batch quantization. Allocating resources based on the index order of the processing blocks can ensure that the output order of the processing results is consistent with the input order, while avoiding resource competition between threads. In the GPU platform, the distribution of the processing kernel function can be controlled through the CUDA block and stream mechanisms, while in the TPU, each processing block execution can be mapped through the matrix multiplication unit (MXU). In this process, the weight quantization parameters and the activation value quantization parameters are synchronously loaded into the corresponding execution context, respectively acting on the quantization mapping of the weight tensor in the model and the dynamic calibration operation of the input tensor.
[0154] The intermediate results after processing need to be checked for the position of tokens within the block. Since some processing blocks contain invalid tokens filled with padding, the calculation results at these positions should be filtered to prevent affecting the subsequent model output. In this step, the token mask recorded in the previous step can be used to perform a screening operation on the intermediate results by position, removing the results corresponding to the invalid tokens and maintaining the position consistency during the splicing process.
[0155] When splicing the valid intermediate results of all processing blocks, they need to be sorted and concatenated according to the starting position index of the original input to ensure that the semantic information is restored in the correct context order. This operation not only requires data reorganization based on the token index but also requires the output formats of different processing blocks to be consistent, including tensor dimensions, sequence alignment methods, etc., to avoid dimension misalignment or information overlap during the splicing process.
[0156] Finally, the concatenated sequence is input to the post-processing module to perform dimension alignment, normalization, and decoding operations. Dimension alignment is used to standardize the tensor shapes of each segment of the output, normalization is used to compress the numerical dynamic range and improve the decoding accuracy, and decoding generates the final inference results such as text, labels, or feature values according to the model task type. The entire post-processing flow can be embedded in the latter part of the inference graph as a unified output layer to ensure efficient execution.
[0157] In some implementations, the weight quantization parameter and the activation value quantization parameter can be obtained through external quantization training, saved in an offline configuration file, and pre-loaded before the start of inference. Another approach is dynamic quantization at runtime, where the scaling ratio is calculated in real-time based on the maximum and minimum value range of the current input activation values. For the assignment of the precision format identifier, it can be represented as an integer encoding during the quantization configuration generation step. For example, 0 represents low precision, 1 represents medium precision, and 2 represents high precision, for the processing kernel function selector to parse.
[0158] The allocation of the processing kernel function can adopt an adaptive kernel call strategy supported by the hardware. For example, in the NVIDIA TensorRT framework, different GEMM implementation modules are automatically matched according to the precision identifier, while on the TPU platform, the precision configuration can be compiled into the corresponding hardware instruction sequence through the XLA compiler. The allocation of processing resources can also be optimized according to the hardware architecture. Some platforms support preferentially scheduling high-precision blocks to the core resource area with stronger computing performance to ensure the accuracy of model inference.
[0159] In actual deployment, to reduce the bandwidth overhead during the stitching process, the valid results of each processing block can be immediately written to the pre-allocated global output buffer after the output of each processing block, arranged according to the token index position, to avoid subsequent sorting operations. During the dimension alignment process, for the output tensor differences caused by different precision formats, a unified mapping strategy or a padding filling mechanism can be introduced to ensure the consistency of the data structure.
[0160] In this embodiment, by loading the corresponding unified quantization configuration for each processing block and dynamically selecting the processing kernel function accordingly, the collaborative optimization of computing precision and resource allocation is achieved. By introducing the precision format identifier to control the kernel function path, while ensuring the processing precision of high-importance regions, the computing resources of low-importance regions are maximally compressed. The parallel processing unit executing by blocks further improves the execution efficiency, and the elimination of invalid token positions ensures the precision and consistency of the stitching result. After the final output undergoes unified post-processing operations, it can not only maintain the context semantic coherence but also effectively reduce the overall inference latency and video memory consumption.
[0161] In one embodiment, a model quantization inference acceleration device is provided, and the model quantization inference acceleration device corresponds one-to-one to the model quantization inference acceleration method in the above embodiment. Refer to Figure 3 , Figure 3 which is a schematic diagram of the functional modules of a preferred embodiment of the model quantization inference acceleration device of the present invention. An input text preprocessing module 10, a self-attention analysis module 20, a precision allocation module 30, a quantization configuration decision module 40, a network module grouping control module 50, a configuration sharing management module 60, and an inference execution module 70. The detailed description of each functional module is as follows: The input text preprocessing module 10 is used to divide the input text into multiple processing blocks, fix the processing precision format of the first processing block to a high-precision format, and disable the quantization processing of the first processing block; The self-attention analysis module 20 is used to generate a self-attention matrix for each of the other processing blocks except the first processing block among the multiple processing blocks through a language model, determine the sum of all element values in the column corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score of each token position; The precision allocation module 30 is used to allocate the token positions with importance scores greater than the first threshold to the high-precision format, allocate the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and allocate the token positions with importance scores less than the second threshold to the low-precision format; The quantization configuration decision module 40 is used to count the number of token positions allocated to the high-precision format, medium-precision format, and low-precision format in each processing block, and select the precision format with the largest number as the unified quantization configuration of the corresponding processing block; The network module grouping control module 50 is used to divide the network modules of the language model into multiple configuration sharing groups, and each configuration sharing group includes at least two network modules; The configuration sharing management module 60 is used to share the unified quantization configuration of the processing block corresponding to the first network module within each configuration sharing group with other network modules within the same configuration sharing group; The inference execution module 70 is used to perform block-level batch quantization on all processing blocks according to the unified quantization configuration corresponding to each processing block and complete model inference, and generate model inference results.
[0162] In an embodiment, the input text preprocessing module 10 is specifically used for: Dividing the input text into multiple equally long processing blocks according to a preset block length; Disabling the adjustment of quantization parameters for all token positions of the first processing block; Fixing the processing precision formats of the embedding layer, self-attention layer, and feed-forward network layer of the first processing block in the language model to the high-precision format; When there are remaining tokens at the end of the input text that are less than the preset block length, forming the text segment composed of the remaining tokens into an independent processing block; Padding the independent processing block with invalid tokens to the preset block length, fixing the processing precision format of the padded independent processing block to the high-precision format, and disabling the quantization operation on the independent processing block; Record the start position index and end position index of all processing blocks.
[0163] In one embodiment, the self-attention analysis module 20 is specifically configured to: Input each other processing block into the multi-head self-attention layer of the language model to obtain the local self-attention matrix of multiple attention heads corresponding to each other processing block; Perform weighted average or arithmetic average processing on the local self-attention matrices of multiple attention heads to generate the final self-attention matrix corresponding to each other processing block; Extract the column vector corresponding to each token position from the final self-attention matrix; Determine the sum of the numerical values of all elements in each column vector, and use the sum as the importance score of the token position corresponding to each column vector; If there are filled invalid token positions in other processing blocks, when determining the importance score of the filled invalid token positions, set all the numerical values of the elements in the column vector corresponding to the invalid token positions to zero.
[0164] In one embodiment, the precision allocation module 30 is specifically configured to: Determine the first threshold and the second threshold according to the length of the processing block; If there are invalid token positions, allocate the processing precision format of the invalid token positions to the low-precision format, and exclude the invalid token positions when counting the number of precision formats of each processing block; Compare the importance score of each valid token position with the first threshold and the second threshold; Allocate the processing precision format of the valid token positions with importance scores greater than the first threshold to the high-precision format; Allocate the processing precision format of the valid token positions with importance scores above the second threshold and below the first threshold to the medium-precision format; Allocate the processing precision format of the valid token positions with importance scores less than the second threshold to the low-precision format.
[0165] In one embodiment, the quantization configuration decision module 40 is specifically configured to: Traverse all token positions of the current processing block to identify the token positions assigned to the high-precision format, medium-precision format, and low-precision format; During the statistical process, if there are invalid token positions, exclude all invalid token positions and only count the precision format allocation results of the valid token positions; Count the numbers of the valid token positions that are assigned to the high-precision format, medium-precision format, and low-precision format respectively, and generate a high-precision count, a medium-precision count, and a low-precision count; Compare the numerical sizes of the high-precision count, the medium-precision count, and the low-precision count; Use the precision format corresponding to the count with the largest value as the unified quantization configuration of the current processing block; If there are multiple counts of precision formats that are the same and are the maximum values, select the highest-precision format among the multiple precision formats as the unified quantization configuration of the current processing block.
[0166] In one embodiment, configure the shared management module 60, which is specifically used for: Assign a unique block identifier to each processing block, and assign a unique group identifier to each configuration sharing group; Establish a binding relationship between the unique group identifier and the unique block identifier in the configuration mapping table; Based on the binding relationship between the unique group identifier and the unique block identifier, associate the first network module of each configuration sharing group with the unified quantization configuration parameters of the corresponding processing block; Write the unified quantization configuration parameters into the shared memory area of each configuration sharing group, and assign the same memory address mapping to all network modules of each configuration sharing group; When other network modules within the same configuration sharing group perform quantization processing, read the unified quantization configuration parameters from the shared memory area through the memory address mapping.
[0167] In one embodiment, the inference execution module 70 is specifically used for: Load the corresponding unified quantization configuration parameters for each processing block, and the unified quantization configuration parameters include a precision format identifier; Configure independent processing kernel functions for different processing blocks according to the precision format identifier; In the parallel processing unit of a graphics processor or a tensor processor, allocate processing resources according to the index order of the processing blocks, and concurrently execute the quantization processing of all processing blocks; Perform in-block token position verification on the processing results of each processing block to eliminate intermediate results corresponding to invalid token positions; Sort the valid intermediate results of all processing blocks according to the start position index, and splice them into a complete model output sequence; Perform post-processing operations on the spliced output sequence to generate the final model inference result.
[0168] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a model quantization inference acceleration method.
[0169] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a model quantization inference acceleration method In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: Divide the input text into multiple processing blocks, fix the processing precision format of the first processing block to a high-precision format, and disable the quantization processing of the first processing block; For the other processing blocks in the multiple processing blocks except the first processing block, generate the self-attention matrix of each other processing block through a language model, determine the sum of all element values in the column corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score of each token position; Assign the token positions with importance scores greater than the first threshold to the high-precision format, assign the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and assign the token positions with importance scores less than the second threshold to the low-precision format; Count the number of token positions assigned to the high-precision format, medium-precision format, and low-precision format in each processing block, and select the precision format with the largest number as the unified quantization configuration of the corresponding processing block; Divide the network modules of the language model into multiple configuration sharing groups, where each configuration sharing group contains at least two network modules; Share the unified quantization configuration of the processing block corresponding to the first network module within each configuration sharing group with other network modules within the same configuration sharing group; According to the unified quantization configuration corresponding to each processing block, perform block-level batch quantization on all processing blocks and complete model inference to generate a model inference result.
[0170] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Divide the input text into multiple processing blocks, fix the processing precision format of the first processing block to a high-precision format, and disable the quantization processing of the first processing block; For the other processing blocks except the first processing block among the multiple processing blocks, generate a self-attention matrix for each other processing block through a language model, determine the sum of all element values in the column corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score for each token position; Allocate the token positions with importance scores greater than the first threshold to the high-precision format, allocate the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and allocate the token positions with importance scores less than the second threshold to the low-precision format; Count the number of token positions allocated to the high-precision format, medium-precision format, and low-precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration for the corresponding processing block; Divide the network modules of the language model into multiple configuration sharing groups, where each configuration sharing group contains at least two network modules; Share the unified quantization configuration of the processing block corresponding to the first network module within each configuration sharing group with other network modules within the same configuration sharing group; According to the unified quantization configuration corresponding to each processing block, perform block-level batch quantization on all processing blocks and complete model inference to generate a model inference result.
[0171] It should be noted that for the functions or steps that can be realized by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0172] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0173] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0174] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.
Claims
1. A model quantization reasoning acceleration method, characterized in that: The following steps are involved: Divide the input text into a plurality of processing blocks, fix the processing precision format of the first processing block to a high precision format, and disable quantization processing of the first processing block; For other processing blocks among the multiple processing blocks except the first processing block, generate a self-attention matrix of each other processing block through a language model, and determine the sum of all element values of a column corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score of each token position; Assign the token position whose importance score is greater than the first threshold to the high-precision format, assign the token position whose importance score is above the second threshold and below the first threshold to the medium-precision format, and assign the token position whose importance score is less than the second threshold to the low-precision format; Count the number of token positions in each processing block that are assigned to high-precision format, medium-precision format, and low-precision format, and select the precision format with the largest number as the unified quantization configuration of the corresponding processing block; Dividing the network modules of the language model into a plurality of configuration sharing groups, each configuration sharing group including at least two network modules; In each configuration sharing group, the unified quantized configuration of the processing block corresponding to the first network module is shared with other network modules in the same configuration sharing group; According to the unified quantization configuration corresponding to each processing block, block-level batch quantization is performed on all processing blocks and model inference is completed to generate model inference results.
2. The model quantization inference acceleration method according to claim 1, characterized in that: Divide the input text into multiple processing blocks, fix the processing precision format of the first processing block to a high-precision format, and disable quantization processing of the first processing block, including: Divide the input text into a plurality of processing blocks of equal length according to a preset block length; Disable quantization parameter adjustment for all token positions in the first processing block; Fixing the processing precision format of the embedding layer, the self-attention layer and the feedforward network layer of the first processing block in the language model to a high-precision format; When there are remaining tokens at the end of the input text that are less than the preset block length, the text segment consisting of the remaining tokens is treated as an independent processing block; Fill the independent processing block with invalid tokens to the preset block length, fix the processing precision format of the filled independent processing block to a high-precision format, and disable the quantization operation on the independent processing block; Record the start position index and end position index of all processing blocks.
3. The model quantization inference acceleration method according to claim 1, characterized in that: For other processing blocks except the first processing block in the multiple processing blocks, a self-attention matrix of each other processing block is generated through a language model, and the sum of all element values of a column corresponding to each token position in the self-attention matrix is determined, and the sum of all element values is used as the importance score of each token position, including: Input each other processing block into the multi-head self-attention layer of the language model to obtain the local self-attention matrix of multiple attention heads corresponding to each other processing block; Perform weighted average or arithmetic average processing on the local self-attention matrices of multiple attention heads to generate the final self-attention matrix corresponding to each other processing block; Extract the column vector corresponding to each token position from the final self-attention matrix; Determine the sum of the values of all elements in each column vector, and use the sum of the values as the importance score of the token position corresponding to each column vector; If there are padded invalid token positions in other processing blocks, when determining the importance score of the padded invalid token position, all element values of the column vector corresponding to the invalid token position are set to zero.
4. The model quantization reasoning acceleration method according to claim 1, characterized in that: Allocating a token position whose importance score is greater than a first threshold to a high-precision format, allocating a token position whose importance score is above a second threshold and below the first threshold to a medium-precision format, and allocating a token position whose importance score is less than the second threshold to a low-precision format, including: determining a first threshold and a second threshold according to a length of the processing block; If there is an invalid token position, the processing precision format of the invalid token position is assigned to a low precision format, and the invalid token position is excluded when counting the number of precision formats of each processing block; Compare the importance score of each valid token position with the first threshold and the second threshold; Assigning the processing precision format of the valid token position whose importance score is greater than the first threshold to the high precision format; Assigning the processing precision format of the valid token position whose importance score is above the second threshold and below the first threshold to the medium precision format; The processing precision format of the valid token position whose importance score is less than the second threshold is assigned to the low precision format.
5. The model quantization reasoning acceleration method according to claim 1, characterized in that: Count the number of token positions in each processing block that are assigned to high-precision format, medium-precision format, and low-precision format, and select the precision format with the largest number as the unified quantization configuration of the corresponding processing block, including: Traverse all token positions of the current processing block and identify token positions that are assigned to high-precision format, medium-precision format, and low-precision format; During the counting process, if there are invalid token positions, all invalid token positions will be excluded, and only the precision format allocation results of valid token positions will be counted; Count the counts of valid token positions that are assigned to high-precision format, medium-precision format, and low-precision format, respectively, to generate high-precision counts, medium-precision counts, and low-precision counts; Comparing the numerical values of the high-precision count, the medium-precision count and the low-precision count; Using the precision format corresponding to the count with the largest value as the unified quantization configuration of the current processing block; If there are multiple precision formats whose counts are the same and are the maximum value, the highest precision format among the multiple precision formats is selected as the unified quantization configuration of the current processing block.
6. The model quantization reasoning acceleration method according to claim 1, characterized in that: In each configuration sharing group, the unified quantization configuration of the processing block corresponding to the first network module is shared with other network modules in the same configuration sharing group, including: assigning a unique block identifier to each processing block and assigning a unique group identifier to each configuration sharing group; Establishing a binding relationship between a unique group identifier and a unique block identifier in a configuration mapping table; Based on the binding relationship between the unique group identifier and the unique block identifier, associating the first network module of each configuration sharing group with the unified quantized configuration parameters of the corresponding processing block; Writing the unified quantized configuration parameters into the shared memory area of each configuration sharing group, and assigning the same memory address mapping to all network modules of each configuration sharing group; When executing quantization processing, other network modules in the same configuration sharing group read the unified quantization configuration parameters from the shared memory area through the memory address mapping.
7. The model quantization reasoning acceleration method according to claim 1, characterized in that: According to the unified quantization configuration corresponding to each processing block, block-level batch quantization is performed on all processing blocks and model inference is completed to generate model inference results, including: Loading corresponding unified quantization configuration parameters for each processing block, wherein the unified quantization configuration parameters include a precision format identifier; According to the precision format identifier, configuring independent processing kernel functions for different processing blocks; In a parallel processing unit of a graphics processor or a tensor processor, processing resources are allocated according to the index order of the processing blocks, and quantization processing of all processing blocks is performed concurrently; The processing results of each processing block are checked for the token position within the block to eliminate the intermediate results corresponding to invalid token positions; Sort the valid intermediate results of all processing blocks by starting position index and concatenate them into a complete model output sequence; Post-processing operations are performed on the concatenated output sequence to generate the final model inference results.
8. A model quantization reasoning acceleration device, characterized in that: The model quantization reasoning acceleration device comprises: An input text preprocessing module, used for dividing the input text into a plurality of processing blocks, fixing the processing precision format of the first processing block to a high precision format, and disabling quantization processing of the first processing block; A self-attention analysis module, for generating a self-attention matrix of each of the other processing blocks except the first processing block among the multiple processing blocks through a language model, and determining the sum of all element values of a column corresponding to each token position in the self-attention matrix, and using the sum of all element values as an importance score of each token position; A precision allocation module, configured to allocate a token position whose importance score is greater than a first threshold value to a high precision format, allocate a token position whose importance score is above a second threshold value and below the first threshold value to a medium precision format, and allocate a token position whose importance score is less than the second threshold value to a low precision format; The quantization configuration decision module is used to count the number of token positions in each processing block that are allocated to high-precision format, medium-precision format, and low-precision format, and select the precision format with the largest number as the unified quantization configuration of the corresponding processing block; A network module grouping control module, used to divide the network modules of the language model into a plurality of configuration sharing groups, each configuration sharing group including at least two network modules; A configuration sharing management module, used to share the unified quantized configuration of the processing block corresponding to the first network module in each configuration sharing group with other network modules in the same configuration sharing group; The inference execution module is used to perform block-level batch quantization on all processing blocks and complete model inference according to the unified quantization configuration corresponding to each processing block, and generate model inference results.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a model quantization reasoning acceleration program stored in the memory and executable on the processor. When the model quantization reasoning acceleration program is executed by the processor, the steps of the model quantization reasoning acceleration method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores a model quantization reasoning acceleration program, which, when executed by a processor, implements the steps of the model quantization reasoning acceleration method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Neural network acceleration hardware architecture and method for quantization bit width dynamic selection
CN113902108A
Text reasoning acceleration method applied to large language model and related device
CN118394895A
Method and system for parallel statistical inference on highly parallel platforms
US20110066578A1
Multi-granular clustering-based solution for key-value cache compression
US20250094712A1
Cited By
Large model reasoning efficiency dynamic optimization and hardware sensing compression method
CN120494006A
Dynamic optimization of large model inference performance and hardware-aware compression methods
CN120494006B
Neural network calculation method and device, electronic equipment and storage medium
CN121070445A
Computing method and device of neural network, electronic equipment and storage medium
CN121070445B
Large language model reasoning effective throughput optimization method, system, equipment and medium
CN121996437A