Model Quantization Inference Acceleration Method, Device, Equipment and Medium
By dividing the input text into multiple processing blocks and calculating importance scores based on the self-attention matrix, combined with the quantitative strategy of configuring shared groups, the problem of low quantitative configuration efficiency in long text inference of large language models is solved, and memory optimization and inference speed are achieved, which is suitable for medical health and financial technology fields.
Patent Information
- Application Number
- CN202510525474.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-25
AI Technical Summary
In the process of long text inference of large language models, the existing technology has low efficiency in quantitative configuration, resulting in large overhead of video memory and affecting the speed and accuracy of inference. Especially in the application of medical health and financial technology, the problem of insufficient video memory is serious.
The input text is divided into multiple processing blocks. The first processing block is fixed to a high-precision format and quantization processing is disabled. Other processing blocks calculate the importance score through the self-attention matrix, allocate the accuracy format based on the threshold, and divide the network module into a unified quantization configuration within the configuration sharing group, performing block-level batch quantization inference.
While ensuring the accuracy of inference, it greatly reduces the overhead of video memory and configuration time, improves the execution efficiency and memory utilization of long-text inference tasks, and is suitable for efficient inference needs in the fields of medical health and financial technology.
Smart Images

Figure CN120086355B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and storage medium for accelerating model quantization inference. Background Art
[0002] In recent years, large language models (LLMs) have performed excellently in natural language processing tasks and are widely used in scenarios such as dialogue systems, machine translation, and question-and-answer systems. However, with the continuous expansion of model scale, they face significant video memory occupancy problems when performing long text inference tasks. Most mainstream LLMs have billions or even hundreds of billions of parameters, and during the inference process, the model and related intermediate states need to be fully loaded into the GPU video memory, which poses extremely high requirements on hardware resources.
[0003] When applied to long text inference tasks, such as long document generation, complex literature abstract extraction, or multi-turn dialogue processing, the input sequence often contains thousands to tens of thousands of tokens. Such extremely long sequences significantly increase the computational complexity and video memory requirements of the attention mechanism. Specifically, the dimension of the attention matrix grows with the square of the input sequence length, and a large number of intermediate activation values also need to be temporarily stored in the video memory, which easily leads to memory overflow or inference failure. Existing video memory optimization means, such as gradient checkpointing, activation value offloading, etc., are mainly for the training stage and are difficult to be directly applied to the inference process. In addition, although distributed inference can alleviate the resource bottleneck of a single device, it has high requirements for infrastructure, complex configuration in actual deployment, and is difficult to operate stably in edge or low-to-medium resource environments.
[0004] In the field of medical and health business, LLMs are used in scenarios such as medical record abstract generation, medical question-and-answer systems, and diagnosis and treatment record analysis. These tasks often involve a large number of medical terms and context-related information, and have a stronger dependence on the model's inference ability for long texts. Although the current general quantization compression method can reduce the video memory occupancy, it often ignores the retention of some key information in medical texts, resulting in a decrease in model accuracy and even a deviation in the understanding of medical information, affecting the credibility of the results.
[0005] In the field of fintech business, large language models are widely used in text-intensive tasks such as contract parsing, financial summary generation, and customer risk assessment. These tasks usually involve compliance reports or historical transaction records with complex structures and extremely long texts, posing extremely high requirements on the video memory efficiency during the inference process. However, the current model is prone to failure due to insufficient video memory when processing long text inputs such as financial documents, seriously affecting the stability and scalability of the model.
[0006] To alleviate the video memory pressure, quantization technology has become one of the mainstream compression means, especially widely adopted in the inference scenario. Existing research attempts to use the mixed-precision quantization method, that is, allocating higher bit widths to high-importance tokens and lower bit widths to low-importance tokens, so as to reduce the overall video memory usage while maintaining the model accuracy. However, the mixed-precision quantization faces the problem of low configuration selection efficiency during the inference process. Since the current method needs to dynamically analyze the importance of each token and determine the bit width strategy in each round of inference, this process itself is time-consuming and weakens the speed advantage brought by quantization, especially more significantly in long text tasks.
[0007] In summary, the existing technologies still have key deficiencies in dealing with the video memory overhead problem during the long text inference process of large language models. Especially in terms of the efficiency of quantization configuration strategies, the importance recognition mechanism across task domains, and block-level inference resource scheduling, further optimization is still needed to meet the actual needs of efficient and accurate inference in fields such as fintech and healthcare. Summary of the Invention
[0008] The main purpose of the present invention is to provide a model quantization inference acceleration method, device, equipment and storage medium, aiming to solve the technical problem that the existing technology needs to determine the quantization bit width for each token during the inference process, resulting in low quantization configuration efficiency, especially seriously affecting the inference speed and video memory optimization effect in long text inference tasks.
[0009] To achieve the above object, the present invention provides a model quantization inference acceleration method, including:
[0010] Divide the input text into multiple processing blocks, fix the processing precision format of the first processing block as the high-precision format, and disable the quantization processing of the first processing block;
[0011] For the other processing blocks except the first processing block among the multiple processing blocks, generate the self-attention matrix of each other processing block through the language model, determine the sum of all element values corresponding to each column at the token position in the self-attention matrix, and use the sum of all element values as the importance score of each token position;
[0012] Allocate the token positions with importance scores greater than the first threshold to the high-precision format, allocate the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and allocate the token positions with importance scores less than the second threshold to the low-precision format;
[0013] Count the number of token positions assigned to high-precision format, medium-precision format, and low-precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration for the corresponding processing block;
[0014] Divide the network modules of the language model into multiple configuration sharing groups, and each configuration sharing group contains at least two network modules;
[0015] Within each configuration sharing group, share the unified quantization configuration of the processing block corresponding to the first network module with other network modules within the same configuration sharing group;
[0016] According to the unified quantization configuration corresponding to each processing block, perform block-level batch quantization on all processing blocks and complete model inference to generate model inference results.
[0017] Furthermore, to achieve the above object, the present invention provides a model quantization inference acceleration device, including:
[0018] An input text preprocessing module, configured to divide the input text into multiple processing blocks, fix the processing precision format of the first processing block as high-precision format, and disable the quantization processing of the first processing block;
[0019] A self-attention analysis module, configured to generate a self-attention matrix for each of the other processing blocks except the first processing block among the multiple processing blocks through a language model, determine the sum of all element values in the column corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score of each token position;
[0020] A precision allocation module, configured to allocate token positions with importance scores greater than the first threshold to high-precision format, allocate token positions with importance scores above the second threshold and below the first threshold to medium-precision format, and allocate token positions with importance scores less than the second threshold to low-precision format;
[0021] A quantization configuration decision module, configured to count the number of token positions assigned to high-precision format, medium-precision format, and low-precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration for the corresponding processing block;
[0022] A network module grouping control module, configured to divide the network modules of the language model into multiple configuration sharing groups, and each configuration sharing group contains at least two network modules;
[0023] A configuration sharing management module, configured to share the unified quantization configuration of the processing block corresponding to the first network module with other network modules within the same configuration sharing group within each configuration sharing group;
[0024] An inference execution module, configured to perform block-level batch quantization on all processing blocks according to the unified quantization configuration corresponding to each processing block, complete model inference, and generate model inference results.
[0025] Furthermore, to achieve the above object, the present invention also provides a computer device, which includes a memory, a processor, and a model quantization inference acceleration program stored on the memory and executable on the processor. When the model quantization inference acceleration program is executed by the processor, the steps of the model quantization inference acceleration method as described above are implemented.
[0026] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a model quantization inference acceleration program is stored. When the model quantization inference acceleration program is executed by a processor, the steps of the model quantization inference acceleration method as described above are implemented.
[0027] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health and fintech. It discloses a model quantization inference acceleration method, including: dividing the input text into multiple processing blocks, fixing the calculation precision format of the first processing block as a high-precision format and disabling the quantization operation; for other processing blocks except the first processing block, generating a self-attention matrix through a language model, and calculating the importance score based on the sum of the values in the corresponding column of the self-attention matrix for each token position; allocating each token position to a high-precision, medium-precision, or low-precision format based on two preset thresholds; counting the number of tokens in each precision format within each processing block, and selecting the precision format with the largest number as the unified quantization configuration; dividing the network module into multiple configuration sharing groups, and sharing the unified quantization configuration of the processing blocks within the group; performing block-level batch quantization according to the unified quantization configuration of the processing blocks and completing model inference to generate inference results. By uniformly determining the quantization configuration of each processing block based on the token importance score and reusing this configuration within the network module group, the present invention realizes precision allocation and parallel quantization inference at the block level, greatly reducing the video memory overhead and configuration time overhead while ensuring the inference precision, and effectively improving the execution efficiency and video memory utilization rate in long text inference tasks. Description of the Drawings
[0028] The following will further illustrate the present invention in conjunction with the drawings. In the drawings:
[0029] Figure 1 is a schematic diagram of an application environment of the model quantization inference acceleration method in an embodiment of the present invention;
[0030] Figure 2 is a schematic flowchart of an embodiment of the model quantization inference acceleration method of the present invention;
[0031] Figure 3 Schematic diagram of functional modules of a preferred embodiment of the model quantization inference acceleration device of the present invention;
[0032] Figure 4 Schematic diagram of a structure of a computer device in an embodiment of the present invention;
[0033] Figure 5 Schematic diagram of another structure of a computer device in an embodiment of the present invention. Detailed implementation manners
[0034] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0035] The model quantization inference acceleration method provided by the embodiments of the present invention can be applied in an application environment such as Figure 1 , where the client communicates with the server through a network. The server can divide the input text into multiple processing blocks through the client, fix the calculation precision format of the first processing block as a high-precision format and disable the quantization operation; for other processing blocks except the first processing block, generate a self-attention matrix through a language model, and calculate the importance score based on the sum of the values in the corresponding column of each token position in the self-attention matrix; allocate each token position to a high-precision, medium-precision or low-precision format based on two preset thresholds; count the number of tokens in each precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration; divide the network module into multiple configuration sharing groups, and share the unified quantization configuration of the processing block within the group; perform block-level batch quantization according to the unified quantization configuration of the processing block and complete model inference to generate an inference result. The present invention realizes precision allocation and parallel quantization inference at the block level by uniformly determining the quantization configuration of each processing block based on the token importance score and reusing this configuration within the network module group, greatly reducing the video memory overhead and configuration time overhead while ensuring the inference precision, and effectively improving the execution efficiency and video memory utilization rate in long text inference tasks. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0036] Please refer to Figure 2 , Figure 2 which is a flowchart of an embodiment of the model quantization inference acceleration method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.
[0037] Such as Figure 2As shown in the figure, the model quantization inference acceleration method proposed by the present invention includes the following steps:
[0038] S10, divide the input text into multiple processing blocks, fix the processing precision format of the first processing block to a high-precision format, and disable the quantization processing of the first processing block;
[0039] In this embodiment, dividing the input text into multiple processing blocks is to control the distribution of computing resources in model inference for long text inputs, and perform local management on the activation values, attention matrices, and video memory usage of the model. In practical applications, text processing blocks can be generated through a sliding window of a fixed length. For example, every 512 tokens are divided into one block, and this length is determined according to the maximum processing length of the training window or inference window of the model. This division method comes from the window sharding mechanism widely used in sequence modeling, and its purpose is to transform long text into local context blocks, which not only ensures local semantic coherence but also avoids the computing bottleneck caused by loading all tokens at once.
[0040] Fixing the processing precision format of the first processing block to a high-precision format is to provide a stable computing starting point for the model at the initial stage of inference. The high-precision format generally refers to the floating-point format, such as FP32 or FP16, which has stronger robustness and fault tolerance in terms of numerical expression range and expression precision compared to low-bit-width fixed-point formats (such as INT8 or INT4).
[0041] Disabling the quantization processing of the first processing block is to further ensure that this segment of the input runs in full-precision expression, and avoid the cumulative interference of quantization errors on the subsequent block propagation path at the initial stage of model propagation. Disabling quantization not only includes not applying quantization operations to the weights and activation values of this processing block, but also includes skipping the quantization configuration generation process of this processing block. Its core purpose is to provide a reference benchmark in quantization decision-making, making the first processing block a relative anchor point for the importance scoring and precision configuration judgment of subsequent blocks.
[0042] In practical applications, quantization processing usually involves discretizing the weight parameters of the model and compressing the activation values into a limited-bit-width expression domain. If this operation is still performed on the first processing block, it will cause the distortion of the initial attention result, thereby reducing the accuracy of precision allocation for subsequent token positions. Therefore, disabling the quantization processing of the first processing block is not only a consideration of numerical stability but also a basic premise for the dynamic precision adjustment strategy in the entire inference path.
[0043] Since models usually adopt residual connections and normalization mechanisms, the high-precision first block helps to stabilize the statistical behavior of the first few layers, thereby improving the discrimination accuracy of subsequent processing blocks in the quantization strategy decision. This feature also has good generalization and does not depend on the specific model structure, so it can be applied to large language models with different architectures such as GPT, T5, and BERT.
[0044] In the specific implementation, after loading the input text as a token sequence, it can be segmented according to the preset processing block length, and each segment forms a processing block. The token range of the first processing block can be the first 512 tokens of the token sequence. After the division, configure the calculation precision format of this processing block as FP16 or FP32. During model inference, skip the quantization configuration generation process of this block, that is, do not participate in importance scoring and do not perform bit-width decision. To disable quantization processing, when calling the model for execution, for the first processing block, forcibly bypass the quantization operator deployed in the model. This can be achieved by configuring the quantization mask parameter of the inference engine, or during the model graph conversion process, setting the data path of this block as a non-quantized path. Further, when this processing block passes through the embedding layer, multi-head attention layer, and feed-forward layer of the model, all can be executed with floating-point precision, without calling low-precision computing kernel functions. In the actual model running environment, if using an inference framework such as TensorRT, the first block can be marked as a resident high-precision area when building the engine, or achieved by adding a first block precision locking instruction during compilation. In the multi-block execution process, the output of the first processing block can also be used as a reference input for the importance analysis and precision allocation strategy generation of subsequent processing blocks.
[0045] Example illustration: In the field of medical and health business, when processing text containing long medical diagnosis records, it is often the case that key disease descriptions appear at the beginning of the text. If the first processing block loses information due to quantization, it may cause the model to fail to correctly extract disease symptoms and corresponding analyses. After adopting fixed first block high-precision processing, the model can more accurately capture the high-value content in the first paragraph, which is beneficial to subsequent generation of accurate diagnosis summaries or disease course predictions.
[0046] In the field of fintech business, when processing a complete risk report, the report usually contains global risk classification and summary information at the beginning. If it is placed in a low-precision calculation path, it is easy to cause the model to misjudge the risk level, thus affecting the subsequent judgment process. By maintaining high-precision calculation and disabling quantization for the first processing block, the core points of the report can be effectively captured, ensuring the inference accuracy and stability of the risk control model, and enhancing the reliability and security of financial service decisions.
[0047] By fixing the first processing block in high-precision format and disabling its quantization during the inference process, the stability and accuracy of subsequent processing blocks in precision strategy discrimination can be significantly improved. The high-precision first block provides complete context representation ability, enabling its output to be used as a benchmark for weight allocation and quantization level selection, while avoiding model drift caused by quantization errors in the propagation of initial information. In long text scenarios, this strategy can effectively reduce the global precision degradation problem caused by misjudgment of the first paragraph, thus improving the overall inference performance and stability.
[0048] S20. For other processing blocks in the multiple processing blocks except the first processing block, generate the self-attention matrix of each other processing block through a language model, determine the sum of all element values corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score of each token position.
[0049] In this embodiment, for other processing blocks in the multiple processing blocks except the first processing block, the self-attention mechanism is introduced to generate the corresponding self-attention matrix, which is mainly used to measure the information dependence relationship between different token positions within the same processing block. In natural language processing tasks, large language models (LLMs) usually adopt the self-attention mechanism based on the Transformer structure to construct the context representation between tokens. By separately constructing the self-attention matrix for each processing block, the calculation range can be localized, thereby effectively controlling the memory occupancy and improving the calculation efficiency.
[0050] The sum of all element values corresponding to each token position is used as the importance score. This operation is essentially a column vector summation, and its technical meaning is to measure the overall attention concentration degree of a certain token in this processing block. The larger this value is, the stronger the connection between this token and other tokens, and the wider the propagation range of its semantic information within the block, so it is considered more "important".
[0051] In self-attention calculation, the column vector corresponding to each token position represents the attention weights assigned by this token as the target token to all source tokens. On this basis, summing the column vectors can obtain its global attention degree. The importance score obtained in this way does not depend on specific task labels, so it has good generality and is applicable to the requirements of token selective precision control in various tasks such as machine translation, question answering systems, and information extraction.
[0052] To avoid introducing unnecessary noise and redundant calculations, the self-attention matrix should skip padding tokens during generation. That is, a mask strategy can be adopted for the positions of padding tokens, forcing their corresponding attention values to zero to prevent them from interfering with the importance evaluation of other tokens. This design also applies to the multi-head attention mechanism (Multi-head Attention), where the local matrices calculated by each attention head can be averaged and fused first to form the final self-attention matrix.
[0053] In a specific implementation, the token sequence of each processing block is fed into the multi-head self-attention module for forward propagation to obtain the attention matrices of multiple attention heads. By performing weighted average or simple arithmetic average on these local matrices, the final self-attention matrix of the current processing block is generated. Subsequently, each column of this matrix is traversed, and the sum of all values in the column vector is calculated as the importance score of the token position corresponding to this column. During this process, the column vector corresponding to the padding token will be replaced with a vector of all zeros, so that its importance score is always zero, ensuring that it does not participate in the subsequent selection of precision configuration.
[0054] For scenarios with high computational performance requirements, the generation of the self-attention matrix and the column vector summation operation can be fused into a GPU kernel function (CUDA Kernel) to batch process multiple processing blocks under the GPU parallel framework, improving the overall processing efficiency. In addition, for cases where the numerical precision in the attention matrix is relatively low (such as FP16), a normalization or numerical smoothing mechanism can be introduced to avoid misjudgment caused by the influence of numerical drift on the token scores.
[0055] Example illustration: In the field of healthcare, texts such as doctors' consultation records, patients' self - reports, and test reports often contain a large number of tokens, while the words that truly affect subsequent decisions are only a very small number of highly specialized and context - highly - related medical terms. Take an electronic medical record as an example. It may contain information fragments such as patient age, underlying diseases, chief complaints, and examination results. When generating the self - attention matrix of this processing block, for example, for the phrase "abnormal liver function", the token positions corresponding to it in the matrix show strong attention connections with the token positions of "elevated ALT", "jaundice", "positive hepatitis B surface antigen", etc. By calculating the sum of the elements in the column corresponding to this token, its semantic importance in the entire fragment can be quantified. Furthermore, these high - score tokens will be marked as important tokens, providing higher computational precision support for subsequent diagnostic classification or automatic medical record summary tasks. For tokens with high generality but low context semantic weight, such as "patient gender" and "this follow - up visit", the sum of the elements in their corresponding columns is smaller, and they are naturally assigned lower importance scores, so they are calculated with lower precision in subsequent inferences, saving computing power resources without affecting the expression of the core semantics.
[0056] In intelligent customer service conversations, user complaint handling, or transaction log parsing in the financial field, important context information is usually distributed in tokens containing words such as "risk", "freeze", "fraud", "delayed payment", etc., and the context dependence of such words is usually relatively high. In the self - attention matrix, the token "fraud" forms strong connections with tokens such as "transfer failure", "funds frozen", and "strange contacts", and the sum of the elements of its corresponding column vector is significantly higher than that of other background words such as "hello", "please", "thank you", etc. tokens, so it is marked as a high - importance position in this step. In the subsequent inference of the risk judgment model, key tokens are assigned high - precision calculations to ensure that the understanding of sensitive expressions is not distorted due to low - bit - width quantization, especially crucial in user inputs with fuzzy semantic boundaries and non - standard expressions. Through this attention - based token - level precision allocation strategy, the video memory occupancy and computational load during inference can be effectively compressed without sacrificing the sensitivity and accuracy of risk control.
[0057] Such example scenarios show that in long - text tasks, by capturing the semantic coupling degree between tokens through the self - attention matrix and determining importance based on column vector summation, not only the technical goal of token - level precision control is achieved, but also cross - domain general adaptability and high inference efficiency are ensured.
[0058] By summing the self-attention matrices of each processing block column by column and calculating the importance scores, the positions of the tokens that contribute the most to semantic expression can be accurately identified based on the context dependence strength of the tokens. When further allocating precision on this basis, it no longer relies on external feature extraction or complex rule matching, significantly reducing the pre-processing computational overhead in the inference stage and achieving the unity of important token recognition and precision control.
[0059] S30, assign the token positions with importance scores greater than the first threshold to the high-precision format, assign the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and assign the token positions with importance scores less than the second threshold to the low-precision format;
[0060] In this embodiment, using the importance scores of each token position to drive the configuration of the subsequent calculation precision is the key path to implementing the differential quantization strategy. The "importance score" here comes from the numerical representation generated based on the self-attention matrix in the previous stage, which is used to reflect the information influence of each token in its context. The higher this score, the higher the degree to which the token is concerned by other tokens in the current processing block, that is, the stronger its role in semantic propagation and the lower the tolerance for semantic masking. Therefore, this value directly determines the computational precision required for the token in the inference.
[0061] The high-precision format usually corresponds to the floating-point calculation format, such as half-precision floating-point (FP16) or single-precision floating-point (FP32), achieving a balance between computational complexity and expression precision. The medium-precision format can be medium-width fixed-point quantization, such as INT4, which can save a part of the video memory resources and computational overhead while ensuring the accuracy of the main semantic expression. The low-precision format usually adopts a smaller-width fixed-point quantization format, such as INT2 or further INT1 (binary) representation, which is suitable for token positions with higher redundancy and strong semantic fault tolerance, significantly compressing the computational load.
[0062] Hierarchical control is achieved by setting the first threshold and the second threshold, where the first threshold is greater than the second threshold. The second threshold is used to identify the tokens with the lowest importance, while the first threshold identifies the most important batch of tokens, and the tokens between the two fall into the medium-precision range. The setting principles of the first threshold and the second threshold can be dynamically set according to the processing block length, model scale, and current hardware resources, or determined through empirical model training. This dual-threshold mechanism makes the precision configuration have both controllable policy flexibility and fine-grained scheduling ability for computational resources.
[0063] The process of comparing a fraction with two thresholds can usually be executed in parallel in the weight calculation thread during the preprocessing stage. The precision level of each token position can be determined through a single floating-point comparison operation, and then it is labeled with the corresponding precision label for subsequent configuration generation and inference allocation. The entire allocation mechanism does not depend on the modification of the model structure and only acts on the quantization strategy layer, with good system compatibility and deployment adaptability.
[0064] Hierarchical precision configuration can be achieved through static threshold setting. For example, the second threshold is set to 0.2, the first threshold is set to 0.8, and the token importance scores are uniformly distributed in the range of [0, 1]. Then, tokens greater than 0.8 are assigned FP16, the middle segment (between 0.2 and 0.8, including 0.8 and 0.2) uses INT4, and those below 0.2 are INT2. Dynamic threshold calculation can also be performed based on the score distribution of each processing block. For example, the quantile method is used, with the top 10% being high precision, the bottom 30% being low precision, and the middle being medium precision. In addition, task-aware factors can be incorporated to preferentially label known important tokens in specific semantic annotation tasks as high precision, that is, to enhance precision under the guidance of attention.
[0065] In the specific implementation, each token position records a precision label (e.g., a 2-bit identifier: 00 for low precision, 01 for medium, 10 for high precision), which is stored in the video memory in the form of a bitmap. After the quantization configuration module reads this identifier, it calls the quantization kernel function with the corresponding bit width to complete the low-level inference binding. Multiple tokens can be batch-labeled, and concurrent allocation is achieved with the help of the SIMD (Single Instruction Multiple Data) instruction set to improve throughput efficiency. For distributed scenarios, the precision allocation and marking can be completed locally on each GPU, and the configuration bitmap is uploaded to the shared control module to achieve cross-block consistent scheduling.
[0066] By allocating precision formats according to the importance scores at the token level, key tilting of computing resources can be achieved, that is, higher computing power is used for critical tokens, and the precision resource investment in low-weight information is reduced. While maintaining a high inference accuracy for the overall model, the video memory overhead and inference latency are effectively controlled.
[0067] S40, count the number of token positions allocated with high-precision format, medium-precision format, and low-precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration for the corresponding processing block;
[0068] In this embodiment, after the multi-level precision allocation is completed, it is necessary to count the distribution of different precision formats within each processing block to provide a decision basis for the subsequent selection of unified quantization configuration. Each token position has been assigned to a high-precision format, a medium-precision format, or a low-precision format after the comparison of the previous importance scores. The goal of this step is to traverse the valid token positions of the entire processing block and count the number of token positions corresponding to the three types of precision formats respectively.
[0069] The quantization configuration of each processing block must be unified to match the scheduling requirements of block-level parallel execution in the hardware computing unit. For example, when the Tensor Core of the GPU executes calculations such as INT8 or FP16, it is required that each operation targets a unified matrix precision. Therefore, it is necessary to determine the overall calculation format of the processing block in advance. This step determines the dominant precision distribution of the current processing block based on the precision statistics results after quantization.
[0070] During the traversal process, invalid token positions do not participate in the statistics to avoid the padding tokens causing bias to the true distribution. In actual implementation, the mask mask recorded during the padding stage is usually used to quickly exclude the padding area, and only the precision labels of the valid positions are counted. After the counting is completed, compare the quantities of the three types of precision formats and select the precision format with the largest quantity as the unified quantization configuration of the processing block.
[0071] In some special cases, if the number of tokens of two or three precision formats is the same and is the maximum value, in order to ensure the stability and conservativeness of the model inference precision, the precision format with a higher precision level is preferentially selected. For example, when the number of high-precision and medium-precision is equal, the high-precision format is selected as the quantization configuration of the processing block. This conservative bias strategy ensures that the model can retain the information expression ability of the key semantic paths as much as possible without increasing the computational burden.
[0072] This operation is not only a configuration strategy based on statistical dominant trends, but also a technical mechanism to ensure the consistency of the quantization strategy within the processing block. Through this mechanism, each processing block is assigned a unique unified quantization configuration, so that the corresponding kernel function can be called and hardware-level block parallel acceleration can be achieved in subsequent inferences.
[0073] A bitmap structure can be used to represent the precision labels of each token position. Each bit represents the precision category of the current token. For example, two-bit binary identifiers are used to represent high, medium, and low precision respectively. During the statistical stage, logical AND operations are used to extract the precision flag bits of each processing block, skip the invalid token positions according to the mask, and perform bit counting on the three precision marks respectively.
[0074] In a multi-core system, statistical operations can be performed in parallel in blocks, with each thread processing the statistical tasks of a processing block and writing the results to the shared memory or quantization configuration register table. After saving the three types of count values in the shared memory, a comparison operation is performed to determine the maximum value, and the precision configuration identifier is generated according to the priority strategy. If a dynamic precision adjustment strategy is adopted, the dynamic quantization bit width setting of the current processing block can be adjusted synchronously after the statistics are completed.
[0075] In order to further improve the intelligence and adaptability of the unified quantization configuration of the processing block, on the basis of selecting the precision configuration based on the statistical results of the precision format, a variety of semantic and historical context features can be introduced to perform weighted fine-tuning on the statistical counting results to reflect the sensitivity of the language model to local semantic changes and dynamically adapt to the actual demand for computing resources in different semantic areas in the reasoning task.
[0076] First, when processing long text input in natural language, there is a phenomenon that the semantic density of tokens changes with the context semantic field and presents an uneven distribution. Taking highly structured documents such as legal contracts and medical test reports as examples, their text fragments may contain a large number of repeated, high-frequency and information-redundant tokens (such as terms, connectors, unit names, etc.). Although such tokens appear frequently in the literal sense, they often do not constitute the focus of information in actual semantics. If the precision format is divided only based on the original importance score calculated by attention, the number of medium or low-precision tokens may be too small in the statistical process, thereby mistakenly configuring such processing blocks as high-precision, increasing computational redundancy.
[0077] To this end, a high-frequency token importance suppression mechanism based on a sliding window can be introduced. This mechanism counts the frequency of tokens that appear in multiple adjacent processing blocks during the statistical process. If the frequency of a token in the sliding window exceeds the preset threshold (such as appearing in 70% of the blocks), the importance score of the token is appropriately reduced by a proportional factor (such as multiplying by 0.8), so that it is more inclined to be classified as a medium or low-precision format during statistics, avoiding high-frequency redundant tokens dominating the overall accuracy configuration selection. The sliding window can be a processing block sequence of length 3 to 5, and the window length can be dynamically adjusted according to the total input length to take into account both response speed and local feature capture capabilities.
[0078] Secondly, the preamble semantic block influence factor can be introduced. Since there is often strong semantic coherence between processing blocks. For example, in financial public opinion analysis, after the "credit card overdue" processing block, semantic blocks such as "collection, default, credit investigation" follow closely. When calculating the statistical accuracy, the accuracy tendency in the previous processing block should be referred to, so as to enhance the accuracy configuration strategy of the semantic consistency between the current block and the above text. In terms of implementation, the proportion of the number of tokens of various accuracy formats in the previous processing block can be recorded, and a weighting factor can be introduced for adjustment based on the accuracy count of the current processing block. For example, if the number of medium-accuracy tokens in the current block is 200, and the proportion of high-accuracy tokens in the previous block is high, the current high-accuracy count can be multiplied by a forward adjustment factor of 1.1 to increase the possibility of its configuration selection.
[0079] Furthermore, for the specific requirements of certain tasks, a context weight awareness model can also be introduced to dynamically generate the accuracy configuration weight vector of the current processing block through a lightweight forward network module (such as one-layer FFN + Softmax). The input features include the original three types of accuracy token counts of this block, the proportion of high-frequency tokens in the sliding window, the accuracy distribution of the previous block, etc. The relative weight of each type of accuracy is adjusted through the weights learned by the forward model, so as to more carefully control the selection logic of the final accuracy configuration. For example, in the scenario of medical and health text generation, this awareness model can identify the points of "pathological state transition" or "change in treatment response", and tend to select a higher accuracy configuration.
[0080] Finally, when the numerical values of the high, medium, and low accuracy format counts are close, in order to avoid the problem of discontinuous context accuracy caused by quantization jitter, a dynamic gating mechanism can be introduced. The gating module judges whether the accuracy configuration distribution of the current processing block is balanced based on the accuracy distribution entropy (such as Shannon entropy or Top-k proportion). When it is determined that there is "no obvious dominant accuracy", the medium accuracy configuration is forcibly selected, so as to achieve a soft regulation of the overall computing resources of the system. This mechanism is particularly applicable to large-scale model deployment scenarios, and can effectively control the power consumption peak and improve the overall inference throughput of the model.
[0081] Example illustration: In the field of medical and health, when facing the task of medical record summary, the input may include content such as diagnosis conclusions, symptom descriptions, treatment plans, etc. For example, the processing block contains "The patient has intermittent headache and blurred vision, suspected of increased intracranial pressure". In the previous steps, "headache", "blurred", and "intracranial pressure" are marked as high accuracy, "patient", "appears", "suspected", etc. are marked as medium accuracy, and the rest of the conjunctions are marked as low accuracy. During the statistical process, it is found that the number of medium-accuracy tokens is slightly more than that of high-accuracy tokens, but the numbers of high-accuracy and medium-accuracy tokens are close. According to the conservative strategy, the system selects high accuracy as the unified configuration of the processing block to ensure the inference accuracy of the subsequent diagnostic model.
[0082] In the financial business scenario, when analyzing the historical transaction behavior of users to generate risk predictions, a certain processing block includes "the user transferred money three times today, and each amount exceeded the limit". In the importance scoring stage, "transfer", "limit", and "amount" are assigned a high-precision format, and the rest are medium or low precision. After precision statistics, the number of high-precision tokens dominates, and the system configures this processing block with FP16 precision and uniformly calls the floating-point kernel function for processing. This strategy ensures that the system can accurately identify and early warn of abnormal transaction patterns.
[0083] By counting the number of tokens in each precision format and using the principle of the largest proportion for configuration selection, the standardization of the internal precision strategy of the processing block is effectively achieved, enabling parallel quantization calculations to be based on the unified configuration at the block level in the subsequent inference stage, improving the system execution efficiency and simplifying the kernel function scheduling path. At the same time, the precision priority rule is introduced to ensure that the semantic expression ability of the model is not damaged first in case of configuration conflicts, thus maintaining the model precision stability while saving resources.
[0084] S50, divide the network module of the language model into multiple configuration sharing groups, and each configuration sharing group contains at least two network modules;
[0085] In this embodiment, the language model is usually composed of multiple continuously stacked network modules, and each module includes components such as an embedding layer, a feed-forward network layer, and a multi-head self-attention layer. In order to control the video memory overhead and improve the inference efficiency in the context of long text input, these network modules can be divided into several "configuration sharing groups" in the processing order. The modules within each group share the same precision configuration and quantization parameters to reduce unnecessary repeated calculations.
[0086] The processing order means grouping according to the natural front and back arrangement of the network modules in the model execution path, rather than randomly shuffling or dynamically organizing based on task allocation. This order division method can ensure the stability and simplicity of the precision configuration propagation logic. Each configuration sharing group contains several adjacent network modules, and the number can be determined by a preset parameter, which can be set according to dimensions such as the depth of the model, the calculation amount of each layer, and the available computing resources. Common preset numbers include 4, 8, 12, etc., and the specific value can be determined through experimental optimization. Each configuration sharing group contains at least two network modules.
[0087] The core of the introduction of the configuration sharing group is to form an execution structure of "block-level inference - group-level reuse". By unifying the quantization schemes of all modules in the configuration sharing group, the number of quantization configurations that need to be independently managed during execution can be significantly reduced, and the video memory and scheduling overhead of frequently loading configurations by different modules in model inference can be reduced. This strategy of dividing in sequence and setting a fixed number within the group can be divisible by the number of modules in terms of structure, thus achieving unified scheduling and allocation in hardware orchestration.
[0088] It should be noted that the configuration sharing group and the "processing block" generated by chunking the previous input text are two-dimensional structures: the former acts inside the model structure and is used for optimizing the quantization configuration between network modules; the latter acts on the organizational structure of the input data and is used to control the input length and computational division. A connection is established between the two through a binding mechanism (such as the first module in the group binding to the corresponding processing block).
[0089] For example, a large language model with 96 network modules can be divided into 12 configuration sharing groups, with each group containing 8 consecutive modules. The division is carried out in the order of module numbers. For example, Group_1 includes Module_1 to Module_8, Group_2 includes Module_9 to Module_16, and so on. During the configuration phase, a unique group identifier is assigned to each sharing group, and a mapping relationship between the group identifier and the bound processing block is established.
[0090] During the execution phase, when the model inference task reaches a certain configuration sharing group, the system first reads the corresponding unified quantization configuration parameters of the group from the bound processing block and maps the parameters for use by all modules within the group. No separate quantization configuration initialization is performed for all modules within the group, thus saving time and storage overhead.
[0091] In large deep models in special scenarios, due to the non-uniformity of network modules at different depths in terms of semantic modeling strength, information flow characteristics, and quantization sensitivity, using a unified number division for configuration sharing groups may not achieve the best balance between accuracy and efficiency. In this case, a non-uniform division method can be introduced, and by flexibly adjusting the number of network modules included in each configuration sharing group, the division strategy can better fit the model hierarchical structure and specific task requirements.
[0092] For example, the modules near the input side usually undertake underlying semantic encoding tasks (such as positional embedding, lexical modeling, etc.). Their computational structures are relatively simple and insensitive to precision loss. Therefore, multiple adjacent modules can be combined into a larger configuration sharing group to improve the parameter reuse efficiency and reduce the configuration switching frequency. On the contrary, the modules near the output side mainly complete high-level semantic abstraction and decision-making information integration. These modules are more sensitive to quantization precision. If a unified configuration sharing strategy is adopted, it is easy to cause a decline in precision. Therefore, these modules can be divided into smaller configuration sharing groups, or even configured individually for every two modules, so as to retain more quantization flexibility in the output layer and adapt to the representation requirements of high-dimensional output features.
[0093] To further improve adaptability, this non-uniform partitioning strategy can be combined with the Neural Architecture Search (NAS) mechanism. The specific approach is to use different partitioning strategies (such as sharing group size, sharing group position distribution, bound block number, etc.) as part of the architecture search space, and introduce control variables during the pre-training or distillation stage for joint performance-efficiency optimization. The architecture search results can feedback which deep regions are more sensitive to precision, thus guiding the dynamic adjustment of the sharing group layout and finally forming an adaptive configuration sharing mechanism that takes into account computing power distribution and precision stability.
[0094] In addition, when deploying deep Transformer architectures for tasks such as medical health dialogue systems or financial transaction behavior modeling, a semantic density-driven dynamic partitioning mechanism can also be introduced: for example, according to the gradient activation distribution caused by the input tokens in the model, the group partitioning method of each segment of the network can be adjusted dynamically in real time. Regions with high semantic density (such as disease diagnosis conclusions or transaction behavior breakpoints) can be configured with a smaller granularity of group partitioning to improve configuration fidelity; while information redundancy regions (such as repeated inquiries or invalid behaviors) can use large group sharing, thus reducing unnecessary waste of computing resources.
[0095] In summary, non-uniform partitioning not only provides a more flexible configuration reuse path, but also provides a structure-aware optimization path for high-performance quantization inference. It is an important enhancement strategy for the efficient deployment of complex models in resource-constrained environments.
[0096] By dividing the network modules of a language model into multiple configuration sharing groups containing a preset number of modules, the time and video memory consumption of repeatedly calculating quantization configurations for each layer of modules during the inference stage can be significantly reduced, and the configuration scheduling logic can be simplified. This strategy effectively introduces a configuration reuse mechanism at the module structure level, realizing batch control of computing resources while ensuring quantization flexibility. By sharing precision configurations within the group, a large number of redundant configuration loading and switching operations are avoided, resulting in an overall improvement in the execution efficiency of long text inference tasks.
[0097] S60. Share the unified quantization configuration of the processing block corresponding to the first network module within each configuration sharing group with other network modules in the same configuration sharing group;
[0098] In this embodiment, in order to improve the execution efficiency and precision consistency of block-level quantization inference, the network modules are divided into multiple configuration sharing groups, and the unified quantization configuration of the processing block associated with the first network module is reused within each group. The essence of this operation is to construct a cross-module parameter sharing mechanism, avoiding generating or loading quantization configurations separately for each network module during the inference process, thereby reducing redundant configuration overhead and improving the overall operation efficiency and memory utilization rate.
[0099] In the specific implementation, first, the network modules of the language model need to be divided into multiple configuration sharing groups according to the processing order. Each sharing group contains a preset number of network modules. For example, each group contains 4 consecutive Transformer sub-layers, or is equally divided according to the model depth. Subsequently, the first network module is selected from each sharing group, and the unified quantization configuration of the processing block associated with it is used as the reference configuration for the group. This processing block may come from the actual inference data in the model input stage, or may be generated by simulated data or task prior data, ensuring that the selected configuration is representative.
[0100] To achieve efficient sharing, the unified quantization configuration needs to be written into a shared memory area, and the same memory address is mapped for all network modules in the same configuration sharing group, thereby avoiding repeated configuration loading and conversion operations. The shared quantization configuration usually includes quantization bit width (such as INT8 or INT4), scale factor (scale), zero point, and calculation precision format identifier (such as whether it is mixed-precision floating-point calculation, etc.). For configuration parameters generated by weighted average, historical accumulation, or semantic importance aggregation strategies, they can also be directly written into the shared area, so that the sharing mechanism does not affect the flexibility of the original configuration generation strategy.
[0101] This configuration sharing mechanism not only improves the operation efficiency in the inference stage, but also enhances the local consistency of precision. Especially in the structural design where there is semantic recursive enhancement between the deep modules of the model, using consistent quantization configurations can reduce error accumulation and quantization jitter effects. To ensure effectiveness in dynamic task scenarios, a lightweight synchronization mechanism can be further introduced into the sharing group. For example, if it is detected that the configuration of the processing block bound to the first network module is updated, the system can automatically trigger the shared memory refresh operation through the configuration mapping table to ensure that the subsequent modules read the latest configuration.
[0102] In the Transformer architecture, assume there are 48 network modules, which are divided into 12 configuration sharing groups, with 4 modules in each group. The first module in each group is bound to a processing block with a representative importance score distribution (such as a block extracted from high-frequency word segments or key areas of conversations). A unified quantization configuration is generated based on the token precision allocation of this processing block. The generated configuration is written to the shared memory area, and the other modules in each shared group access this configuration through the same memory pointer during the inference process to complete the unified quantization operation of weights and activation values. In a multi-GPU deployment environment, the quantization configuration of the group can also be synchronously stored in the shared memory of each GPU, reducing communication latency through pointer or IPC mapping. If a drastic change in the structure of the processing block or a cross-scenario input switch is detected, the system can dynamically update the bound processing block and configuration, and remap the shared memory address to ensure the adaptability of configuration sharing.
[0103] By constructing a configuration sharing group based on processing blocks and reusing the quantization configuration associated with the first network module within the group, it is possible to significantly reduce the generation and switching frequency of quantization parameters during the model inference process, while reducing the repetitive calculation overhead while ensuring semantic fidelity. It strengthens the computational consistency within the model and improves the execution stability and scalability of large model inference in multi-scenario deployments.
[0104] S70, according to the unified quantization configuration corresponding to each processing block, perform block-level batch quantization on all processing blocks and complete the model inference to generate the model inference result.
[0105] In this embodiment, in order to improve the computational efficiency and video memory utilization of the large language model in long text inference tasks, it is necessary to perform block-level batch quantization calculations based on the unified quantization configuration corresponding to each processing block and complete the corresponding model inference process. This operation integrates the capabilities of configuration-driven quantization calculation and processing block parallel scheduling in the technical path, ensuring that each processing block completes refined inference processing according to its importance precision requirements, and finally splicing to generate the complete inference result.
[0106] The unified quantization configuration of each processing block is derived from the previous importance score analysis and precision statistics operations, and usually includes precision format identifiers (such as FP16, INT8, INT4), quantization information of weight parameters (such as scale factor scale_w and zero point zero_point_w), and quantization information of activation values (scale_a and zero_point_a). The execution of quantization calculations usually depends on the selection of efficient computational kernel functions. After the precision configuration of the processing block takes effect, the corresponding execution path needs to be allocated according to the precision format. For example, floating-point calculations use FP16 kernel functions implemented based on CUDA or ROCm, and the fixed-point quantization path calls INT kernel functions while loading relevant quantization parameters.
[0107] To improve the operation efficiency, all processing blocks are mapped to the parallel processing units (SM, Streaming Multiprocessor) on the Graphics Processing Unit (GPU) or Tensor Processing Unit (TPU), and task scheduling is performed according to the starting index or processing order of the processing blocks. During the scheduling process, each processing block loads its own bound unified quantization configuration and executes corresponding operations such as matrix multiplication, normalization, activation function, and residual connection. During the process, the activation value quantization parameter is used to dynamically adjust the input range, and the weight quantization parameter is used to restore the model weights from the fixed-point representation to the computable quantization expression.
[0108] After the inference calculation is completed, it is necessary to perform accuracy verification and invalid position clearing operations on the intermediate results of each processing block. This step is usually completed according to the recorded invalid token position index table, setting the intermediate tensor results at the positions corresponding to the invalid tokens to zero or deleting them to avoid interference with the final output. All valid calculation results are then reconstructed into a complete model output sequence according to the starting position index of the processing blocks.
[0109] Finally, a standard post-processing process is performed on the concatenated output sequence, including dimension alignment (such as padding tensors of different blocks to a unified shape), normalization (such as LayerNorm normalizing the activation output), and decoding (such as generating natural language text through Greedy Decoding or Beam Search), and finally outputting the inference result text or structured vector that can be directly used by the user.
[0110] For example, in a text generation task, the input text is split into a sequence of processing blocks of 512 tokens each. After analysis in steps 4 and 5, each processing block obtains a corresponding unified quantization configuration. For example, processing block A is in high-precision format (FP16), block B is in medium precision (INT8), and block C is in low-precision (INT4). During the inference execution phase, after the GPU loads the unified quantization configuration of each processing block, it allocates processing block A to the high-precision execution channel, and blocks B and C to the fixed-point execution channel, and performs matrix multiplication and non-linear transformation calculations through different kernel functions. After the calculation of each processing block is completed, the system clears the calculation results corresponding to the padding tokens according to the token position index recorded during preprocessing. For example, if block C contains 12 padding tokens, the outputs at the last 12 positions will be set to zero. Then, the valid outputs of blocks A, B, and C are sorted and concatenated according to the starting index, and post-processing (such as normalization and decoding) is performed to obtain the complete text generation result. In a batch inference scenario with multiple inputs, a micro-batch scheduling mechanism can also be combined to schedule the processing blocks of multiple input samples on the same GPU, and share the weight cache of some high-frequency structures when loading the shared configuration to further improve the throughput rate.
[0111] Example illustration: In the field of healthcare business, an intelligent assisted diagnosis platform hopes to deploy a large language model for structured analysis and reasoning of extremely long electronic medical record texts. The goal is to extract the main diagnosis, the evolution process of key symptoms, and the conclusion of etiological inference from tens of thousands of words of hospitalization records. In this task, the length of the original input text exceeds the standard window limit of the model, so the processing block method needs to be used for segmented reasoning. First, the entire medical record text is divided into multiple processing blocks according to a preset length (for example, every 512 tokens). The system marks the first processing block as a high-precision processing area, which usually contains high-semantic-density parts such as the patient's basic information, chief complaint, and current medical history, and directly determines the credibility of subsequent diagnostic reasoning. Therefore, the quantization operation of this processing block is disabled, and only high-precision operators are used for floating-point calculation. For the remaining processing blocks, the platform automatically generates the self-attention matrix corresponding to each processing block before the model execution, and sums each column in the matrix to obtain the degree of attention of each token in the entire context, that is, the importance score. On this basis, the system compares the score with two dynamic thresholds, configures the tokens with higher importance scores as high-precision formats, generally corresponding to core disease course nodes or etiological information, configures the tokens with medium scores as medium-precision formats for processing descriptive paragraphs or disease observation records, and configures the tokens with lower scores such as format information and table data as low-precision formats. After each processing block completes the token-level precision annotation, the system counts the number of tokens in different precision formats and selects the precision level with the largest number as the unified quantization configuration for this processing block. For example, in a certain processing block, the tokens in the medium-precision format account for the highest proportion, then this block is overall configured as a medium-precision calculation path. This strategy avoids frequent precision switching and improves execution efficiency. Then, the system divides the inference module of the model into several configuration sharing groups in sequence, and each group contains several consecutive network modules. The block processed by the first module in each group is bound and configured with precision parameters, and the remaining modules share this configuration, avoiding repeating the precision selection logic for each layer. Finally, all processing blocks call the quantization calculation kernel functions of the corresponding precision to execute the inference in parallel according to their unified quantization configurations. The high-precision blocks use the FP calculation kernel function, and the medium- and low-precision blocks call the fixed-point kernel function, running concurrently in the GPU or TPU parallel units. The positions corresponding to the padding tokens will be automatically skipped during the inference to ensure the validity of the output results. After the inference is completed, the system concatenates the valid results of all processing blocks into a whole inference result text according to the original token position index, and performs unified dimension alignment and normalization, and finally decodes to generate the diagnostic conclusion, the key symptom timeline, and the list of suspected etiologies. The entire inference process ensures the precision and stability of the diagnostic core information while controlling the video memory occupancy, and realizes the deployment ability of high-performance long-text medical understanding tasks.
[0112] In the field of fintech business, a credit assessment platform for small and medium-sized enterprises needs to use large language model inference to judge the credit risk of enterprises based on long text data such as complete business information submitted by enterprises, historical credit contracts, public annual reports, and customer feedback records. Since these documents often have loose structures and complex content, with the total number of tokens easily reaching tens of thousands, directly inputting them into the model will face serious problems such as out-of-memory and long response delays. The system first divides all text information into multiple processing blocks according to a preset length (such as every 1024 tokens), ensuring that the length of each block is controllable and the computing resources are acceptable. The first processing block usually contains high-weight fields such as enterprise entity information, registered capital, core business, and operating income in the past three years, which are the core affecting the credit score. Therefore, the platform strategy is to fix this block to use a high-precision calculation path and disable quantization processing to retain the expression details of key financial features and legal statements to the greatest extent. For the remaining processing blocks, the platform introduces a multi-head self-attention mechanism to analyze the degree of attention received by each token in the context. After constructing the attention matrix, it extracts the sum of the global attention weights received by each token in each column as an importance indicator for the token. Subsequently, the system assigns different precision levels to each token according to two dynamic thresholds, ensuring that high-risk factors such as tax anomalies, contract defaults, and bank credit records are processed in high-precision format, while relatively less important information such as financial statement annotations and customer descriptions uses medium or low precision. Then, the system counts the number of tokens of various precision levels in each processing block and selects the precision level with the highest frequency as the unified quantization configuration for that block, thereby simplifying the execution path and reducing the scheduling overhead. If there is a boundary situation where the numbers of high, medium, and low precision are similar, high precision is preferred to avoid weakening of key semantics in credit judgment. The platform divides the entire large language model into multiple configuration sharing groups at different levels. Each group contains a fixed number of network modules, and the quantization configuration used for the block processed by the first module in the group is used as the shared configuration, which is recorded in the configuration mapping table where the group identifier is bound to the block identifier, ensuring that all modules use the same quantization path when processing the tasks of this group. During the execution phase, the system loads the corresponding unified quantization configuration parameters for each processing block, including the scaling ratio of weights, the calibration range of activation values, and the precision format identifier of the current block. Different blocks are scheduled to the GPU parallel units that support mixed precision. Among them, high-precision blocks call floating-point processing units, and medium- and low-precision blocks use INT4 / INT8 instruction sets to perform inference. At the same time, the intermediate results corresponding to the positions of invalid tokens are removed to avoid introducing redundant interference. Finally, the platform sorts and concatenates all valid inference results according to the original token index to form a complete output sequence, and through the post-processing module, performs structure alignment, risk label normalization, and multi-label score decoding to generate a credit rating, risk exposure range, and warning dimension recommendation list for the enterprise in the current financial scenario, assisting credit review personnel or automated credit engines to complete the decision-making.
[0113] By performing block-level batch quantization calculations based on a unified quantization configuration, not only is a precision-aware heterogeneous execution strategy achieved, but also the parallel computing power of multi-core processors is fully unleashed. While ensuring that different token segments perform differential calculations according to importance, it effectively reduces the overhead of repeated configuration loading, dynamic quantization generation, etc. It reduces the average inference video memory occupancy and improves the processing throughput rate, and is applicable to resource-sensitive large-scale deployment scenarios in long text tasks.
[0114] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health and fintech. It discloses a method for accelerating model quantization inference, including: dividing the input text into multiple processing blocks, fixing the calculation precision format of the first processing block as a high-precision format and disabling quantization operations; for other processing blocks except the first processing block, generating a self-attention matrix through a language model, and calculating the importance score based on the sum of the values in the corresponding column of each token position in the self-attention matrix; allocating each token position to a high-precision, medium-precision, or low-precision format based on two preset thresholds; counting the number of tokens in each precision format within each processing block, and selecting the precision format with the largest number as the unified quantization configuration; dividing the network module into multiple configuration sharing groups, and sharing the unified quantization configuration of the processing block within the group; performing block-level batch quantization according to the unified quantization configuration of the processing block and completing model inference to generate an inference result. The present invention realizes precision allocation and parallel quantization inference at the block level by uniformly determining the quantization configuration of each processing block based on the token importance score and reusing this configuration within the network module group, greatly reducing the video memory overhead and configuration time overhead while ensuring the inference precision, and effectively improving the execution efficiency and video memory utilization rate in long text inference tasks.
[0115] In one embodiment, the above step S10 includes:
[0116] S101, dividing the input text into multiple equally long processing blocks according to a preset block length;
[0117] S102, disabling quantization parameter adjustment for all token positions in the first processing block;
[0118] S103, fixing the processing precision format of the embedding layer, self-attention layer, and feed-forward network layer of the first processing block in the language model as a high-precision format;
[0119] S104, when there are remaining tokens at the end of the input text that are less than the preset block length, forming the text segment composed of the remaining tokens as an independent processing block;
[0120] S105, Pad the independent processing block with invalid tokens to the preset block length, fix the processing precision format of the padded independent processing block to the high-precision format, and disable the quantization operation on the independent processing block;
[0121] S106, Record the start position index and end position index of all processing blocks.
[0122] In this embodiment, to reduce the video memory pressure of the large language model during long text inference, a processing block (Processing Block) partitioning strategy is adopted to slice the input text into multiple sub-sections according to the preset block length. Each processing block is the smallest quantization scheduling unit in model inference and has clear start and end position indexes. This partitioning is not only used to control the locality of video memory usage but also provides boundary support for subsequent quantization strategy generation.
[0123] The processing precision format of the first processing block is fixed to the high-precision format, usually the FP16 or FP32 floating-point format, which means that the model uses the full-precision calculation path when processing this block and explicitly prohibits this block from participating in any form of quantization parameter adjustment, including but not limited to bit-width compression of weight parameters and activation values, dynamic scaling factor generation, etc., thus skipping the process of this block in quantization configuration generation.
[0124] The embedding layer, self-attention layer, and feed-forward network layer together constitute the core module for token representation transformation in the language model and play a key role in the semantic representation of tokens. Using the high-precision format for these core structures of the first processing block helps to stabilize the context awareness ability of the model at the initial stage of inference and improve the nested inference performance of subsequent tokens, especially in scenarios where key topic cues are included at the initial stage of long text input.
[0125] If the text length cannot be evenly divided by the preset block length, the remaining tokens at the end will be assembled into an independent processing block. To maintain the consistency of the block-level inference structure and the alignment requirements for parallel scheduling, this independent processing block will be padded with invalid tokens (Paddingtoken) to complete the full block length and will also be processed using the high-precision format. The padding tokens do not participate in the effective calculation of the model output during actual inference and are only used to keep the input dimensions consistent.
[0126] After the system finishes text chunking, it needs to record the start and end position index information of each processing chunk, which serves as a crucial auxiliary identifier for subsequent steps such as computational scheduling, output splicing, and precision allocation. This index information can be constructed into a bidirectional mapping table (such as a hash table) to quickly locate the processing chunk corresponding to a token and its precision control strategy during runtime.
[0127] The input text can be sequentially partitioned by setting the chunk length as a subset of the maximum sequence length supported by the model (such as 512 or 1024 tokens). The partitioning operation can be completed offline during the preprocessing stage or dynamically based on streaming input during runtime. In the GPU inference framework, position encoding masks can be used to distinguish actual tokens from padding tokens, and high-precision kernel functions can be used to process the specified chunks.
[0128] In the current mainstream quantization inference hardware systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), and TPUs (Tensor Processing Units) are widely used in the deployment of high-performance large language models. For these platforms, flexible precision scheduling capabilities need to be provided during actual deployment to accommodate the high-precision processing requirements of the first processing chunk.
[0129] In programmable chip architectures (such as FPGAs or AI acceleration chips), generally, each operation in the model execution path is mapped to a series of reconfigurable logic units or operator chains. To achieve high-precision execution of the first processing chunk, the computational precision mode of the current chunk can be set through a bit-width control register or an operator scheduling instruction table on the intermediate computational path. When it is recognized that the token position being processed belongs to the first processing chunk, the system controller can, through a combination of hardware and software, set the precision control field of the corresponding operator chain to a high-precision flag, such as enabling the FP16 or FP32 channels, and bypass the quantization scaling unit, skipping operations such as zero-point adjustment and discrete mapping, so that the data is directly transmitted to the operator input port in its original floating-point format, implementing the so-called "precision passthrough path".
[0130] The key to this mechanism lies in binding the location information of processing blocks with the register configuration strategy. By maintaining a processing block - precision strategy mapping table in the scheduling unit, the goal of dynamically switching precision paths at the block - level granularity can be achieved. When processing streaming token inputs, this mechanism can update the runtime configuration through the boundary detection module.
[0131] In highly integrated heterogeneous computing platforms such as TPU, it usually contains multiple processing cores (Cores) that support different computing precisions. Some of these cores are high - computing - power main cores (High - Precision Cores) that support floating - point computing, while other cores are optimized for dedicated fixed - point computing paths. The TPU architecture often integrates a scheduler and a resource allocator to dynamically schedule workloads among different Cores.
[0132] When processing the first processing block, the scheduler can identify it as a high - priority computing unit and directly allocate its tasks to the main core, using the floating - point computing core to execute the calculations of the Embedding, self - attention, and feed - forward network modules, thus avoiding the introduction of the quantization process; while the inference tasks of the remaining processing blocks can be divided into other low - bit - width Cores and executed in a low - precision high - throughput mode for quantized inference. To achieve dynamic resource reuse, the scheduling strategy of TPU can also support time slicing and task prefetching to ensure low latency and high parallelism in the scheduling between multiple processing blocks.
[0133] If deployed on a general - purpose GPU platform, the high - precision execution flag can also be passed through CUDA kernel launch parameters to disable the quantization process in a specific kernel path. For example, launch a group of kernels that do not contain quantization kernel functions for the first processing block, or force the scheduling of the float path instead of the int path through the precision tag injection mechanism in the general - purpose model framework.
[0134] In addition, in a higher - order system, a dynamic fine - tuning logic (precision adaptation controller) can also be introduced to automatically determine whether to enable the high - precision path based on the importance score distribution or semantic weight of the first processing block, thus achieving a flexible balance between performance and precision.
[0135] In this embodiment, by dividing the input text into multiple equally-sized processing blocks and disabling quantization processing for the first processing block, more original semantic and context information can be retained in the initial stage of model inference, thereby improving the stability of subsequent token importance evaluation and precision allocation. This helps to improve the starting inference precision of long text input without significantly increasing the overall video memory overhead, providing a more reliable baseline representation for subsequent chunk quantization configuration.
[0136] In one embodiment, the above step S20 includes:
[0137] S201, input each of the other processing blocks into the multi-head self-attention layer of the language model to obtain the local self-attention matrices of multiple attention heads corresponding to each of the other processing blocks;
[0138] S202, perform weighted average or arithmetic average processing on the local self-attention matrices of multiple attention heads to generate the final self-attention matrix corresponding to each of the other processing blocks;
[0139] S203, extract the column vectors corresponding to each token position from the final self-attention matrix;
[0140] S204, determine the sum of the values of all elements in each column vector, and use the sum as the importance score of the token position corresponding to each column vector;
[0141] S205, if there are filled invalid token positions in the other processing blocks, when determining the importance score of the filled invalid token positions, set all the element values of the column vector corresponding to the invalid token position to zero.
[0142] In this embodiment, before precision allocation for the processing blocks, it is necessary to first obtain the importance of each token in the current semantic context. To achieve this goal, the processing flow starts from the multi-head self-attention layer and constructs an attention matrix for each processing block. After each processing block is input into the language model, it passes through multiple attention heads, and each attention head independently calculates a local self-attention matrix. Each matrix takes all the tokens in the processing block as units and describes the mutual attention relationship between them. These matrices reflect the language model's ability to model the internal structure of the sentence from different attention perspectives.
[0143] To obtain a unified and comparable representation of attention, these local matrices need to be fused. The fusion method can be flexibly set according to the actual task requirements. Common methods include taking the arithmetic mean of the matrices of all attention heads, that is, simply calculating the mean; or weighted average can be achieved by configuring different weight coefficients for each attention head. This fusion process generates the final self-attention matrix for each processing block, which serves as a comprehensive representation of the interdependence of tokens within the processing block.
[0144] In this final matrix, each column vector represents the degree of attention focus that a specific token receives from other tokens. Specifically, fixing the position of a certain token and extracting the column vector corresponding to that position can be regarded as "the set of degrees to which this token is concerned by all other tokens". This is an important basis for evaluating whether a token occupies an important semantic position in the current processing block. That is:
[0145] Current token: The token for which the attention output is being calculated, denoted as the token at the i-th position.
[0146] Other tokens: All valid tokens in the same processing block except the token at the i-th position, that is, the position numbers within the block are from 0 to N - 1 (where N is the length of the processing block); excluding the i-th position itself; excluding the positions of invalid tokens (such as padding tokens).
[0147] To extract this importance signal, the sum of all elements in each column vector is calculated to obtain an aggregated value for measuring the global attention received by this token. This aggregated value is the importance score of this token. The higher the score, the greater the impact of this token on the overall context semantic structure. In the subsequent precision allocation process, this score will be directly used to determine whether this token should be processed using higher-precision computing resources.
[0148] Considering that the processing block may contain invalid tokens padded for the filling length, these tokens do not carry semantic content. If they participate in the calculation of the attention matrix, they may interfere with the importance score. To prevent such interference, after generating the final self-attention matrix, it is necessary to identify the positions corresponding to all invalid tokens and set their column vectors in the matrix to zero uniformly. This can effectively avoid invalid tokens obtaining false high attention values and ensure the accuracy and robustness of importance calculation. This operation can be achieved through a structured mask or an explicit indexing mechanism, and the specific method is adjusted according to different model architectures and deployment platforms.
[0149] In actual deployment, when a processing block is fed into the language model for inference, the self-attention mechanism of each layer of the model will output attention matrices corresponding to a number of attention heads. These matrices are usually organized in a three-dimensional structure, where the first dimension is the number of attention heads, and the second and third dimensions respectively represent the number of tokens in the processing block, forming a three-dimensional tensor with a structure of "number of attention heads × number of tokens × number of tokens".
[0150] During the forward inference process of the model, through the intermediate layer interception mechanism, these attention tensors can be cached at the output of the self-attention module of each layer. To reduce the interruption overhead during inference, the caching operation should be inserted into the computational graph of the model in a non-blocking manner and its lifecycle should be uniformly managed by the scheduler. After caching, different fusion strategies can be achieved by operating on the first dimension (i.e., the attention head dimension): one way is to simply average the matrices of all attention heads to obtain a two-dimensional matrix; another way is to assign different weighting coefficients to each head and perform weighted average processing, which is applicable when the differences in the semantic perception capabilities of different attention heads have been evaluated in advance and strategic adjustments are made in the weight setting to enhance the expressive power of the fusion matrix.
[0151] After completing the fusion of attention heads, a two-dimensional matrix of number of tokens × number of tokens will be obtained, representing the comprehensive attention relationship between tokens within the processing block. The global attention degree of each token is determined by the sum of the elements in the corresponding column of this matrix. In specific operations, by fixing a certain column index, the entire matrix can be sliced to extract the column vector, and this operation can be implemented through standard matrix slicing functions in the tensor operation framework (such as PyTorch, TensorFlow, JAX, etc. all support such slicing access). Subsequently, all the values in this column vector are accumulated to obtain the attention aggregation value of the token at this position, that is, its importance score. Since vector addition is highly optimized in modern tensor computing libraries, the calculation of the importance score can be completed with extremely low latency on the GPU / TPU, with good engineering feasibility.
[0152] Regarding the processing of padding tokens, during forward inference, the model usually generates a padding mask, which is used to mark which tokens are invalid positions generated by padding. The mask is generally a boolean vector or a 0 / 1 tensor of the same length as the processing block, where 1 represents a valid position and 0 represents a padding position. This mask can be extended to a matrix form and aligned with the column structure of the final self-attention matrix. In implementation, this mask can be broadcast into a two-dimensional matrix and applied to each column of the attention matrix to uniformly set all elements of the column vector corresponding to the padding position to zero at the matrix level.
[0153] To achieve optimal efficiency, the mask application process can be embedded in the quantization preprocessing module during the model graph construction phase and attached as a post-operation after the attention matrix calculation node. Depending on the running platform, the column-level zeroing logic can be implemented through CUDA kernels on GPUs; on TPUs, the built-in high-throughput tensor mapping operation can be relied on to complete batch column masking within a single clock cycle.
[0154] In addition, to ensure that special structures appearing in the context scenario do not cause misjudgment, such as in some tasks, the padding tokens may be in the middle of the sequence rather than at the end (such as multi-segment concatenated input), it is necessary to record the original input structure and the padding mask generation logic in advance to ensure the accuracy of the mask. An optional method is to introduce a position label assistance mechanism, which jointly uses the position semantics encoding of each token and the mask to enhance the recognition stability of invalid tokens.
[0155] In this embodiment, by extracting column vectors from the fused self-attention matrix and summing the total weights, the importance scores of tokens are obtained, and a quantization accuracy perception mechanism starting from the language model structure is constructed. Without introducing an external scoring module, the attention information already learned by the model itself is fully utilized to internalize the basis for accuracy allocation. It avoids explicit label dependence and has good computational graph compatibility, which helps with subsequent efficient deployment on acceleration chips.
[0156] In one embodiment, the above step S30 includes:
[0157] S301, determining a first threshold and a second threshold according to the length of the processing block;
[0158] S302, if there are invalid token positions, assigning the processing precision format of the invalid token positions to a low-precision format and excluding the invalid token positions when counting the number of precision formats of each processing block;
[0159] S303, comparing the importance score of each valid token position with the first threshold and the second threshold;
[0160] S304, assigning the processing precision format of the valid token positions with importance scores greater than the first threshold to a high-precision format;
[0161] S305, assigning the processing precision format of the valid token positions with importance scores above the second threshold and below the first threshold to a medium-precision format;
[0162] S306, assigning the processing precision format of the valid token positions with importance scores less than the second threshold to a low-precision format.
[0163] In this embodiment, allocating the token positions with importance scores greater than the first threshold to the high-precision format is an important operation for differentially configuring the processing precision after evaluating the attention information of all valid tokens in each processing block. The high-precision format refers to using a higher-precision representation and calculation method during the inference process, such as using a numerical representation with a higher bit width or floating-point operations that disable quantization operations. Its purpose is to retain the expressive ability of key semantic tokens and avoid negative impacts on the model prediction results caused by precision degradation. The medium-precision format and the low-precision format correspond to the medium and low bit-width quantization configurations in resource-constrained scenarios and are used to process tokens with lower importance to save video memory and computing resources.
[0164] The length of the processing block has a dynamic impact on the setting of the threshold. Since the number of tokens in a long processing block increases and its overall information density distribution tends to be average, in order to ensure that the screening mechanism has relatively consistent discrimination ability in processing blocks of different lengths, the values of the first threshold and the second threshold need to be correspondingly lowered to expand the determination range of high-precision tokens. Dynamic adaptation can be achieved by setting an initial standard threshold and linearly or non-linearly scaling the two thresholds in combination with the ratio of the current block length to the standard length.
[0165] In the scenario where there are invalid token positions in the processing block, special judgment needs to be added to the processing precision allocation. Since invalid tokens do not participate in semantic modeling, there is no semantic contribution in their attention information itself, and their corresponding importance scores can be considered as structural null values or forced to zero. Allocating them uniformly to the low-precision format can save computing resources and actively exclude them when counting the precision ratio later to avoid interfering with the unified quantization configuration decision of the current processing block.
[0166] The importance score of each valid token will be compared with the two thresholds in sequence. When its value is higher than the first threshold, it means that it plays a significant semantic connection role in the global attention aggregation and needs to be allocated the high-precision format; when its score is between the first and second thresholds, it is considered to be in the sub-important information interval and can be processed using the medium-precision format; when it is lower than the second threshold, it indicates that the token has a weak role in context construction and can use the low-precision format to perform inference calculations. This multi-level precision configuration scheme ensures the balance of semantic expression, performance control, and computing efficiency.
[0167] If there are no invalid token positions in the processing block, all token positions are valid token positions.
[0168] In a specific implementation, when initializing the model for running, by reading the length of the processing block, a preset dynamic threshold calculation function can be called to generate the first threshold and the second threshold for the current block. This function can be implemented by means of table lookup or formula calculation, and the specific implementation can support mechanisms such as interval interpolation, exponential decay, and minimum threshold protection. Subsequently, all token positions of the current processing block are traversed in sequence. During the traversal process, the system calls the padding mask or the token validity label to determine whether the current token is an invalid token. If it is an invalid token, its processing precision format is directly marked as low precision, and the importance score comparison logic is skipped, and the process proceeds to the next token judgment. For valid tokens, the importance score of the token is read from the cache or the current attention module. This score is then compared with the two thresholds corresponding to the current processing block. If the score is greater than the first threshold, it is marked as high precision in the processing block precision allocation vector; if the score is greater than the second threshold and less than or equal to the first threshold, it is marked as medium precision; if the score is less than or equal to the second threshold, it is marked as low precision. After processing, the precision labels of all valid tokens are saved in the precision control vector inside the processing block and are used for subsequent unified quantization configuration decisions and kernel function selection. If acceleration of processing is required, the importance score comparison logic can be vectorized, and the judgment of the allocation strategy can be completed in parallel on the GPU to improve the execution efficiency.
[0169] Through the above steps, in this embodiment, without significantly increasing the model complexity, the calculation precision of tokens can be differentially configured according to the actual semantic importance of the tokens, thereby greatly reducing the overall video memory consumption and calculation load while ensuring the inference precision of the model. The dynamic threshold mechanism further improves the adaptation ability of the algorithm when processing input sequences of different lengths, and avoids the problem of uneven precision allocation caused by fixed thresholds. The special processing of invalid tokens ensures the accuracy of resource allocation and simplifies the subsequent statistical logic.
[0170] In one embodiment, the above step S40 includes:
[0171] S401, traverse all token positions of the current processing block to identify the token positions assigned as high-precision format, medium-precision format, and low-precision format;
[0172] S402, during the statistical process, if there are invalid token positions, exclude all invalid token positions and only count the precision format allocation results of valid token positions;
[0173] S403. Respectively count the numbers of tokens with positions assigned to high-precision format, medium-precision format, and low-precision format among the valid token positions, and generate a high-precision count, a medium-precision count, and a low-precision count.
[0174] S404. Compare the numerical values of the high-precision count, the medium-precision count, and the low-precision count.
[0175] S405. Use the precision format corresponding to the count with the largest numerical value as the unified quantization configuration for the current processing block.
[0176] S406. If there are multiple counts with the same precision format and they are the maximum values, select the highest precision format among the multiple precision formats as the unified quantization configuration for the current processing block.
[0177] In this embodiment, counting the number of token positions assigned to different precision formats within each processing block is a key analysis operation performed after the token precision assignment to determine the overall quantization strategy for the current processing block. This process first requires traversing all token positions in the current processing block to identify the precision label corresponding to each position, which may be in high-precision, medium-precision, or low-precision format.
[0178] During the traversal, if the processing block contains invalid tokens, such as padding tokens filled at the end, they need to be excluded from the statistics. This is because invalid tokens do not participate in the actual inference calculation and should not affect the overall precision distribution. Only counting the precision labels of valid tokens can ensure the rationality and accuracy of subsequent unified quantization configuration decisions.
[0179] Counting each precision label is the core of this step. The system will generate three counting results respectively: high-precision count, medium-precision count, and low-precision count. Each counting result corresponds to the number of positions where valid tokens in the processing block are assigned to this precision format, and is used to depict the precision distribution trend within the block.
[0180] Subsequently, the system compares the sizes of the three counting results, finds the precision format with the largest quantity, and uses it as the unified quantization configuration for the current processing block. That is, the current processing block will be uniformly quantized and the kernel function will be called in this format subsequently. The design purpose of the unified configuration mechanism is to avoid frequent switching of precision calculation paths within the same processing block, thereby reducing the context switching cost and improving the inference throughput efficiency.
[0181] If the results of multiple precision counts are the same and are the maximum value, such as when the high-precision and medium-precision counts are exactly the same, the system should use the precision priority principle and select the one with a higher precision level among the tied maximum values as the final configuration. This strategy reflects the tendency to protect semantic quality, that is, to preferentially retain higher computational precision within the scope allowed by resources, thereby reducing the risk of token misrecognition or inference deviation.
[0182] For example, during the execution of the model, after the token precision allocation for each processing block is completed, the system will start the precision statistics module. First, it will traverse all token index positions in this processing block through loop or vectorization logic. The precision label corresponding to each token is stored in the precision marking vector of the previous step. Each element of the vector is the precision identifier of the token, for example, 0 represents low precision, 1 represents medium precision, and 2 represents high precision.
[0183] During the traversal process, the system checks whether each token is a valid token. The judgment basis can be the padding mask or the token validity vector. If it is an invalid token, its corresponding position will be skipped during the precision statistics process and will not participate in any precision counting.
[0184] For valid tokens, they will be added to the corresponding precision counter according to their precision identifiers. The system also maintains three independent integer variables to accumulate the number of tokens with high, medium, and low precision respectively. After the traversal is completed, the system compares the values of the three counters to find the precision format corresponding to the maximum value.
[0185] If the counts of two or three precisions are both the maximum value, for example, both high precision and medium precision are 100, the system performs a priority judgment. It can compare through the precision level values and preferentially select the precision format with a higher level. The system can also expand and introduce additional factors in this logic, such as the inertia of historical block configuration, the trend of context semantic density, etc., to enhance the stability and adaptability of the precision configuration.
[0186] The finally determined precision format will be written into the processing block control structure as the unified quantization configuration of the current processing block and passed to the subsequent quantization kernel function selection logic. This configuration can be applied in batch form on the GPU to improve the efficiency of the quantization preparation stage.
[0187] In this embodiment, by comprehensively counting and analyzing the precision tags inside each processing block, a unified quantization configuration that best suits its semantic density and precision requirements can be assigned to each block in a data-driven manner, significantly reducing the computational path complexity and precision switching overhead. While retaining the high-precision processing ability of key tokens, the inference efficiency is also improved and the video memory is compressed through the block-level merging strategy. The exclusion process of invalid tokens ensures that the statistical results are not interfered by redundant tokens, further improving the configuration accuracy and execution stability of the system. The introduction of the precision-first strategy enables the system to still guarantee the processing quality of high-semantic contribution tokens when facing scenarios with fuzzy precision distributions, which helps to improve the stability and reliability of the model output.
[0188] In one embodiment, the above step S60 includes:
[0189] S601, assign a unique block identifier to each processing block and a unique group identifier to each configuration sharing group;
[0190] S602, establish a binding relationship between the unique group identifier and the unique block identifier in the configuration mapping table;
[0191] S603, based on the binding relationship between the unique group identifier and the unique block identifier, associate the first network module of each configuration sharing group with the unified quantization configuration parameters of the corresponding processing block;
[0192] S604, write the unified quantization configuration parameters into the shared memory area of each configuration sharing group, and assign the same memory address mapping to all network modules of each configuration sharing group;
[0193] S605, when other network modules within the same configuration sharing group perform quantization processing, read the unified quantization configuration parameters from the shared memory area through the memory address mapping.
[0194] In this embodiment, in order to achieve compression optimization of video memory resources and computational paths during the inference process of large language models (LLMs), a configuration sharing mechanism is introduced to reuse the unified quantization configuration among multiple network modules within the same group. The key to this mechanism lies in the structured organization of the mapping relationship between processing blocks and network modules, and the reuse of parameter access paths using shared memory.
[0195] First, the system will generate a globally unique block identifier for each processing block during the processing block division stage, which is used to identify the group of token sequences and their quantization configurations corresponding to the block. In different implementation platforms, this identifier can be an integer index, a hash result, or a combined encoding that combines the processing block position and length, with stability and traceability, facilitating quick positioning of the quantization configuration ownership in subsequent steps.
[0196] Meanwhile, the system also divides the network modules into several configuration sharing groups according to their processing order in the model inference graph. A configuration sharing group is a logical structure that contains multiple adjacent or semantically related network modules, usually set to a fixed number (such as 4, 8, or 16 modules per group). Each configuration sharing group is also assigned a unique group identifier, and its generation mechanism should maintain the same uniqueness and mapability as the processing block identifier.
[0197] Based on the establishment of these two types of identifiers, the system forms a binding relationship between the group identifier and the block identifier by constructing a configuration mapping table. This mapping table can be in the form of key-value pairs, with the group identifier as the key and the corresponding block identifier as the value. This structure allows the model to quickly find the processing block and its quantization configuration bound to the current configuration sharing group during runtime, thereby determining the shared access path.
[0198] Furthermore, to achieve the reuse of quantization configuration among multiple modules within a group, the system writes the unified quantization configuration parameters of the processing block into a shared memory area after the binding is completed. This shared area can be a shared cache (shared memory) in the GPU memory at the implementation level, or a tensor accelerator register area that supports concurrent access, or a high-bandwidth cache block allocated by middleware in a system that supports heterogeneous architectures.
[0199] The unified quantization configuration parameters usually include weight quantization parameters, activation value quantization parameters, processing precision format information, bit-width encoding strategy, quantization range information, and auxiliary biases. These parameters are stored in a structured format, such as a key-value dictionary, a compressed tensor, or a configuration structure, and are exposed to all network modules within the configuration sharing group through the shared memory address.
[0200] In terms of address mapping, the system points the quantization configuration access paths of all modules within the configuration sharing group to a unified shared memory address. This mapping process can be statically completed when the model is loaded, or dynamically maintained during the inference process, depending on the time coupling relationship between the processing block generation and the module calls within the group. It should be noted that to ensure the performance and consistency of multi-module concurrent reading, the shared memory area should be ensured to have thread safety and non-blocking characteristics, which can be achieved through lock mechanisms, read-write caches, or communication channels.
[0201] Once the mapping is completed, when the remaining network modules within the configuration sharing group perform quantization processing, they will no longer independently generate or load quantization configurations, but directly read the required parameter data from the shared area by looking up the mapped address. This method avoids repeated calculations, memory copying, and configuration switching, greatly compresses the execution graph width of the inference path, and improves the overall computational density and throughput.
[0202] Furthermore, if this mechanism is combined with quantization-aware training (QAT) or post-training quantization (PTQ) strategies, configuration sharing not only reduces the memory consumption at the hardware layer, but also improves semantic consistency, allowing modules in the same group to collaboratively express local semantics under consistent quantization configuration, effectively avoiding representation jitter or semantic drift caused by differences in quantization strategies.
[0203] In addition, the configuration sharing mechanism has good scalability and hardware compatibility. In the framework that supports graph execution (such as TensorRT, ONNX Runtime), it can be implemented through a unified layer parameter injection strategy, and in the low-level hardware layer (such as TPU, NPU), dynamic sharing can be achieved through the memory address binding table and interrupt synchronization mechanism.
[0204] This embodiment uses a configuration sharing mechanism to achieve efficient reuse of processing block quantization configurations within a configuration sharing group, significantly reducing video memory consumption and computational burden caused by repeated configuration loading. Compared with the traditional method of loading configurations for each module independently, the use of a unified memory address mapping can reduce configuration access latency to a constant level and reduce the overall peak memory usage of the system. The unified quantization strategy within the group also improves the consistency of multi-module collaborative execution, effectively reduces quantization jitter in the reasoning path, and enhances model output stability. In the scenario of multi-block concurrent execution, this mechanism can effectively avoid semantic drift or output fault problems caused by inconsistent quantization strategies.
[0205] In one embodiment, the above step S70 includes:
[0206] S701, loading corresponding unified quantization configuration parameters for each processing block, wherein the unified quantization configuration parameters include a precision format identifier;
[0207] S702, configuring independent processing kernel functions for different processing blocks according to the precision format identifier;
[0208] S703, in a parallel processing unit of a graphics processor or a tensor processor, allocating processing resources according to the index order of the processing blocks, and concurrently executing quantization processing of all processing blocks;
[0209] S704, performing token position verification on the processing result of each processing block to eliminate the intermediate results corresponding to invalid token positions;
[0210] S705, sorting the valid intermediate results of all processing blocks according to the starting position index and splicing them into a complete model output sequence;
[0211] S706, performing post-processing operations on the spliced output sequence to generate a final model inference result.
[0212] In this embodiment, loading unified quantization configuration parameters for each processing block is the preparation stage before model execution. The unified quantization configuration parameters include three key contents. One is the weight quantization parameters, which are a set of parameters used to map the model weights from floating-point representation to quantized representation, usually including scaling factors and zeros, and these parameters are widely used in both static quantization and dynamic quantization processes. The second is the activation value quantization parameters, which control the compression accuracy of the input activation values and are used to maintain the stability of the input feature distribution during inference. The third is the precision format identifier, which is a marker guiding the processing block to select the corresponding calculation path, usually an enumeration type or a control register flag, and is used for subsequent dynamic kernel function scheduling. The specific implementation of loading these parameters can be achieved through configuration table lookup, memory prefetching, or configuration caching to reduce the memory access latency during the inference stage.
[0213] Configuring independent processing kernel functions according to the precision format identifier is the key link to achieve block-level heterogeneous computing. The processing kernel functions are execution paths customized for different precision formats. High-precision formats usually correspond to floating-point kernel functions of FP16 or FP32, retaining higher numerical precision to process semantically dense or position-critical tokens. Medium and low-precision formats correspond to fixed-point quantization kernel functions such as INT8 and INT4, which use mechanisms such as tensor quantization and weight sharing to accelerate matrix calculations and memory transfers. The scheduling strategy of the computing kernel functions can be implemented through a function pointer mapping table or a kernel selection mechanism supported by the hardware, enabling the processing block execution paths to be allocated on demand and improving the overall throughput.
[0214] In a Graphics Processing Unit (GPU) or a Tensor Processing Unit (TPU), allocating parallel processing resources for all processing blocks is the physical guarantee for achieving block-level batch quantization. Allocating resources based on the index order of the processing blocks can ensure that the output order of the processing results is consistent with the input order, while avoiding resource competition between threads. In the GPU platform, the distribution of the processing kernel functions can be controlled through the CUDA block and stream mechanisms, while in the TPU, each processing block execution can be mapped through the Matrix Multiply Unit (MXU). During this process, the weight quantization parameters and the activation value quantization parameters are synchronously loaded into the corresponding execution contexts and are respectively applied to the quantization mapping of the weight tensors and the dynamic calibration operation of the input tensors in the model.
[0215] The intermediate results after processing need to be checked for the token positions within the block. Since some processing blocks contain invalid tokens filled with padding, the calculation results at these positions should be filtered to prevent affecting the subsequent model output. In this step, the token mask recorded in the previous step can be used to perform a screening operation on the intermediate results by position, removing the results corresponding to the invalid tokens and maintaining their position consistency during the splicing process.
[0216] When stitching together the valid intermediate results of all processing blocks, they need to be sorted and concatenated according to the starting position index of the original input to ensure that the semantic information is restored in the correct contextual order. This operation not only requires data reorganization based on token indices but also demands consistency in the output formats of different processing blocks, including tensor dimensions, sequence alignment methods, etc., to avoid dimension misalignment or information overlap during the stitching process.
[0217] Finally, the stitched sequence is input into the post-processing module to perform dimension alignment, normalization, and decoding operations. Dimension alignment is used to standardize the tensor shapes of each segment's output, normalization is used to compress the numerical dynamic range and improve decoding accuracy, and decoding generates the final inference results such as text, labels, or feature values according to the model task type. The entire post-processing flow can be embedded as a unified output layer at the back end of the inference graph to ensure efficient execution.
[0218] In some implementations, the weight quantization parameters and activation value quantization parameters can be obtained through external quantization training, saved in an offline configuration file, and pre-loaded before the start of inference. Another approach is runtime dynamic quantization, which calculates the scaling ratio in real time according to the maximum and minimum value ranges of the current input activation values. For the assignment of precision format identifiers, they can be represented as integer encodings during the quantization configuration generation step. For example, 0 represents low precision, 1 represents medium precision, and 2 represents high precision, for the processing kernel function selector to parse.
[0219] The allocation of processing kernel functions can adopt an adaptive kernel call strategy supported by the hardware. For example, in the NVIDIA TensorRT framework, different GEMM implementation modules are automatically matched according to the precision identifier, and on the TPU platform, the precision configuration can be compiled into the corresponding hardware instruction sequence through the XLA compiler. The allocation of processing resources can also be optimized according to the hardware architecture. Some platforms support preferentially scheduling high-precision blocks to the core resource area with stronger computing performance to ensure the accuracy of model inference.
[0220] In actual deployment, to reduce the bandwidth overhead during the stitching process, the valid results of each processing block can be immediately written into a pre-allocated global output buffer after the output of each block, arranged according to the token index position, to avoid subsequent sorting operations. During the dimension alignment process, for the differences in output tensors caused by different precision formats, a unified mapping strategy or padding filling mechanism can be introduced to ensure consistent data structures.
[0221] In this embodiment, by loading the corresponding unified quantization configuration for each processing block and dynamically selecting the processing kernel function accordingly, the collaborative optimization of computing precision and resource allocation is achieved. By introducing a precision format identifier to control the kernel function path, while ensuring the processing precision of high-importance regions, the computing resources of low-importance regions are maximally compressed. The parallel processing unit executes by blocks, further improving the execution efficiency, and the elimination of invalid token positions ensures the precision and consistency of the splicing result. After the final output undergoes unified post-processing operations, it can not only maintain the context semantic coherence but also effectively reduce the overall inference latency and video memory consumption.
[0222] In one embodiment, a model quantization inference acceleration device is provided, and this model quantization inference acceleration device corresponds one-to-one with the model quantization inference acceleration method in the above embodiment. Referring to Figure 3 , Figure 3 is a schematic diagram of the functional modules of a preferred embodiment of the model quantization inference acceleration device of the present invention. An input text preprocessing module 10, a self-attention analysis module 20, a precision allocation module 30, a quantization configuration decision module 40, a network module grouping control module 50, a configuration sharing management module 60, and an inference execution module 70. The detailed description of each functional module is as follows:
[0223] The input text preprocessing module 10 is used to divide the input text into multiple processing blocks, fix the processing precision format of the first processing block as a high-precision format, and disable the quantization processing of the first processing block;
[0224] The self-attention analysis module 20 is used to generate a self-attention matrix for each of the other processing blocks except the first processing block among the multiple processing blocks through a language model, determine the sum of all element values in the column corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score of each token position;
[0225] The precision allocation module 30 is used to allocate the token positions with importance scores greater than the first threshold to the high-precision format, allocate the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and allocate the token positions with importance scores less than the second threshold to the low-precision format;
[0226] The quantization configuration decision module 40 is used to count the number of token positions allocated to the high-precision format, medium-precision format, and low-precision format in each processing block, and select the precision format with the largest number as the unified quantization configuration of the corresponding processing block;
[0227] The network module grouping control module 50 is used to divide the network modules of the language model into multiple configuration sharing groups, and each configuration sharing group includes at least two network modules;
[0228] The configuration sharing management module 60 is used to share the unified quantization configuration of the processing block corresponding to the first network module within each configuration sharing group to other network modules within the same configuration sharing group;
[0229] The inference execution module 70 is used to perform block-level batch quantization on all processing blocks according to the unified quantization configuration corresponding to each processing block and complete model inference to generate a model inference result.
[0230] In one embodiment, the input text preprocessing module 10 is specifically used for:
[0231] Dividing the input text into multiple equally long processing blocks according to a preset block length;
[0232] Disabling the adjustment of quantization parameters for all token positions of the first processing block;
[0233] Fixing the processing precision format of the embedding layer, self-attention layer, and feed-forward network layer of the first processing block in the language model to a high-precision format;
[0234] When there are remaining tokens at the end of the input text that are less than the preset block length, forming the text segment composed of the remaining tokens into an independent processing block;
[0235] Padding the independent processing block with invalid tokens to the preset block length, fixing the processing precision format of the padded independent processing block to a high-precision format, and disabling the quantization operation on the independent processing block;
[0236] Recording the start position index and end position index of all processing blocks.
[0237] In one embodiment, the self-attention analysis module 20 is specifically used for:
[0238] Inputting each other processing block into the multi-head self-attention layer of the language model to obtain the local self-attention matrices of multiple attention heads corresponding to each other processing block;
[0239] Performing weighted average or arithmetic average processing on the local self-attention matrices of multiple attention heads to generate the final self-attention matrix corresponding to each other processing block;
[0240] Extracting the column vector corresponding to each token position from the final self-attention matrix;
[0241] Determine the sum of the values of all elements in each column vector, and use the sum of the values as the importance score of the token position corresponding to each column vector;
[0242] If there are filled invalid token positions in other processing blocks, when determining the importance score of the filled invalid token positions, set all element values of the column vector corresponding to the invalid token positions to zero.
[0243] In one embodiment, the precision allocation module 30 is specifically configured to:
[0244] Determine a first threshold and a second threshold according to the length of the processing block;
[0245] If there are invalid token positions, allocate the processing precision format of the invalid token positions to the low-precision format, and exclude the invalid token positions when counting the number of precision formats of each processing block;
[0246] Compare the importance score of each valid token position with the first threshold and the second threshold;
[0247] Allocate the processing precision format of the valid token positions with importance scores greater than the first threshold to the high-precision format;
[0248] Allocate the processing precision format of the valid token positions with importance scores above the second threshold and below the first threshold to the medium-precision format;
[0249] Allocate the processing precision format of the valid token positions with importance scores less than the second threshold to the low-precision format.
[0250] In one embodiment, the quantization configuration decision module 40 is specifically configured to:
[0251] Traverse all token positions of the current processing block to identify the token positions allocated to the high-precision format, medium-precision format, and low-precision format;
[0252] During the statistical process, if there are invalid token positions, exclude all invalid token positions and only count the precision format allocation results of the valid token positions;
[0253] Respectively count the counts of the valid token positions allocated to the high-precision format, medium-precision format, and low-precision format to generate a high-precision count, a medium-precision count, and a low-precision count;
[0254] Compare the numerical sizes of the high-precision count, medium-precision count, and low-precision count;
[0255] Use the precision format corresponding to the count with the largest value as the unified quantization configuration for the current processing block;
[0256] If there are multiple counts with the same precision format and they are the maximum values, select the highest precision format among the multiple precision formats as the unified quantization configuration for the current processing block.
[0257] In one embodiment, the configuration sharing management module 60 is specifically configured to:
[0258] Allocate a unique block identifier for each processing block and a unique group identifier for each configuration sharing group;
[0259] Establish a binding relationship between the unique group identifier and the unique block identifier in the configuration mapping table;
[0260] Based on the binding relationship between the unique group identifier and the unique block identifier, associate the first network module of each configuration sharing group with the unified quantization configuration parameters of the corresponding processing block;
[0261] Write the unified quantization configuration parameters into the shared memory area of each configuration sharing group, and allocate the same memory address mapping for all network modules of each configuration sharing group;
[0262] When other network modules within the same configuration sharing group perform quantization processing, read the unified quantization configuration parameters from the shared memory area through the memory address mapping.
[0263] In one embodiment, the inference execution module 70 is specifically configured to:
[0264] Load the corresponding unified quantization configuration parameters for each processing block, and the unified quantization configuration parameters include precision format identifiers;
[0265] Configure independent processing kernel functions for different processing blocks according to the precision format identifiers;
[0266] In the parallel processing unit of the graphics processor or tensor processor, allocate processing resources according to the index order of the processing blocks, and concurrently execute the quantization processing of all processing blocks;
[0267] Perform in-block token position verification on the processing results of each processing block to eliminate intermediate results corresponding to invalid token positions;
[0268] Sort the valid intermediate results of all processing blocks according to the starting position index and splice them into a complete model output sequence;
[0269] Perform post-processing operations on the spliced output sequence to generate the final model inference result.
[0270] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in Figure 4 . The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a model quantization inference acceleration method.
[0271] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as shown in Figure 5 . The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a model quantization inference acceleration method.
[0272] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:
[0273] Divide the input text into multiple processing blocks, fix the processing precision format of the first processing block to a high-precision format, and disable the quantization processing of the first processing block;
[0274] For the other processing blocks except the first processing block among the multiple processing blocks, generate the self-attention matrix of each other processing block through a language model, determine the sum of all the element values in the column corresponding to each token position in the self-attention matrix, and use the sum of all the element values as the importance score of each token position;
[0275] Assign the token positions with importance scores greater than the first threshold to the high-precision format, assign the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and assign the token positions with importance scores less than the second threshold to the low-precision format;
[0276] Count the number of token positions assigned to the high-precision format, medium-precision format, and low-precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration for the corresponding processing block;
[0277] Divide the network modules of the language model into multiple configuration sharing groups, and each configuration sharing group contains at least two network modules;
[0278] Within each configuration sharing group, share the unified quantization configuration of the processing block corresponding to the first network module with other network modules within the same configuration sharing group;
[0279] According to the unified quantization configuration corresponding to each processing block, perform block-level batch quantization on all processing blocks and complete model inference to generate model inference results.
[0280] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0281] Divide the input text into multiple processing blocks, fix the processing precision format of the first processing block to the high-precision format, and disable the quantization processing of the first processing block;
[0282] For the other processing blocks except the first processing block among the multiple processing blocks, generate the self-attention matrix of each other processing block through the language model, determine the sum of all element values in the column corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score of each token position;
[0283] Assign the token positions with importance scores greater than the first threshold to the high-precision format, assign the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and assign the token positions with importance scores less than the second threshold to the low-precision format;
[0284] Count the number of token positions assigned to the high-precision format, medium-precision format, and low-precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration for the corresponding processing block;
[0285] Divide the network modules of the language model into multiple configuration sharing groups, and each configuration sharing group contains at least two network modules;
[0286] Share the unified quantization configuration of the processing block corresponding to the first network module within each configuration sharing group with other network modules in the same configuration sharing group;
[0287] According to the unified quantization configuration corresponding to each processing block, perform block-level batch quantization on all processing blocks and complete model inference to generate a model inference result.
[0288] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0289] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to memory, storage, database, or other media used in the various embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0290] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0291] It should be noted that in the embodiments of this application, if there are software tools or components of other companies, they are only used for illustrative introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for accelerating model quantization inference, characterized in that Including the following steps: Dividing the input text into multiple processing blocks, fixing the processing precision format of the first processing block to a high-precision format, and disabling the quantization processing for the first processing block; For the other processing blocks in the multiple processing blocks except the first processing block, generating the self-attention matrix of each other processing block through a language model, determining the sum of all element values in the column corresponding to each token position in the self-attention matrix, and using the sum of all element values as the importance score for each token position; Assigning the token positions with importance scores greater than the first threshold to the high-precision format, assigning the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and assigning the token positions with importance scores less than the second threshold to the low-precision format; Counting the number of token positions assigned to the high-precision format, medium-precision format, and low-precision format in each processing block, and selecting the precision format with the largest number as the unified quantization configuration for the corresponding processing block; Dividing the network modules of the language model into multiple configuration sharing groups, and each configuration sharing group contains at least two network modules; Sharing the unified quantization configuration of the processing block corresponding to the first network module within each configuration sharing group with the other network modules within the same configuration sharing group; Performing block-level batch quantization on all processing blocks according to the unified quantization configuration corresponding to each processing block and completing model inference to generate a model inference result.
2. The model quantization inference acceleration method according to claim 1, wherein Dividing the input text into multiple processing blocks, fixing the processing precision format of the first processing block to a high-precision format, and disabling the quantization processing for the first processing block, including: Dividing the input text into multiple equally long processing blocks according to a preset block length; Disabling the adjustment of quantization parameters for all token positions in the first processing block; Fixing the processing precision format of the embedding layer, self-attention layer, and feed-forward network layer of the first processing block in the language model to a high-precision format; When there are remaining tokens at the end of the input text that are less than the preset block length, forming the text segment composed of the remaining tokens as an independent processing block; Padding the independent processing block with invalid tokens to the preset block length, fixing the processing precision format of the padded independent processing block to a high-precision format, and disabling the quantization operation for the independent processing block; Recording the start position index and end position index of all processing blocks.
3. The model quantization inference acceleration method according to claim 1, wherein, For the other processing blocks in the multiple processing blocks except the first processing block, generating the self-attention matrix of each other processing block through a language model, determining the sum of all element values in the column corresponding to each token position in the self-attention matrix, and using the sum of all element values as the importance score for each token position, including: Inputting each other processing block into the multi-head self-attention layer of the language model to obtain the local self-attention matrices of multiple attention heads corresponding to each other processing block; Performing weighted average or arithmetic average processing on the local self-attention matrices of multiple attention heads to generate the final self-attention matrix corresponding to each other processing block; Extract the column vectors corresponding to each token position from the final self-attention matrix; Determine the sum of the values of all elements in each column vector, and use the sum as the importance score of the token position corresponding to each column vector; If there are filled invalid token positions in other processing blocks, when determining the importance scores of the filled invalid token positions, set all element values of the column vectors corresponding to the invalid token positions to zero.
4. The model quantization inference acceleration method according to claim 1, wherein Assign the token positions with importance scores greater than the first threshold to the high-precision format, assign the token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and assign the token positions with importance scores less than the second threshold to the low-precision format, including: Determine the first threshold and the second threshold according to the length of the processing block; If there are invalid token positions, assign the processing precision format of the invalid token positions to the low-precision format, and exclude the invalid token positions when counting the number of precision formats of each processing block; Compare the importance scores of each valid token position with the first threshold and the second threshold; Assign the processing precision format of the valid token positions with importance scores greater than the first threshold to the high-precision format; Assign the processing precision format of the valid token positions with importance scores above the second threshold and below the first threshold to the medium-precision format; Assign the processing precision format of the valid token positions with importance scores less than the second threshold to the low-precision format.
5. The model quantization inference acceleration method according to claim 1, characterized in that, Count the number of token positions assigned to the high-precision format, medium-precision format, and low-precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration corresponding to the processing block, including: Traverse all token positions of the current processing block to identify the token positions assigned to the high-precision format, medium-precision format, and low-precision format; During the counting process, if there are invalid token positions, exclude all invalid token positions and only count the precision format assignment results of valid token positions; Count the number of valid token positions assigned to the high-precision format, medium-precision format, and low-precision format respectively, and generate the high-precision count, medium-precision count, and low-precision count; Compare the numerical sizes of the high-precision count, medium-precision count, and low-precision count; Take the precision format corresponding to the count with the largest value as the unified quantization configuration of the current processing block; If there are multiple counts of precision formats that are the same and the maximum value, select the highest precision format among the multiple precision formats as the unified quantization configuration of the current processing block.
6. The model quantization inference acceleration method according to claim 1, wherein, Share the unified quantization configuration of the processing block corresponding to the first network module within each configuration sharing group to other network modules within the same configuration sharing group, including: Assign a unique block identifier to each processing block and a unique group identifier to each configuration sharing group; Establish a binding relationship between the unique group identifier and the unique block identifier in the configuration mapping table; Based on the binding relationship between the unique group identifier and the unique block identifier, associate the first network module of each configuration sharing group with the unified quantization configuration parameters of the corresponding processing block; Write the unified quantization configuration parameters into the shared memory area of each configuration sharing group, and allocate the same memory address mapping for all network modules of each configuration sharing group; When other network modules within the same configuration sharing group perform quantization processing, read the unified quantization configuration parameters from the shared memory area through the memory address mapping.
7. The model quantization inference acceleration method according to claim 1, wherein According to the unified quantization configuration corresponding to each processing block, perform block-level batch quantization on all processing blocks and complete model inference to generate model inference results, including: Load the corresponding unified quantization configuration parameters for each processing block, and the unified quantization configuration parameters include precision format identifiers; According to the precision format identifier, configure independent processing kernel functions for different processing blocks; In the parallel processing unit of the graphics processor or tensor processor, allocate processing resources according to the index order of the processing blocks, and concurrently execute the quantization processing of all processing blocks; Perform in-block token position verification on the processing results of each processing block to eliminate intermediate results corresponding to invalid token positions; Sort the valid intermediate results of all processing blocks according to the start position index and splice them into a complete model output sequence; Perform post-processing operations on the spliced output sequence to generate the final model inference result.
8. A model quantization inference acceleration device, characterized in that, The model quantization inference acceleration device includes: An input text preprocessing module, which is used to divide the input text into multiple processing blocks, fix the processing precision format of the first processing block to a high-precision format, and disable the quantization processing of the first processing block; A self-attention analysis module, which is used to generate the self-attention matrix of each other processing block except the first processing block among the multiple processing blocks through a language model, determine the sum of all element values in the column corresponding to each token position in the self-attention matrix, and use the sum of all element values as the importance score of each token position; A precision allocation module, which is used to allocate token positions with importance scores greater than the first threshold to the high-precision format, allocate token positions with importance scores above the second threshold and below the first threshold to the medium-precision format, and allocate token positions with importance scores less than the second threshold to the low-precision format; A quantization configuration decision module, which is used to count the number of token positions allocated to the high-precision format, medium-precision format, and low-precision format within each processing block, and select the precision format with the largest number as the unified quantization configuration of the corresponding processing block; A network module grouping control module, which is used to divide the network modules of the language model into multiple configuration sharing groups, and each configuration sharing group includes at least two network modules; A configuration sharing management module, which is used to share the unified quantization configuration of the processing block corresponding to the first network module within each configuration sharing group with other network modules within the same configuration sharing group; The inference execution module is used to perform block-level batch quantization on all processing blocks according to the unified quantization configuration corresponding to each processing block, complete model inference, and generate model inference results.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a model quantization inference acceleration program stored in the memory and executable on the processor. When the model quantization inference acceleration program is executed by the processor, it implements the steps of the model quantization inference acceleration method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a model quantization inference acceleration program. When the model quantization inference acceleration program is executed by the processor, it implements the steps of the model quantization inference acceleration method according to any one of claims 1-7.
Citation Information
Patent Citations
Neural network acceleration hardware architecture and method for quantization bit width dynamic selection
CN113902108A
Text reasoning acceleration method applied to large language model and related device
CN118394895A