Large model task execution method, task processing method, device, intelligent agent, equipment, medium and product
Patent Information
- Application Number
- CN202610882350.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-04
Smart Images

Figure CN122693731A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of artificial intelligence, deep learning, large model technology, and model quantization technology. More specifically, it relates to a method for executing large model tasks, a task processing method, a device, an intelligent agent, an equipment, a medium, and a product. Background Technology
[0002] Large Language Models (LLMs) have demonstrated outstanding performance in various tasks such as natural language processing, code generation, and multimodal understanding. How to store and deploy models for inference in order to cope with their massive number of parameters has become an important research direction. Summary of the Invention
[0003] In view of this, this disclosure provides a method for executing large-scale model tasks, a task processing method, an apparatus, an intelligent agent, a device, a medium, and a product.
[0004] According to one aspect of this disclosure, a method for executing a large model task is provided, comprising: quantizing first model parameters of a network layer of a model to be quantized, indicated by a quantization task, based on a plurality of candidate quantization parameters, to obtain a plurality of candidate model parameters; determining a target model parameter among the plurality of candidate model parameters with the objective of minimizing the quantization error between the first model parameters and the candidate model parameters; and performing channel dimension compensation on the target model parameter based on sample data to obtain channel compensation parameters, and storing the target model parameter, the target quantization parameter corresponding to the target model parameter, and the channel compensation parameters as quantization results corresponding to the model identifier of the model to be quantized in a storage unit.
[0005] According to another aspect of this disclosure, a method for processing large model tasks is provided, comprising: obtaining a quantization result from a storage unit according to a model identifier indicated by a task to be processed, wherein the quantization result is obtained using a task execution method; and processing the task to be processed using a target model to obtain a task processing result, wherein the target model is obtained by configuring a model for the task to be processed using the quantization result.
[0006] According to another aspect of this disclosure, a large model task execution apparatus is provided, comprising: a quantization module, configured to quantize first model parameters of a network layer of a model to be quantized, indicated by a quantization task, based on a plurality of candidate quantization parameters, to obtain a plurality of candidate model parameters; a determination module, configured to determine a target model parameter among the plurality of candidate model parameters with the objective of minimizing the quantization error between the first model parameters and the candidate model parameters; and a compensation module, configured to perform channel dimension compensation on the target model parameter based on sample data to obtain channel compensation parameters, and store the target model parameter, the target quantization parameter corresponding to the target model parameter, and the channel compensation parameters as quantization results corresponding to the model identifier of the model to be quantized in a storage unit.
[0007] According to another aspect of this disclosure, a large model task processing apparatus is provided, comprising: an acquisition module for acquiring quantization results from a storage unit based on a model identifier indicated by a task to be processed, wherein the quantization results are obtained using a task execution device; and a processing module for processing the task to be processed using a target model to obtain a task processing result, wherein the target model is obtained by configuring a model for the task to be processed using the quantization results.
[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0009] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program or instructions thereon, which, when executed by a processor, implement the steps of the above-described method.
[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0013] Figure 1This illustration schematically shows a system architecture for applying large model task execution methods and large model task processing methods according to embodiments of the present disclosure;
[0014] Figure 2 A flowchart illustrating a large-scale task execution method according to an embodiment of the present disclosure is shown schematically.
[0015] Figure 3 This illustration schematically shows an example diagram of a process for determining a target model parameter from a plurality of candidate model parameters according to an embodiment of the present disclosure;
[0016] Figure 4 This illustration schematically depicts an example of a process for determining a target model parameter from a plurality of candidate model parameters with the aim of minimizing the quantization error between a first model parameter and candidate model parameters, according to an embodiment of the present disclosure.
[0017] Figure 5 This illustration schematically shows an example of the process of performing channel dimension compensation on target model parameters based on sample data to obtain channel compensation parameters according to an embodiment of the present disclosure;
[0018] Figure 6 A flowchart illustrating a large model task processing method according to an embodiment of the present disclosure is shown schematically.
[0019] Figure 7 A block diagram of a large-scale task execution apparatus according to an embodiment of the present disclosure is shown schematically;
[0020] Figure 8 A block diagram of a large-scale task processing apparatus according to an embodiment of the present disclosure is shown schematically;
[0021] Figure 9 A schematic diagram illustrating the structure of a large-model-based intelligent agent according to embodiments of the present disclosure; and
[0022] Figure 10 A block diagram of an electronic device suitable for implementing a large model task execution method and a large model task processing method according to embodiments of the present disclosure is shown schematically. Detailed Implementation
[0023] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0024] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0026] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0027] Large language models have billions or even trillions of weight parameters. Even when stored in 16-bit floating-point format, a single model requires tens of gigabytes of storage space. During inference, weight parameters need to be repeatedly read from GPU memory to perform matrix multiplication operations. The bandwidth for reading weight parameters becomes the main bottleneck restricting inference throughput in autoregressive decoding scenarios. Therefore, how to perform low-bit quantization compression of model weight parameters to reduce storage consumption and memory bandwidth requirements is a key way to achieve efficient deployment of large language models.
[0028] In one example, model weight quantization methods can include scalar quantization and lookup table-based codebook quantization. Scalar quantization maps each weight value independently to a finite number of discrete values. However, this method can only represent four discrete values with extremely low bit widths, resulting in insufficient expressive power and a sharp drop in model accuracy after quantization. Lookup table-based codebook quantization constructs a codebook through methods such as clustering, storing a corresponding codebook index for each weight, and retrieving the weight value from the codebook during dequantization. However, this method requires frequent access to the codebook storage area during inference, introducing additional memory read latency; moreover, the storage overhead of the codebook itself becomes non-negligible with extremely low bit widths, contradicting the goal of compression. Furthermore, the above methods typically employ a uniform quantization strategy across all network layers during quantization, ignoring the differences in weight statistical distributions between different layers, leading to excessive quantization errors in some layers.
[0029] Furthermore, in extremely low bit quantization scenarios, due to the extremely limited availability of discrete values, simple element-wise mapping will produce obvious stripe-like, grid-like, or slice-like artifacts in the reconstructed weight distribution, which seriously deviates from the high-dimensional distribution structure of the original weights, resulting in irreversible degradation of the semantic quality of the model output, thus hindering the practical application of extremely low bit quantization.
[0030] To address this, this disclosure proposes a task execution scheme. For example, based on multiple candidate quantization parameters, the first model parameters of the network layer of the model to be quantized, as indicated by the quantization task, are quantized to obtain multiple candidate model parameters. With the goal of minimizing the quantization error between the first model parameters and the candidate model parameters, a target model parameter is determined from the multiple candidate model parameters. Channel dimension compensation is performed on the target model parameter based on sample data to obtain channel compensation parameters. The target model parameter, the target quantization parameter corresponding to the target model parameter, and the channel compensation parameters are stored in a storage unit as quantization results corresponding to the model identifier of the model to be quantized.
[0031] According to embodiments of this disclosure, by employing multiple candidate quantization parameters to quantize the first model parameters respectively, diverse candidate model parameters can be generated in parallel within the same network layer, expanding the search coverage of the optimal compression scheme. By determining the target model parameters with the objective of minimizing the quantization error between the first model parameters and the candidate model parameters, the selected quantization representation exhibits the smallest overall deviation from the original parameters among all candidate schemes, thereby maximally suppressing the parameter information loss introduced by the quantization process.
[0032] Based on this, channel compensation parameters are obtained by compensating the target model parameters based on sample data in the channel dimension. This allows the compensation parameters to fit the statistical distribution characteristics of the real inference data in the channel dimension, and specifically corrects the deviation of quantization in the channel dimension. Thus, the consistency of the quantization model's output to real data can be further improved while keeping the existing structure of the target model parameters unchanged.
[0033] Furthermore, by storing the target model parameters, target quantization parameters, and channel compensation parameters in the storage unit in correspondence with the model identifier, the correspondence between each parameter and the model is uniquely determined. During deployment and loading, the complete parameters of the model can be obtained based on the model identifier, thereby ensuring the accurate reuse of the quantization model.
[0034] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution of this invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0035] In the technical solution of the present invention, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.
[0036] Figure 1 The illustration schematically depicts a system architecture for applying large-scale task execution methods and large-scale task processing methods according to embodiments of the present disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0037] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, a processing unit 105, and a storage unit 106. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, the processing unit 105, and the storage unit 106. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0038] Users can interact with at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 via the network 104, the processing unit 105, and the storage unit 106 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0039] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0040] Processing unit 105 can be a processing unit that provides various services, such as a background management processing unit that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The background management processing unit can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0041] In one embodiment, processing unit 105 may be a central processing unit (CPU). Storage unit 106 may store the original model parameters of the network layers of the model to be quantized. Processing unit 105 may obtain the original model parameters of the network layers of the model to be quantized from storage unit 106, execute the large model task execution method provided in this embodiment of the disclosure, obtain target model parameters, target quantization parameters, and channel compensation parameters, and store the target model parameters, target quantization parameters, and channel compensation parameters as quantization results corresponding to the model identifier of the model to be quantized in storage unit 106.
[0042] It should be noted that the large model task execution method provided in this embodiment can generally be executed by the processing unit 105. Accordingly, the large model task execution device provided in this embodiment can generally be disposed in the processing unit 105.
[0043] Alternatively, the large model task processing method provided in this disclosure can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or it can be executed by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the large model task execution device and the large model task processing device provided in this disclosure can also be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0044] It should be understood that Figure 1 The number of terminal devices, networks, processing units, and storage units shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, processing units, and storage units can be included.
[0045] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.
[0046] Figure 2 A flowchart illustrating a large-scale task execution method according to an embodiment of the present disclosure is shown schematically.
[0047] like Figure 2 As shown, the large model task execution method 200 may include operations S210~S230.
[0048] In operation S210, the first model parameters of the network layer of the model to be quantized, as indicated by the quantization task, are quantized according to multiple candidate quantization parameters to obtain multiple candidate model parameters.
[0049] In operation S230, with the goal of minimizing the quantization error between the first model parameters and the candidate model parameters, the target model parameters are determined from multiple candidate model parameters.
[0050] In operation S230, channel dimension compensation is performed on the target model parameters based on the sample data to obtain channel compensation parameters. The target model parameters, the target quantization parameters corresponding to the target model parameters, and the channel compensation parameters are stored in the storage unit as quantization results corresponding to the model identifier of the model to be quantized.
[0051] A quantization task is an instruction to perform low-bit weight compression on a specified network layer in a large, trained language model. This quantization task specifies at least the model identifier, the range of network layers to be quantized, and the required quantization bit width. The quantization task can be initiated by the model compression scheduler, the model deployment pipeline, or manually by the user; no specific restrictions apply here.
[0052] The network layers of the model to be quantized are the specific layers in the large language model that need to undergo weight quantization and compression after training. For example, they may include linear projection layers such as query projection matrix layer q_proj, key projection matrix layer k_proj, value projection matrix layer v_proj, output projection matrix layer o_proj, gated projection layers gate_proj, up_proj, and down_proj in the Transformer architecture.
[0053] Candidate quantization parameters are discrete parameter combinations selected from a pre-built candidate pool to recode model parameters into a low-bit compressed form. Candidate quantization parameters can include multiplication parameters and shift parameters. The multiplication parameters are 32-bit odd-number multiplication constants, and the shift parameters are the number of right shifts that control the strength of the feedback from higher bits to lower bits. It should be noted that different candidate quantization parameters represent computational paths that generate different approximate Gaussian distributions. After traversing all possible parameter combinations through an offline heuristic search, the candidate pool can select the most frequently occurring parameter sets in each network layer of the model.
[0054] The first model parameters are the original high-precision weight matrix stored on the storage medium in the network layer, typically represented in FP16 (16-bit floating-point) or BF16 (16-bit floating-point) format. After obtaining multiple candidate quantization parameters, the first model parameters can be quantized using each candidate quantization parameter to obtain the candidate model parameters corresponding to each candidate quantization parameter. These candidate model parameters are the approximate weight matrix obtained by quantizing, encoding, and decoding the first model parameters using a specific candidate quantization parameter, expressed in low-bit compressed form, and recovered through dequantization. It should be noted that the differences between different candidate model parameters stem from the differences in Gaussian source morphology and mesh topology generated by different candidate quantization parameters, resulting in varying approximation accuracy of the first model parameters.
[0055] Quantization error is a metric that measures the degree of deviation between the parameters of a candidate model and those of the first model. Quantization error can be used to apply differentiated penalties to errors in different directions of the model parameters; for example, errors in sensitive directions that have a greater impact on the model output are given higher weights, while less sensitive directions are tolerated more leniently. The method for determining quantization error can be configured according to actual business needs and is not limited here. For example, quantization error can be determined using mean squared error, L2 distance, or reconstruction error, etc.
[0056] The target model parameters are the model parameters that minimize the quantization error among multiple candidate model parameters. Structurally, the target model parameters maintain the same dimensions as the first model parameters, but their values are determined by the optimal path under low-bit trellis coding constraints. At this point, these parameters have not yet undergone channel-continuous compensation, and there may still be systematic deviations between them and the layer outputs of the original weights at extremely low bit depths.
[0057] The sample data is an unlabeled calibration dataset used to guide channel dimension compensation optimization. It should be noted that this sample data does not require annotation and can be user query text, program code snippets, natural language paragraphs, or other modal data such as images and audio; no specific restrictions are imposed here. The sample data is used to drive the forward propagation of network layers, collect the statistical distribution of layer input activations, and construct the optimization objective for channel compensation based on this distribution: minimizing the squared Frobenius norm between the quantized layer output and the original FP16 layer output.
[0058] Channel dimension compensation, under the premise that the target model parameters and target quantization parameters are frozen, applies learnable diagonal scaling factors to the input and output channels of the target model parameters to correct the systematic output bias introduced by extremely low bit quantization. The channel compensation parameters are the input channel scaling vector S_in and the output channel scaling vector S_out, learned during the channel dimension compensation process described above. The length of each vector is equal to the dimension corresponding to the weight matrix.
[0059] The model identifier is a unique identifier that identifies the model to be quantized. It can be the model name, hash value, version number, or storage path. The model identifier can be used to bind the quantization result to the model to be quantized, ensuring that the correspondence between the target model parameters, target quantization parameters, and channel compensation parameters is not confused during deployment and loading.
[0060] A storage unit is a physical or virtual storage medium used to persistently or temporarily store quantization results. For example, a storage unit may include, but is not limited to, GPU memory, CPU memory, solid-state drive (SSD), hard disk drive (HDD), network attached storage (NAS), or cloud object storage bucket, etc., without limitation.
[0061] In the embodiments of this disclosure, by using multiple candidate quantization parameters to quantize the first model parameters respectively, diverse candidate model parameters can be generated in parallel within the same network layer, expanding the search coverage of the optimal compression scheme. By determining the target model parameters with the goal of minimizing the quantization error between the first model parameters and the candidate model parameters, the selected quantization representation has the smallest overall deviation from the original parameters among all candidate schemes, thereby maximally suppressing the parameter information loss introduced by the quantization process.
[0062] Based on this, channel compensation parameters are obtained by compensating the target model parameters based on sample data in the channel dimension. This allows the compensation parameters to fit the statistical distribution characteristics of the real inference data in the channel dimension, and specifically corrects the deviation of quantization in the channel dimension. Thus, the consistency of the quantization model's output to real data can be further improved while keeping the existing structure of the target model parameters unchanged.
[0063] Furthermore, by storing the target model parameters, target quantization parameters, and channel compensation parameters in the storage unit in correspondence with the model identifier, the correspondence between each parameter and the model is uniquely determined. During deployment and loading, the complete parameters of the model can be obtained based on the model identifier, thereby ensuring the accurate reuse of the quantization model.
[0064] According to embodiments of this disclosure, the large model task execution method 200 may further include the following operation: processing sample data using network layers to obtain reference sensitivity information, wherein the reference sensitivity information characterizes the degree of influence of quantization error on the output of the network layer.
[0065] A network layer is a layered structural unit in a trained large language model that performs forward computation. It has learnable weight parameters and can receive input activations and generate output activations. By processing sample data using network layers and observing their input-output behavior on the sample data, we can infer the strength of the influence of each weight direction on the final output, thereby obtaining reference sensitivity information.
[0066] Reference sensitivity information is a reference indicator extracted from the statistical distribution of layer input activations after the network layer has been propagated forward using sample data. It measures the impact of quantization error on the network layer output. Reference sensitivity information can provide quantitative directional constraints for the targeted allocation of bit precision during subsequent quantization processes, ensuring that a limited bit budget prioritizes the protection of sensitive weight directions that have a significant impact on the output.
[0067] In one embodiment, the reference sensitivity information can be determined based on the mean vector of the layer input activations and the symmetric compressed representation of the layer input activations. The reference sensitivity information can be a symmetric positive semi-definite matrix with the same dimension as the original model parameters of the network layer, whose quadratic form in any direction approximates the deviation of the layer output if a unit perturbation is applied to the weights in that direction.
[0068] In the embodiments of this disclosure, by utilizing network layers to process sample data, the extraction process of sensitivity information takes the actual forward computation behavior of the network layer itself as the observation window. The information obtained is directly derived from the response pattern of the layer to the real data distribution, rather than based on external assumptions or theoretical approximations unrelated to the layer structure, thereby ensuring the intrinsic consistency between the reference sensitivity information and the specific computational characteristics of the layer.
[0069] According to embodiments of this disclosure, processing sample data using a network layer to obtain reference sensitivity information may include the following operations: determining intermediate sensitivity information based on the distribution information obtained by processing sample data using a network layer; performing incoherent processing on the intermediate sensitivity information to make the parameter distribution approximate a Gaussian distribution, thereby obtaining the reference sensitivity information.
[0070] Distribution information is a quantitative description of the statistical distribution characteristics of activation values extracted from the layer input activations after the network layer processes sample data. Distribution information can include the mean vector of the layer input activations and the symmetric compressed representation of the layer input activations. The mean vector of the layer input activations reflects the central position of the activations in each input dimension, and the symmetric compressed representation of the layer input activations is the encoding form of the second-order moment matrix of the activations.
[0071] Intermediate sensitivity information refers to the initial curvature proxy matrix constructed from the distribution information, which has not yet undergone incoherent processing. It is obtained by decompressing the compressed representation back to the complete second-order moment matrix and adding it to the mean outer product matrix. Since the intermediate sensitivity information is still under the original coordinate basis, its sensitivity distribution in each direction is affected by the coherent structure of the weight matrix, and may exhibit highly non-uniform anisotropy.
[0072] In one embodiment, intermediate sensitivity information can be obtained using the following formula (1).
[0073] (1)
[0074] in, Characterizing intermediate sensitivity information, and Represents distribution information.
[0075] For example, sample data can be fed into the model to be quantized, and the input activation vectors generated in the forward computation flow of the network layer can be extracted. Statistical analysis is performed on all input activation vectors to extract distribution information: the first-order statistic is calculated to obtain the mean vector, and the second-order moment matrix of the activation is symmetrically compressed to obtain the encoded form of the second-order statistic. The compressed representation is restored to the complete second-order moment matrix through decompression, and then added to the mean outer product matrix to construct the intermediate sensitivity information. This intermediate sensitivity information completely encodes the first-order and second-order statistical features of the network layer input activations, and each element quantifies the joint influence strength between perturbations in the i-th row and j-th column directions of the model parameters.
[0076] Incoherence processing is a mathematical operation that reduces the coherence between the weight directions and the sensitivity directions by applying coordinate basis transformations to the intermediate sensitivity information and corresponding weight parameters. For example, a random sign vector can be first applied to flip the sign of each row / column, and then the Fast Walsh–Hadamard Transform (FWHT) can be applied for orthogonal basis rotation. The transformed reference sensitivity information is under the new matching basis, the coherence between the weight directions and the sensitivity directions is weakened, and the overall distribution approaches an independent Gaussian distribution.
[0077] Parameter distribution refers to the statistical distribution of the original model parameters of a network layer in various directions. Under the original coordinate basis, the original model parameters may exhibit strong structure and anisotropy, meaning that values in some directions are greater than those in others, or there is a strong correlation between adjacent elements. This non-uniform distribution poses a challenge to subsequent low-bit trellis coding because trellis coding assumes that the encoded signal approximately follows an independent Gaussian distribution. After incoherent processing, the parameter distribution of the weights under the new basis is forced to approach an approximately Gaussian form of independent and identically distributed parameters, thus matching the theoretical assumptions of low-bit trellis coding.
[0078] In the embodiments of this disclosure, intermediate sensitivity information is determined by using the distribution information obtained from processing sample data using network layers. This ensures that the construction process of the sensitivity matrix uses the actual activation statistics endogenous to the network layers as the data source, rather than relying on external assumptions or theoretical approximations. This ensures that the representation of the relationship between the intermediate sensitivity information and the network layer's perturbation output response is data-driven and directly corresponds to the actual behavior of the layer on the samples.
[0079] Based on this, by performing incoherent processing on the intermediate sensitivity information, the sensitivity energy that was originally concentrated in a few directions is forcibly diffused to all orthogonal directions, which reduces the difference in weights of each direction in the subsequent quantization error measurement and avoids a few extremely sensitive directions from dominating the quantization decision. This provides an isotropic error evaluation benchmark for grid coding under uniform bit allocation, which helps to improve the accuracy of the target model parameters obtained by subsequent quantization.
[0080] According to an embodiment of this disclosure, operation S220 may include the following operations: dequantizing candidate model parameters in a sliding window manner to obtain dequantized model parameters; and determining target model parameters from multiple candidate model parameters with the goal of minimizing the difference between the dequantized model parameters and the first model parameters.
[0081] The sliding window method refers to a technique used when decoding and dequantizing a compressed bitstream. A fixed-length window slides across the bitstream segment by segment, truncating sub-bits. Adjacent windows share some bits, thus converting the low-bit compressed representation to a sequence of floating-point code values. The parameters of the sliding window can include window length L, number of bits per symbol k, number of decoded values per step V, and window movement step size Δ = kV.
[0082] It should be noted that in the embodiments of this disclosure, adjacent windows x_t and x_{t+1} share L-Δ bits. This overlapping design creates a structured dependency between consecutive decoded values, thereby forming a bitshift trellis grid, which provides a graph structure basis for subsequent optimal path search.
[0083] For example, with L=16, k=2, V=4, and Δ=8, the first window x1 extracts bits 0 to 15 of the candidate model parameters, corresponding to four 2-bit code values. After the window slides forward by Δ=8 bits, the second window x2 extracts bits 8 to 23 of the candidate model parameters, where bits 8 to 15 are shared with x1. If the candidate model parameters have a total of 32 bits, then three sliding windows can be generated, with the starting bits of the windows being 0, 8, and 16, respectively, resulting in 12 code values after decoding.
[0084] Dequantization is the process of restoring the low-bit compressed representation to floating-point precision weight values through a preset calculation method; it is the inverse operation of quantization. This dequantization process does not require runtime table lookups; all code values are calculated in real-time through a preset calculation method, without needing to read a pre-stored codebook.
[0085] The dequantized model parameters are floating-point precision reconstructed weight matrices obtained by dequantizing the candidate model parameters using a sliding window method. These dequantized model parameters have the exact same dimensions and numerical precision format as the first model parameters and can directly participate in the forward matrix multiplication calculations of network layers. While the dequantized model parameters are an approximation of the first model parameters, they suffer from information loss due to the significantly smaller number of bits required for storage compared to the first model parameters. The difference between the dequantized model parameters and the first model parameters is the quantization error.
[0086] After all candidate model parameters have been dequantized to obtain their respective dequantized model parameters, the deviation between each dequantized model parameter and the first model parameter can be calculated using a preset difference metric function. After obtaining the deviation values for each dequantized model parameter, they can be compared and sorted, and the candidate model parameter with the smallest difference value can be selected as the target model parameter.
[0087] In the embodiments of this disclosure, candidate model parameters are dequantized using a sliding window approach. A window overlap mechanism establishes structured dependency constraints between adjacent decoded code values, enabling the dequantization process to expand the low bitstream into a high-dimensional weight matrix while preserving the correlation between code values, thus obtaining the dequantized model parameters. Based on this, target model parameters are determined by minimizing the difference between the dequantized model parameters and the first model parameters. The target model parameters are selected from multiple candidate model parameters. Through a combination of multi-candidate parallel trial quantization and a unified difference metric, different network layers can independently select the optimal parameters based on their own weight statistical characteristics, achieving hierarchical adaptive quantization without manual intervention.
[0088] Figure 3 The illustration shows an example diagram of a process for determining a target model parameter from a plurality of candidate model parameters according to an embodiment of the present disclosure.
[0089] like Figure 3As shown, in embodiment 300 for determining target model parameters, taking the model 310 to be quantized as indicated by the quantization task as including network layer 311, network layer 312, ..., network layer 31M, and N candidate quantization parameters, where M and N are both positive integers, the following describes the process of determining the target model parameters of network layer 311.
[0090] The first model parameters of network layer 311 are quantized according to multiple candidate quantization parameters to obtain multiple candidate model parameters. For example, the first model parameters of network layer 311 are quantized according to candidate quantization parameter 321 to obtain candidate model parameter 331; the first model parameters of network layer 311 are quantized according to candidate quantization parameter 322 to obtain candidate model parameter 332; and so on, the first model parameters of network layer 311 are quantized according to candidate quantization parameter 32N to obtain candidate model parameter 33N.
[0091] With the goal of minimizing the quantization error between the first model parameter and the candidate model parameter, the target model parameter 340 is determined from the candidate model parameter 331, candidate model parameter 332, ..., candidate quantization parameter 32N.
[0092] According to embodiments of this disclosure, dequantizing candidate model parameters using a sliding window to obtain dequantized model parameters may include the following operations: truncating candidate model parameters in a sliding manner according to a preset window width to obtain multiple sub-model parameters, wherein adjacent sub-model parameters have a preset number of identical parameters; for each sub-model parameter, dequantizing the sub-model parameter according to the candidate quantization parameter corresponding to the candidate model parameter to obtain dequantized sub-model parameters; and concatenating the dequantized sub-model parameters according to their order of arrangement in the candidate model parameters to obtain dequantized model parameters.
[0093] The preset window width (L) is the number of consecutive bits (L) truncated from the candidate model parameters at each step during sliding window dequantization. The preset window width is a hyperparameter determined and fixed during the offline quantization stage, and its value determines the amount of bit information contained in each sliding window. The preset window width L is related to the number of bits per symbol (k) and the number of decoded values per step (V); that is, each L-bit window can decode V low-bit code values, with each code value occupying k bits.
[0094] It should be noted that the choice of preset window width affects the balance between dequantization accuracy and compression ratio; that is, the larger the window width, the more overlapping bits between adjacent windows, the stronger the mesh constraint, and the higher the reconstruction accuracy. The preset window width remains unchanged after model deployment, and the same window parameter is used for dequantization of all layers during the inference phase.
[0095] The sliding method involves moving the starting position of the window segment by segment along the bitstream of candidate model parameters with a fixed step size Δ to generate a truncation pattern for the continuous sub-model parameter sequence. The sliding step size Δ = kV, which is equal to the number of effective bits decoded in each step. The sliding process can be as follows: the first window truncates L bits from bit offset 0 in the bitstream; after the window slides forward Δ bits, the second window truncates L bits from bit offset Δ; and so on until the end of the candidate model parameters.
[0096] For example, for each candidate model parameter, sub-model parameters can be extracted from the candidate model parameters according to the preset sliding window parameters L, k, and V, with adjacent windows sharing L-Δ bits; each window x_t is used as input and calculated in a preset way, and after modular multiplication, XOR shift high-bit feedback, four-byte summation, and hierarchical affine normalization, the dequantized sub-model parameters are output; the dequantized sub-model parameters obtained by decoding all windows are reorganized into a two-dimensional matrix form according to the original model parameters of the network layer, thus obtaining the dequantized model parameters.
[0097] The sub-model parameters are continuous bit segments truncated from the candidate model parameters using a sliding window with a preset window width L. Each sub-model parameter is an L-bit unsigned integer, used as input for a preset calculation method. Adjacent sub-model parameters share a preset number of bits, which can be L-Δ.
[0098] The dequantized sub-model parameters are scalar floating-point code values output after processing individual sub-model parameters using a preset calculation method. The order of arrangement is the order in which each sub-model parameter appears in the bitstream of the candidate model parameters, from smallest to largest, according to the starting offset of the sliding window. All dequantized sub-model parameters are combined sequentially according to the above arrangement to form a complete two-dimensional floating-point weight matrix, thus obtaining the dequantized model parameters.
[0099] In the embodiments of this disclosure, multiple sub-model parameters are obtained by truncating the candidate model parameters in a sliding manner according to a preset window width, and there are a preset number of identical parameters between adjacent sub-model parameters. A structured dependency constraint is established between the sub-model parameters through a window overlap mechanism, so that the influence of each bit in the bit stream spans multiple output code values, thereby enhancing the expressive continuity of the compressed representation without changing the equivalent bit rate.
[0100] Building upon this foundation, for each sub-model parameter, dequantization is performed based on the candidate quantization parameters corresponding to the candidate model parameters. This ensures that the dequantization path of each sub-model parameter precisely matches the decoder parameters used in its encoding stage, avoiding systematic deviations in dequantized code values caused by inconsistencies in encoding and decoding parameters. By independently converting each window segment in the candidate model parameters into scalar floating-point code values and concatenating the dequantized sub-model parameters according to their order in the candidate model parameters, the dequantized model parameters are obtained. This maintains a deterministic and ordered correspondence between the window-level code value sequence and the matrix space position, ensuring that the element arrangement structure of the dequantized weight matrix is isomorphic to the original weight matrix, guaranteeing the index correctness and computational compatibility of subsequent matrix multiplication calculations.
[0101] According to embodiments of this disclosure, dequantizing sub-model parameters based on candidate quantization parameters corresponding to candidate model parameters to obtain dequantized sub-model parameters may include the following operations: transforming sub-model parameters using candidate quantization parameters to obtain intermediate parameters; extracting intermediate parameters into multiple byte values and summing them to aggregate information distributed across different bytes into a Gaussian distribution to obtain aggregated parameters; and performing affine normalization on the aggregated parameters according to the distribution information corresponding to the network layer to obtain dequantized sub-model parameters.
[0102] The intermediate parameters are unsigned integers obtained by modular multiplication and XOR high-bit feedback transformation of the sub-model parameters. The high and low bit information of the input window has been fully mixed in each bit of the intermediate parameters. The high bits are fed back to the low byte region through a shifted XOR operation, so that subsequent summation of the low bytes is sufficient to extract the approximately Gaussian distributed statistics without accessing the complete lookup table for the corresponding bits.
[0103] For example, modular multiplication can refer to multiplying a zero-extended sub-model parameter x (with a preset number of bits) with an odd-numbered multiplication parameter a from the candidate quantization parameters, and then taking the result modulo 2³² to obtain x1 = (a·x) mod 2³². Modulo multiplication is used to spread the high and low bit information of x throughout the entire 32-bit result space through multiplication. XOR high-bit feedback can refer to performing a bitwise XOR operation on x1 with its logically right-shifted value (by s bits) to obtain the intermediate parameter x2 = x1 XOR (x1 >> s). XOR shift is used to feed back the high-bit information of x1 to the low-bit byte region, causing information originally only distributed in the high bits to be folded into the low-bit bytes to be extracted later.
[0104] Multiple byte values are obtained by splitting the intermediate parameter into multiple independent unsigned integer values according to byte boundaries (e.g., in groups of 8 bits). Since the XOR shift operation has fed back the high-order bits to the low-order bits, each of these multiple byte values carries mixed information from different bit ranges. Based on this, by summing the multiple byte values, an aggregate parameter that approximately follows a Gaussian distribution is obtained.
[0105] For example, the 32-bit intermediate parameter x2 can be split into four byte values along 8-bit boundaries: b0 = x2 & 255 (bits 0-7), b1 = (x2 >> 8) & 255 (bits 8-15), b2 = (x2 >> 16) & 255 (bits 16-23), and b3 = (x2 >> 24) & 255 (bits 24-31). Adding these four byte values directly yields the aggregated parameter u = b0 + b1 + b2 + b3. Since the XOR shift in operation one has fed back the high-order information to the low-order bytes, these four byte values are not simply independent intervals, but each carries mixed high- and low-order information. Therefore, according to the Central Limit Theorem, the sum of multiple relatively independent random variables tends to have a Gaussian distribution, allowing information originally distributed in a 32-bit discrete space to be aggregated into an approximately Gaussian scalar through byte summation.
[0106] The distribution information corresponding to each network layer is an affine statistical constant determined independently for each network layer. This constant can be used to linearly map the aggregated parameters from their original value range to a standard numerical range that matches the distribution of the first model parameters. The affine statistical constant can include the mean shift and scaling factor for that network layer. This distribution information is determined during the offline quantization phase by statistically matching the distribution of the aggregated parameters with the transformed weight distribution. After deployment, it is embedded as a layer constant into the inverse quantization kernel.
[0107] Affine normalization is an operation that maps aggregated parameters to dequantized submodel parameters by performing a linear transformation based on distribution information. In one embodiment, affine normalization may include a translation process and a scaling process: adjusting the distribution center to near zero based on a mean translation amount; and adjusting the distribution dispersion to the order of magnitude based on a scaling factor.
[0108] For example, the distribution information (μ_l, σ_l) bound to the current network layer can be obtained, and an affine transformation y=(u-μ_l) / σ_l can be performed on the aggregation parameter u. The center position of the aggregation parameter is adjusted to near zero by the translation operation (u-μ_l), and the dispersion of the aggregation parameter is adjusted to the order of unit variance by the scaling operation (divided by σ_l), so that the distribution of the output inverse quantized sub-model parameter y matches the distribution of the first model parameter.
[0109] In one embodiment, the process of dequantizing the sub-model parameters according to the candidate quantization parameters corresponding to the candidate model parameters to obtain the dequantized sub-model parameters can be shown by the following formulas (2) to (6).
[0110] (2)
[0111] (3)
[0112] (4)
[0113] (5)
[0114] (6)
[0115] in, Characterizing sub-model parameters, Characterizing the multiplication parameter in the candidate quantization parameters, The shift parameter characterizes the candidate quantization parameters. Characterizing intermediate parameters, Characterizing aggregation parameters, The translation components that represent the distribution information corresponding to the network layers. The scaling component represents the distribution information corresponding to the network layer. The parameters of the dequantized sub-model are represented by XOR, which represents bitwise XOR, >> represents logical right shift, and &255 represents truncating a byte.
[0116] In the embodiments of this disclosure, intermediate parameters are obtained by transforming the sub-model parameters using candidate quantization parameters. A combination of modular multiplication and XOR shift operations is used to fully diffuse and mix the high and low bit information of the input window into the entire bit space. This allows for the subsequent extraction of only a portion of the bits to recover the global information features, avoiding the memory access latency of reading the complete codebook or lookup table at runtime. By extracting the intermediate parameters into multiple byte values and summing them, the information distributed across different bytes is aggregated into a Gaussian distribution to obtain aggregate parameters. This allows the information originally discretely distributed in each bit integer space to naturally coalesce into an approximately Gaussian scalar through the statistical effect of the central limit theorem. Thus, an intermediate representation compatible with the Gaussian assumption is obtained without the need for explicit probability modeling or distribution fitting.
[0117] Based on this, the dequantized sub-model parameters are obtained by affine normalization of the aggregation parameters according to the distribution information corresponding to the network layers. This allows each network layer to apply translation and scaling to the aggregation parameters according to its unique weight statistical characteristics, ensuring that the distribution of dequantized code values of different layers is precisely aligned with the weight distribution of the corresponding layer. This avoids the problem of partial layer distribution mismatch caused by uniform normalization parameters, thus achieving layer-by-layer adaptation of dequantized code values to the original model parameters at the layer-customized granularity.
[0118] According to embodiments of this disclosure, determining a target model parameter from multiple candidate model parameters with the goal of minimizing the difference between the inverse quantization model parameters and the first model parameters may include the following operations: determining a first difference between the inverse quantization model parameters and the first model parameters based on reference sensitivity information; and determining the target model parameter from multiple candidate model parameters with the goal of minimizing the first difference.
[0119] The first difference is a quantization error metric obtained by directionally weighting the element-wise deviations between the dequantized model parameters W_i and the first model parameters W_r using the reference sensitivity information H_r as the weighting matrix. It should be noted that the first difference does not treat all reconstruction deviations equally. Instead, through quadratic weighting of H_r, deviations in sensitive directions that have a significant impact on the model output are amplified, while deviations in insensitive directions are suppressed. This reflects not only the element-wise accuracy of the reconstruction but also the actual fidelity of the reconstruction weights to the network layer output.
[0120] In one embodiment, the first difference can be determined using the following formulas (7) and (8).
[0121] (7)
[0122] (8)
[0123] in, A set representing candidate quantization parameters. Characterize candidate quantization parameters, Characterize the parameters of the candidate model. Characterizing the parameters of the inverse quantization model, Characterizing the first difference, Characterizing the parameters of the first model, Characterizes reference sensitivity information.
[0124] For example, for each candidate model parameter Z_i obtained by inverse quantization, the weighted reconstruction error between it and the first model parameter W_r can be calculated using the reference sensitivity information H_r as a weighted benchmark. Specifically, the element-wise deviation matrix Δ_i = W_r - W_i can be calculated; the weighted quadratic form of H_r can be calculated: multiplying Δ_i^T on the left to obtain the row vector, multiplying it on the right to obtain the weighted row vector, and then multiplying it on the right to each column of Δ_i and taking the trace, i.e., Tr(Δ_i^T·H_r·Δ_i); normalizing the denominator using the norm of W_r under H_r Tr(W_r^T·H_r·W_r) to obtain the first difference J_i.
[0125] It should be noted that since the reference sensitivity information H_r is a symmetric positive semi-definite matrix, it can be decomposed into H_r = L·L^T. That is, the Frobenius norm ratio can be calculated by left-multiplying Δ_i and W_r by L^T.
[0126] After obtaining the first difference J_i corresponding to each candidate model parameter, all candidate difference values can be sorted and compared. Using the minimization of difference as the criterion, the candidate model parameter Z_i with the smallest J_i value is selected as the target model parameter. Since different network layers have different weight statistical characteristics, different target quantization parameters and target model parameters may be determined, allowing each layer to obtain the most suitable quantization configuration for itself.
[0127] In the embodiments of this disclosure, a first difference between the dequantized model parameters and the first model parameters is determined based on reference sensitivity information. This ensures that the difference measurement does not assign equal weights to all weight positions, but rather uses reference sensitivity information to directionally weight the quantization error, thereby ensuring that the magnitude of the first difference directly reflects the actual impact of the reconstruction weights on the fidelity of the network layer output. Based on this, the target model parameters are determined from multiple candidate model parameters with the goal of minimizing the first difference. This allows different network layers to make independent decisions based on their own weight distribution and sensitivity information, converging to their respective optimal compression configurations at each layer. This achieves a balance between accuracy loss and compression ratio at the overall model level.
[0128] Figure 4 The illustration shows an example of a process for determining a target model parameter from a plurality of candidate model parameters with the aim of minimizing the quantization error between a first model parameter and candidate model parameters, according to an embodiment of the present disclosure.
[0129] like Figure 4As shown, in embodiment 400 for determining target model parameters, candidate model parameters 410 can be truncated in a sliding manner according to a preset window width to obtain multiple sub-model parameters. For example, taking a preset window width of 12, sub-model parameters 411 and 412 can be truncated from candidate model parameters 410. Adjacent sub-model parameters 411 and 412 have a preset number of identical parameters. For example, the preset number can be 10.
[0130] For each sub-model parameter, the sub-model parameter can be dequantized according to the candidate quantization parameter corresponding to the candidate model parameter 410 to obtain the dequantized sub-model parameter. For example, the sub-model parameter 411 is dequantized according to the candidate quantization parameter corresponding to the candidate model parameter 410 to obtain the dequantized sub-model parameter 421; the sub-model parameter 412 is dequantized according to the candidate quantization parameter corresponding to the candidate model parameter 410 to obtain the dequantized sub-model parameter 422.
[0131] Based on this, the dequantization sub-model parameters 421 and 422 can be concatenated according to their order in the candidate model parameters 410 to obtain the dequantization model parameters 430.
[0132] Based on the distribution information obtained by processing sample data 450 using network layer 440, intermediate sensitivity information 460 is determined; the intermediate sensitivity information 460 is subjected to incoherent processing to make the parameter distribution approach a Gaussian distribution, thereby obtaining reference sensitivity information 461.
[0133] Based on the reference sensitivity information 461, a first difference 480 is determined between the inverse quantization model parameter 430 and the first model parameter 470; with the goal of minimizing the first difference 480, a target model parameter 490 is determined from multiple candidate model parameters.
[0134] According to embodiments of this disclosure, the channel dimension includes an input channel and an output channel; operation S230 may include the following operations: using each row of the target model parameters as the output channel and each column as the input channel, determining the reconstructed model parameters based on the output channel and the input vector, and the input channel and the output vector; using sample data, determining a second difference between the output of the network layer based on the original weight parameters and the output based on the reconstructed model parameters; adjusting the input vector and the output vector with the goal of minimizing the second difference to obtain channel compensation parameters.
[0135] In a Transformer linear layer, the row direction of the target model parameters corresponds to the output channels, meaning each row vector corresponds to the connection weights between an output neuron and all input channels; the column direction of the target model parameters corresponds to the input channels, meaning each column vector corresponds to the projection weights from an input neuron to all output neurons. Channel dimension compensation refers to applying scaling corrections along both the row and column directions.
[0136] The input channel is the column dimension of the network layer weight matrix, corresponding to each component of the input vector. All elements in the j-th column of the target model parameters collectively determine the contribution weight of the j-th input component to all output components. The input channel can be scaled using the corresponding element S_in[j] of the input vector S_in; that is, this scaling factor uniformly adjusts the influence intensity of the j-th input component on the overall output, thereby correcting the systematic bias generated on this input channel during quantization. The input vector is a learnable scaling parameter vector that corresponds one-to-one with the input channels, and its length can be the number of input channels, d_in.
[0137] The output channel represents the row dimension of the network layer weight matrix, corresponding to each component of the output vector. All elements in the i-th row of the target model parameters collectively determine the contribution weight of all input components to the i-th output component. The output channel can be scaled using the corresponding element S_out[i] of the output vector S_out; that is, this scaling factor uniformly adjusts the aggregate response intensity received by the i-th output component from all inputs, thereby correcting the systematic bias generated on this output channel during quantization. The output vector is a learnable scaling parameter vector that corresponds one-to-one with each output channel, and its length can be the number of output channels, d_out.
[0138] For example, we can first determine the output channels (d_out) and input channels (d_in) corresponding to the row direction of the target model parameters; construct diagonal matrices diag(S_out) and diag(S_in), where S_out∈R^(d_out) and S_in∈R^(d_in) are the output vector and input vector, respectively; and determine the reconstructed model parameters through matrix multiplication: W_tilde=diag(S_out)·W·diag(S_in). Specifically, we multiply all elements in the i-th row of the target model parameters by S_out[i] to scale the output channels; and we multiply all elements in the j-th column of the target model parameters by S_in[j] to scale the input channels.
[0139] The reconstructed model parameters are the final model parameters obtained after channel compensation of the target model parameters. Each element of the reconstructed model parameters, W_tilde[i,j] = S_out[i]·W[i,j]·S_in[j], means that each element of the original target model parameters is jointly adjusted by the output scaling factor of its row and the input scaling factor of its column.
[0140] The second difference occurs during the channel compensation stage, where the network layer measures the deviation between its forward output f_l(X; W_l) based on the original model parameters W_l and its forward output f_l(X; W_tilde_l) based on the reconstructed model parameters W_tilde. For example, sample data X can be fed into the network layer batch by batch, with forward computation performed using both the original weight parameters W_l and the reconstructed model parameters W_tilde. For each batch of samples, the network layer calculates the original output Y_orig=f_l(X; W_l) using W_l and the reconstructed output Y_recon=f_l(X; W_tilde) using W_tilde; then it calculates the squared difference of the Frobenius norm between the two outputs ||Y_orig-Y_recon||²_F, and takes the expected or average value over the entire sample dataset to obtain the second difference.
[0141] In one embodiment, the second difference can be determined using the following formulas (9) and (10).
[0142] (9)
[0143] (10)
[0144] in, Reconstructing model parameters Characterize the parameters of the target model. Represents the output vector. Representing the input vector, Characterizing sample data, Characterize the sample dataset.
[0145] After obtaining the second difference, the second difference L can be used as the loss function, and the input vector S_in and output vector S_out as the optimizable variables. Initially, S_in and S_out are usually set to vectors of all 1s. The gradient optimizer is used for iterative updates, that is, L is calculated by forward propagation and backward propagation for each batch of samples. L / S_in[j] and L / S_out[i] updates each element of the vector along the negative gradient direction. The optimization continues until L converges or the preset number of iterations is reached. At this point, S_in and S_out are the channel compensation parameters.
[0146] In the embodiments of this disclosure, the channel dimension is divided into input channels and output channels, with each row of the target model parameters serving as the output channel and each column as the input channel. This ensures that the application range of channel compensation has a clear structural belonging in the row and column directions of the weight matrix, thereby guaranteeing the spatial structure of the compensation operation. The reconstructed model parameters are determined based on the output channels and input vectors, and the input channels and output vectors. The channel-level correction is structurally integrated into the weight representation in the form of matrix multiplication through the left and right products of two diagonal scaling matrices. This allows the compensated weight matrix to maintain the original relative relationships between the elements of the target model parameters while only adjusting the magnitude at the channel granularity.
[0147] Building upon this, a second difference is determined using sample data between the network layer's output based on the original weight parameters and its output based on the reconstructed model parameters. This second difference reflects the output bias introduced by the quantized weights in the actual forward computation. The input and output vectors are adjusted to minimize this second difference to obtain channel compensation parameters. This allows the learning of the compensation parameters to directly use layer output consistency as the optimization criterion, adaptively identifying and specifically correcting channel-level systematic biases introduced by quantization from the data. This improves the forward computation fidelity of the network layer while freezing the discrete quantization structure.
[0148] Figure 5 The illustration shows an example of a process for obtaining channel compensation parameters by performing channel dimension compensation on target model parameters based on sample data according to an embodiment of the present disclosure.
[0149] like Figure 5 As shown in Example 500 of obtaining channel compensation parameters, the method for determining channel compensation parameters is explained using the example of channel dimensions including input channels and output channels.
[0150] Using each row of the target model parameters as the output channel and each column of the target model parameters as the input channel, the reconstructed model parameters 531 are determined based on the output channel and input vector 540, and the input channel and output vector 550.
[0151] Using sample data 510, a second difference 560 is determined between the output of network layer 520 based on the original weight parameters 521 and the output of network layer 530 based on the reconstructed model parameters 531; with the goal of minimizing the second difference 560, the input vector 540 and the output vector 550 are adjusted to obtain the channel compensation parameters.
[0152] According to embodiments of this disclosure, the above-described large model task execution method 200 may further include the following operations: using the model to be quantized as the teacher model and the model corresponding to the target model parameters as the student model, determining a third difference between the output of the teacher model and the output of the student model; and updating the input vector and the output vector with the goal of minimizing the third difference while keeping the target model parameters and the target quantization parameters unchanged.
[0153] The teacher model f_T is a reference model that provides supervision signals within the knowledge distillation framework. The student model f_Q is a lightweight model that receives supervision signals from the teacher model within the knowledge distillation framework and improves its performance by mimicking the teacher's output. In one embodiment, the teacher model f_T can be the model to be quantized itself, with its forward output serving as the soft target for the student model f_Q to learn; the student model f_Q can be a quantized low-bit model.
[0154] It should be noted that during the distillation process, the student model only unlocks and updates the affine statistics and channel compensation parameters of each network layer, while the target model parameters and target quantization parameters remain frozen.
[0155] For example, the model to be quantized can be loaded as the teacher model f_T, and the quantized model with embedded target model parameters and initial channel compensation parameters can be loaded as the student model f_Q. Unlabeled sample data x are simultaneously input into the teacher model and the student model in batches, and the two complete the full forward propagation to obtain their respective final outputs. The output f_T(x) of the teacher model is usually the logits on the vocabulary or the probability distribution after softmax. The output f_Q(x) of the student model is the prediction distribution in the same format. The third difference between the two output distributions is calculated.
[0156] The third difference is a measure of the distributional difference between the output of the teacher model f_T and the output of the student model f_Q during the end-to-end distillation stage. The method for determining the third difference can be configured according to actual business needs and is not limited here. For example, the third difference can be expressed as Kullback-Leibler divergence to measure the information loss of the student model's output distribution relative to the teacher model's output distribution.
[0157] In one embodiment, the third difference can be determined using the following formula (11).
[0158] (11)
[0159] in, Teacher representation model Student representation model Characterize the parameters of the target model. Characterizing the target quantization parameters, and Characterize updatable affine statistics. Characterizing channel compensation parameters, Characterize the third difference.
[0160] It should be noted that the target model parameters and target quantization parameters of each network layer do not participate in gradient calculation and parameter updates; only the input vector S_in, output vector S_out, and affine statistics of each network layer are used as trainable parameters. Based on this, the third difference is used as the loss function, and a gradient optimizer is employed for iterative updates. For example, L is calculated during forward propagation and backward propagation for each batch of samples. L / S_in、 L / S_out、 L / μ_l、 L / σ_l, these continuous parameters are adjusted according to the gradient direction.
[0161] In the embodiments of this disclosure, the model to be quantized is used as the teacher model, and the model corresponding to the parameters of the target model is used as the student model. This ensures that the supervision signal for distillation originates from the full-precision version of the model itself, and the supervision target and the task space of the quantized model are completely isomorphic, avoiding the distribution bias that may be introduced by cross-model distillation. By determining the third difference between the output of the teacher model and the output of the student model, which directly reflects the magnitude of the impact of quantization on the final prediction behavior of the model, a global loss function aligned with the task objective is provided for the optimization of continuous parameters.
[0162] Based on this, while keeping the target model parameters and target quantization parameters unchanged, the input and output vectors are updated with the goal of minimizing the third difference. This ensures that the target model parameters and target quantization parameters remain absolutely stable throughout the process. The prediction distribution drift introduced by quantization is absorbed and compensated only by adjusting the channel scaling parameters and affine statistics. Thus, end-to-end reuse of the overall prediction accuracy of the model is achieved while retaining the constraints of table-free deployment and bitstream compatibility.
[0163] The above are merely exemplary embodiments, but are not limited thereto. Other large model task execution methods known in the art may also be included, as long as they can ensure the accurate reuse of the quantized model.
[0164] Figure 6 A flowchart illustrating a large model task processing method according to an embodiment of the present disclosure is shown schematically.
[0165] like Figure 6 As shown, the large model task processing method 600 may include operations S610~S620.
[0166] In operation S610, the quantization result is obtained from the storage unit according to the model identifier indicated by the task to be processed, wherein the quantization result is obtained using the large model task execution method.
[0167] When operating the S620, the target model is used to process the task to be processed and obtain the task processing result. The target model is obtained by configuring the model used for the task to be processed using the quantization result.
[0168] A pending task is a specific processing request received by the inference service that needs to be executed by the large language model. A pending task may include a model identifier and input information. The task specifies which model should be used to complete the task, and the input information includes text generation suggestions, dialogue context, code completion prefixes, etc. The initiator of a pending task can be a user terminal, upstream microservice, batch processing scheduler, or automated pipeline, etc., without limitation here.
[0169] It should be noted that the type of task to be processed can be configured according to actual business needs, and is not limited here. For example, the task to be processed can be a text generation task, a multi-turn dialogue response task, a code auto-completion task, a text summarization task, a machine translation task, a sentiment analysis task, etc.
[0170] The model used for the task to be processed is a model architecture registered in the inference service, corresponding to the model identifier specified by the task. The model architecture can include network structure, number of layers, dimensions, etc., as described in the model configuration file config.json. The model architecture defines structural meta-information such as the number of Transformer Blocks, number of attention heads, hidden state dimensions, and activation function types, but has not yet attached specific weight parameters.
[0171] The target model is a low-bit quantization deployment model that can be directly used for inference computation, obtained by configuring the original model architecture using quantization results obtained from the storage unit. The configuration process may include: replacing the original model parameters of each network layer with the target model parameters for the corresponding network layer; replacing weight reads in the forward computation with calls to the General Matrix Multiply (GEMM) function; and embedding channel compensation parameters and affine statistics into the constant region of the inverse quantization kernel. During inference, the target model does not require access to runtime lookup tables or codebooks; all code values are generated on the fly using integer calculations according to the aforementioned preset method.
[0172] The task processing result is the output generated by the target model after performing forward inference on the input data of the task to be processed. The form of the task processing result depends on the type of task to be processed. For example, for text generation tasks, the task processing result can be a sequence of generated tokens and their corresponding logits; for multi-turn dialogue tasks, the task processing result can be a dialogue increment containing response content; for code completion tasks, the task processing result can be a completed code snippet, etc., without limitation.
[0173] In the embodiments of this disclosure, the quantization result is retrieved from the storage unit according to the model identifier indicated by the task to be processed. This allows the retrieval and loading of quantization parameters to be completed automatically with the model identifier as the unique index key, avoiding the risk of mismatch due to manual specification of quantization file paths or version numbers, and ensuring the correspondence between the model required by the task and the loaded compressed parameters. Since the quantization result is obtained using the large model task execution method, the quantization result contains all decoding parameters of the lookup table-free code value generator and the structured organization of the compressed bitstream. The deployment end can use it directly without any additional codebook compilation or lookup table generation steps, thus seamlessly connecting the offline quantization output with the online inference input.
[0174] Based on this, the task processing results are obtained by using the target model to process the task to be processed. This allows the inference computation to be executed by the configured low-bit quantization model. The weights of each layer are reconstructed in real time through sliding window truncation and integer arithmetic, avoiding the bandwidth bottleneck of repeatedly reading the complete floating-point weights from the storage unit. While maintaining the semantic consistency of task processing, the inference latency and memory consumption are reduced, thereby improving the efficiency of task processing.
[0175] The above are merely exemplary embodiments, but are not limited thereto. Other task processing methods known in the art may also be included, as long as they can improve the efficiency of task processing.
[0176] Figure 7 A block diagram of a task execution apparatus according to an embodiment of the present disclosure is shown schematically.
[0177] like Figure 7 As shown, the task execution device 700 may include a quantization module 710, a determination module 720, and a compensation module 730.
[0178] The quantization module 710 is used to quantize the first model parameters of the network layer of the model to be quantized indicated by the quantization task according to multiple candidate quantization parameters, so as to obtain multiple candidate model parameters.
[0179] The determination module 720 is used to determine the target model parameters from multiple candidate model parameters with the goal of minimizing the quantization error between the first model parameters and the candidate model parameters.
[0180] The compensation module 730 is used to perform channel dimension compensation on the target model parameters based on sample data to obtain channel compensation parameters, and store the target model parameters, the target quantization parameters corresponding to the target model parameters, and the channel compensation parameters as quantization results corresponding to the model identifier of the model to be quantized in the storage unit.
[0181] According to embodiments of this disclosure, the determining module 720 may include an inverse quantization submodule and a first determining submodule.
[0182] The dequantization submodule is used to dequantize the candidate model parameters using a sliding window method to obtain the dequantized model parameters.
[0183] The first determination submodule is used to determine the target model parameters from multiple candidate model parameters with the goal of minimizing the difference between the inverse quantization model parameters and the first model parameters.
[0184] According to embodiments of this disclosure, the dequantization submodule may include a truncation unit, a dequantization unit, and a splicing unit.
[0185] The truncating unit is used to truncate the candidate model parameters in a sliding manner according to a preset window width to obtain multiple sub-model parameters, wherein there is a preset number of identical parameters among adjacent sub-model parameters.
[0186] The dequantization unit is used to dequantize each sub-model parameter according to the candidate quantization parameter corresponding to the candidate model parameter, so as to obtain the dequantized sub-model parameter.
[0187] The splicing unit is used to splice the parameters of each inverse sub-model according to the order in which the parameters of each sub-model are arranged in the candidate model parameters, so as to obtain the inverse model parameters.
[0188] According to embodiments of this disclosure, the inverse quantization unit may include a transformation subunit, a aggregation subunit, and an affine subunit.
[0189] The transformation subunit is used to transform the sub-model parameters using candidate quantization parameters to obtain intermediate parameters.
[0190] The aggregation subunit is used to extract intermediate parameters into multiple byte values and sum them to aggregate information distributed on different bytes into a Gaussian distribution, thus obtaining the aggregation parameters.
[0191] Affine sub-units are used to perform affine normalization on the aggregate parameters according to the distribution information corresponding to the network layers, so as to obtain the inverse quantized sub-model parameters.
[0192] According to embodiments of this disclosure, the task execution device 700 may further include an acquisition module.
[0193] The module is used to process sample data using network layers to obtain reference sensitivity information, which characterizes the degree of influence of quantization error on the output of the network layer.
[0194] According to embodiments of this disclosure, the first determining submodule may include a first determining unit and a second determining unit.
[0195] The first determining unit is used to determine the first difference between the inverse quantization model parameters and the first model parameters based on reference sensitivity information.
[0196] The second determining unit is used to determine the target model parameters from multiple candidate model parameters with the goal of minimizing the first difference.
[0197] According to embodiments of this disclosure, the obtaining module may include a second determining submodule and a processing submodule.
[0198] The second determination submodule is used to determine intermediate sensitivity information based on the distribution information obtained by processing sample data using network layers.
[0199] The processing submodule is used to perform incoherent processing on the intermediate sensitivity information to make the parameter distribution approximate a Gaussian distribution and obtain the reference sensitivity information.
[0200] According to embodiments of this disclosure, the channel dimension includes an input channel and an output channel; the compensation module 730 may include a third determining submodule, a fourth determining submodule, and an adjustment submodule.
[0201] The third determination submodule is used to determine the reconstructed model parameters based on the output channels and input vectors, and the input channels and output vectors, using each row of the target model parameters as the output channel and each column as the input channel.
[0202] The fourth determination submodule is used to determine a second difference between the network layer's output based on the original weight parameters and its output based on the reconstructed model parameters, using sample data.
[0203] The adjustment submodule is used to adjust the input and output vectors to obtain the channel compensation parameters with the goal of minimizing the second difference.
[0204] According to embodiments of this disclosure, the compensation module 730 may further include a fifth determining submodule and an updating submodule.
[0205] The fifth determination submodule is used to determine the third difference between the output of the teacher model and the output of the student model, using the model to be quantized as the teacher model and the model corresponding to the parameters of the target model as the student model.
[0206] The update submodule is used to update the input and output vectors with the goal of minimizing the third difference, while keeping the target model parameters and target quantization parameters unchanged.
[0207] Figure 8 A block diagram of a task processing apparatus according to an embodiment of the present disclosure is shown schematically.
[0208] like Figure 8 As shown, the task processing device 800 may include an acquisition module 810 and a processing module 820.
[0209] The acquisition module 810 is used to acquire quantization results from the storage unit according to the model identifier indicated by the task to be processed, wherein the quantization results are obtained using the task execution device.
[0210] The processing module 820 is used to process the task to be processed using the target model and obtain the task processing result. The target model is obtained by configuring the model used for the task to be processed using the quantization result.
[0211] Figure 9 A schematic diagram illustrating the structure of a large-model-based intelligent agent according to an embodiment of the present disclosure is shown.
[0212] In embodiments of this disclosure, the von Neumann architecture in modern computer theory is inspired, such as... Figure 9 As shown, the AI agent 900 may include five core modules: input module 910, processing module 920 and output module 930.
[0213] The input module 910 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment) and converting it into a format that the AI agent 900 can understand and process. The input module 910 is the primary link for the AI agent 900 to interact with the outside world, enabling it to efficiently and accurately acquire necessary "sensory" information and respond to it. In the example, the input module 910 can input the quantization task, training task, and task to be processed described above.
[0214] In embodiments of this disclosure, the processing module 920 may include a control module 921, a storage module 922, and a computation module 923. The processing module 920 is configured to determine a target task based on the input information received by the input module 910, determine a target large model based on the target task, and obtain output information by calling the corresponding method executed by the target large model.
[0215] The control module 921 is the core support for the AI agent 900's ability to handle complex tasks. The control module 921 can execute the large model task execution method and the large model task processing method described above.
[0216] In the example, the control module 921 will continuously interact with the storage module 922, the arithmetic module 923, and / or the output module 930 during operation. However, it should be noted that in the embodiments of this disclosure, the control module 921 initiates communication with the storage module 922, the arithmetic module 923, and / or the output module 930 as a single initiator, and there is no communication coupling between the storage module 922, the arithmetic module 923, and the output module 930.
[0217] In the example, the performance of the control module 921 is closely related to the large model on which the AI agent 900 is based. To fully leverage the capabilities of the large language model, the internal structure of the control module 921 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.
[0218] The storage module 922 is responsible for storing the generated target sample set. The original model parameters, target model parameters, and trained model parameters, as mentioned earlier, can be included in the storage module 922.
[0219] In the example, after receiving the quantization task, the AI agent 900 can retrieve the original model parameters from the storage module 922 and trigger the task execution process to obtain the target model parameters, which are then fed back to the control module 921. The control module 921 can then pass the returned target model parameters to the output module 930.
[0220] The computation module 923 can be viewed as a predefined tool library. Tools for quantization, as described above, can be included in the computation module 923.
[0221] In the example, when the AI agent 900 needs to perform a linear transformation on features, it can invoke relevant tools from the computation module 923 and feed them back to the control module 921. The control module 921 can then use the fed-back tools to process the relevant information. It's understandable that while large language models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks without any tools is limited. When the AI agent 900 is given the ability to invoke tools, it can perform tasks such as quantization using tools designed for quantization.
[0222] The output module 930 can output the target model parameters and task processing results described above.
[0223] The AI agent 900 according to embodiments of this disclosure can simply and effectively improve the level of intelligence, as well as enhance flexibility and versatility.
[0224] Figure 10The diagram schematically illustrates an electronic device suitable for implementing a large-scale task execution method and a large-scale task processing method according to embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0225] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded into random access memory (RAM) 1003 from storage unit 1008. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0226] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0227] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as large model task execution methods and large model task processing methods. For example, in some embodiments, the large model task execution methods and large model task processing methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the large model task execution methods and large model task processing methods described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute a large model task execution method or a large model task processing method by any other suitable means (e.g., by means of firmware).
[0228] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0229] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0230] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0231] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0232] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0233] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0234] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0235] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for executing large-scale model tasks, comprising: Based on multiple candidate quantization parameters, the first model parameters of the network layer of the model to be quantized indicated by the quantization task are quantized to obtain multiple candidate model parameters. With the goal of minimizing the quantization error between the first model parameters and the candidate model parameters, a target model parameter is determined from among the multiple candidate model parameters; as well as Channel dimension compensation is performed on the target model parameters based on sample data to obtain channel compensation parameters. The target model parameters, the target quantization parameters corresponding to the target model parameters, and the channel compensation parameters are stored in the storage unit as quantization results corresponding to the model identifier of the model to be quantized.
2. The method according to claim 1, wherein, The step of determining the target model parameter from among the multiple candidate model parameters with the goal of minimizing the quantization error between the first model parameter and the candidate model parameter includes: The candidate model parameters are dequantized using a sliding window method to obtain the dequantized model parameters; and With the goal of minimizing the difference between the inverse quantization model parameters and the first model parameters, the target model parameter is determined from among the multiple candidate model parameters.
3. The method according to claim 2, wherein, The step of dequantizing the candidate model parameters using a sliding window method to obtain dequantized model parameters includes: According to the preset window width, the candidate model parameters are truncated in a sliding manner to obtain multiple sub-model parameters, wherein there are a preset number of identical parameters between adjacent sub-model parameters; For each of the sub-model parameters, the sub-model parameter is dequantized according to the candidate quantization parameter corresponding to the candidate model parameter to obtain the dequantized sub-model parameter; and The inverse quantization sub-model parameters are concatenated according to their order of arrangement in the candidate model parameters to obtain the inverse quantization model parameters.
4. The method according to claim 3, wherein, The step of dequantizing the sub-model parameters according to the candidate quantization parameters corresponding to the candidate model parameters to obtain dequantized sub-model parameters includes: The candidate quantization parameters are used to transform the sub-model parameters to obtain intermediate parameters; The intermediate parameters are extracted byte by byte and summed to aggregate the information distributed across different bytes into a Gaussian distribution, thus obtaining the aggregation parameters; and The aggregation parameters are affinely normalized according to the distribution information corresponding to the network layer to obtain the inverse quantization sub-model parameters.
5. The method according to claim 2, further comprising: The sample data is processed using the network layer to obtain reference sensitivity information, wherein the reference sensitivity information characterizes the degree of influence of quantization error on the output of the network layer; The step of determining the target model parameter from among a plurality of candidate model parameters with the objective of minimizing the difference between the inverse quantization model parameters and the first model parameters includes: Based on the reference sensitivity information, a first difference between the inverse quantization model parameters and the first model parameters is determined; and With the goal of minimizing the first difference, the target model parameter is determined from among the multiple candidate model parameters.
6. The method according to claim 5, wherein, The process of processing the sample data using the network layer to obtain reference sensitivity information includes: Based on the distribution information obtained by processing the sample data using the network layers, intermediate sensitivity information is determined; and The intermediate sensitivity information is subjected to incoherent processing to make the parameter distribution approximate a Gaussian distribution, thereby obtaining the reference sensitivity information.
7. The method according to any one of claims 1 to 6, wherein, The channel dimension includes input channels and output channels; The channel dimension compensation of the target model parameters based on sample data is used to obtain channel compensation parameters, including: Using each row of the target model parameters as the output channel and each column as the input channel, the reconstructed model parameters are determined based on the output channel and the input vector, and the input channel and the output vector. Using the sample data, determine a second difference between the network layer's output based on the original weight parameters and its output based on the reconstructed model parameters; and With the goal of minimizing the second difference, the input vector and the output vector are adjusted to obtain the channel compensation parameters.
8. The method according to claim 7, further comprising: Using the model to be quantified as the teacher model and the model corresponding to the parameters of the target model as the student model, a third difference between the output of the teacher model and the output of the student model is determined. as well as While keeping the target model parameters and the target quantization parameters unchanged, the input vector and the output vector are updated with the goal of minimizing the third difference.
9. A method for processing large model tasks, comprising: Based on the model identifier indicated by the task to be processed, the quantization result is obtained from the storage unit, wherein the quantization result is obtained using the method described in any one of claims 1 to 8; and The target model is used to process the task to be processed to obtain the task processing result. The target model is obtained by configuring the model used for the task to be processed using the quantization result.
10. A large-scale model task execution device, comprising: The quantization module is used to quantize the first model parameters of the network layer of the model to be quantized indicated by the quantization task according to multiple candidate quantization parameters, so as to obtain multiple candidate model parameters. The determining module is configured to determine the target model parameter from among the multiple candidate model parameters with the goal of minimizing the quantization error between the first model parameter and the candidate model parameter; as well as The compensation module is used to perform channel dimension compensation on the target model parameters based on sample data to obtain channel compensation parameters, and store the target model parameters, the target quantization parameters corresponding to the target model parameters, and the channel compensation parameters as quantization results corresponding to the model identifier of the model to be quantized in the storage unit.
11. A large-scale model task processing device, comprising: The acquisition module is configured to acquire a quantization result from a storage unit based on a model identifier indicated by a task to be processed, wherein the quantization result is obtained using the apparatus of claim 10; and The processing module is used to process the task to be processed using the target model to obtain the task processing result, wherein the target model is obtained by configuring the model used for the task to be processed using the quantization result.
12. An intelligent agent based on a large model, comprising: The input module is used to receive input information; The processing module is configured to determine a target task based on the input information received by the input module, determine a target large model based on the target task, and execute the method described in claims 1 to 9 by calling the target large model to obtain output information; as well as An output module is used to output the output information obtained by the processing module.
13. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 9.
14. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 9.
15. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 9.