A model processing method and a model processing apparatus

CN122735809APending Publication Date: 2026-09-11LCFC HEFEI ELECTRONICS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610869850.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

尽管已有研究尝试通过稀疏注意力、分层处理或动态计算来缓解这一问题,但在边缘设备上实现精度与效率的平衡仍面临巨大困难

Benefits of technology

本申请实施例提供的一种模型处理方法,通过获取神经网络模型各层特征通道的激活值并构建激活分布图谱,能够精准刻画各特征通道激活的分布规律与稳定程度。基于激活分布图谱识别对模型预测精度影响显著的关键特征通道,避免轻量化过程中关键特征信息被过度压缩而导致精度下降,通过梯度下降方法确定关键特征通道的缩放因子,基于该缩放因子能够针对性优化权重的数值分布,有效降低量化过程中的舍入误差,提升关键特征通道的量化精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122735809A_ABST
    Figure CN122735809A_ABST
Patent Text Reader

Abstract

The application provides a model processing method and device, the model processing method comprises: obtaining activation values generated by feature channels of each layer of a neural network model, and determining an activation distribution atlas based on the activation values; identifying key feature channels that significantly affect model prediction accuracy based on the activation distribution atlas; determining a scaling factor of the key feature channels by a gradient descent method; performing sparse processing on a correlation score matrix of a query and a key in an attention layer to obtain an attention weight distribution; and performing low-bit quantization on target data of each feature channel based on the scaling factor and the attention weight distribution to obtain a lightweight model. In this way, while reducing model storage and calculation overhead, the loss of quantization accuracy can be effectively suppressed, the resource-limited scenario of an edge device can be adapted, and the balance between lightweight and inference accuracy can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a model processing method and a model processing device. Background Technology

[0002] With the widespread application of artificial intelligence (AI) technologies in fields such as computer vision and natural language processing, deploying complex deep learning models to resource-constrained edge devices (such as smartphones, IoT terminals, and embedded systems) has become a key challenge for the industry. Edge computing environments are typically limited by computing resources, storage space, and energy budgets, while existing mainstream AI models (especially those based on the Transformer architecture) are difficult to run efficiently on these devices due to their massive number of parameters and high computational complexity. In recent years, the industry has proposed various model lightweighting techniques, including but not limited to network pruning, knowledge distillation, low-bit quantization, and attention mechanism optimization, aiming to reduce the model's hardware resource requirements while maintaining its inference accuracy as much as possible.

[0003] As a core component of the Transformer model, the computational cost of the attention mechanism increases quadratically with the length of the input sequence, which is particularly pronounced when processing high-resolution images or long text sequences. Although research has attempted to alleviate this problem through sparse attention, hierarchical processing, or dynamic computation, achieving a balance between accuracy and efficiency on edge devices remains extremely challenging. The traditional Softmax function forcibly assigns weight probabilities to all inputs during attention computation, resulting in non-zero activations even when faced with irrelevant or low-relevance information. During model compression, even if two features are completely unrelated and have extremely low correlation, Softmax will still assign them a tiny probability weight, failing to directly mask invalid information. This results in the absence of true zero values ​​in the attention weight matrix, making effective sparsity difficult to achieve. The non-sparse distribution of attention weights further hinders efficient compression of key parameters during quantization. In low-bit quantization scenarios, a large number of non-zero secondary weights occupy limited numerical representation space, weakening the model's accuracy in representing important features. Ultimately, this leads to a significant decrease in accuracy during inference on edge devices. This problem is further amplified in models such as the visual Transformer due to the complexity of spatial dimensions. Summary of the Invention

[0004] This application addresses the aforementioned technical problems in the existing technology. The purpose of this application is to provide a model processing method and apparatus that, under resource-constrained conditions on edge devices, achieves synergistic optimization of attention mechanism sparsity and adaptive low-bit quantization. This significantly reduces model storage and computational overhead while effectively suppressing quantization accuracy loss, improving model inference efficiency and stability, and meeting the practical deployment needs of edge devices for lightweight, high-precision, and low-latency AI inference.

[0005] According to the first aspect of this application, a model processing method is provided, the model processing method comprising: obtaining activation values ​​generated by feature channels of each layer of a neural network model, and determining an activation distribution map based on the activation values; identifying key feature channels that significantly affect the prediction accuracy of the model based on the activation distribution map; determining a scaling factor for the key feature channels using a gradient descent method, the scaling factor being used to adjust the dynamic range of the weights of each feature channel; performing sparsification processing on the correlation score matrix between the query and the key in the attention layer to obtain an attention weight distribution; and performing low-bit quantization on the target data of each feature channel based on the scaling factor and the attention weight distribution to obtain a lightweight model.

[0006] According to a second aspect of this application, a model processing apparatus is provided, the model processing apparatus including a processor configured to execute the steps of the model processing methods described in various embodiments of this application.

[0007] Compared with the prior art, the beneficial effects of the embodiments of this application are as follows: This application provides a model processing method that, by acquiring the activation values ​​of feature channels in each layer of a neural network model and constructing an activation distribution map, can accurately characterize the distribution pattern and stability of activation in each feature channel. Based on the activation distribution map, key feature channels that significantly affect model prediction accuracy are identified, avoiding excessive compression of key feature information during the quantization process that could lead to decreased accuracy. A scaling factor for key feature channels is determined using a gradient descent method. Based on this scaling factor, the numerical distribution of weights can be optimized in a targeted manner, effectively reducing rounding errors during quantization and improving the quantization accuracy of key feature channels.

[0008] Sparsification of the query-key correlation score matrix at the attention layer removes invalid weights corresponding to irrelevant or low-relevance features, resulting in a sparse distribution of attention weights and reducing unnecessary computational and storage burdens. Low-bit quantization of the target data for each feature channel is then performed based on the scaling factor and the sparsified attention weight distribution. This allows for dynamic adaptation of the quantization strategy to the importance of feature channels and attention sparsity, maximizing model size compression while ensuring high-precision representation of key feature channels. This achieves the overall technical effect of reducing storage and computational overhead, suppressing inference accuracy decay, and improving model running efficiency and stability on edge devices.

[0009] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above description and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0010] In drawings that are not necessarily drawn to scale, the same reference numerals may describe similar parts in different views. The same reference numerals with or without letter suffixes may indicate different instances of similar parts. The drawings illustrate various embodiments generally by way of example rather than limitation, and are used, together with the description and claims, to explain the disclosed embodiments. Where appropriate, the same reference numerals are used in all drawings to refer to the same or similar parts. Such embodiments are illustrative and not intended to be exhaustive or exclusive embodiments of the apparatus or method.

[0011] Figure 1 A flowchart illustrating a model processing method according to an embodiment of this application is shown.

[0012] Figure 2 A flowchart illustrating the construction of an activation distribution map according to an embodiment of this application is shown.

[0013] Figure 3 A flowchart illustrating the identification of key feature channels according to an embodiment of this application is shown.

[0014] Figure 4 A flowchart illustrating the sparsification process according to an embodiment of this application is shown.

[0015] Figure 5 A flowchart illustrating the dynamic adjustment monitoring process according to an embodiment of this application is shown.

[0016] Figure 6 A schematic diagram of the structure of a neural network model according to an embodiment of this application is shown. Detailed Implementation

[0017] To enable those skilled in the art to better understand the technical solutions of this application, the application will be described in detail below with reference to the accompanying drawings and specific embodiments. The embodiments of this application will be further described in detail below with reference to the accompanying drawings and specific examples, but these are not intended to limit the scope of this application.

[0018] The terms "first," "second," and similar words used in this application do not indicate any order, quantity, or importance, but are merely used for distinction. The terms "including" or "comprising," etc., used in this application mean that the element preceding the word encompasses the elements listed after the word, and do not exclude the possibility of encompassing other elements. In this application, the arrows shown in the figures for each step are merely examples of the execution order, not limitations. The technical solution of this application is not limited to the execution order described in the embodiments. The steps in the execution order can be combined, broken down, or rearranged, as long as the logical relationship of the executed content is not affected.

[0019] All terms used in this application (including technical or scientific terms) have the same meaning as understood by one of ordinary skill in the art to which this application pertains, unless otherwise specifically defined. It should also be understood that terms defined in general dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and not as idealized or highly formalized, unless expressly defined herein. Technologies and equipment known to one of ordinary skill in the art may not be discussed in detail, but where appropriate, such technologies and equipment should be considered part of the specification.

[0020] Figure 1 A flowchart of a model processing method according to an embodiment of this application is shown. Specifically, the model processing method executes steps S101-S105 via a processor. In this application, the arrows shown in the figure for each step are merely examples of the execution order and not limitations. The technical solution of this application is not limited to the execution order described in the embodiments. The steps in the execution order can be combined, decomposed, or rearranged, as long as the logical relationship of the executed content is not affected.

[0021] The model includes a neural network model based on the Transformer architecture, preferably a visual Transformer model. The neural network model includes, but is not limited to, deep learning models suitable for computer vision tasks such as image classification models, object detection models, and semantic segmentation models. The neural network model contains at least one of a feature extraction layer, an attention layer, a fully connected layer, and a convolutional layer, wherein the attention layer includes an attention mechanism module for calculating the relevance between the query and the key.

[0022] In step S101, the activation values ​​generated by the feature channels of each layer of the neural network model are obtained, and the activation distribution map is determined based on the activation values.

[0023] The input data is fed into the neural network model. The input data includes at least one of image data, video frame data, text data, and voice data. It is preferably real-time images or continuous video frames collected from edge devices such as smartphones, security monitoring, and vehicle sensors. It is suitable for computer vision reasoning tasks such as image classification, object detection, and semantic segmentation.

[0024] In this embodiment, each layer of the neural network model contains multiple feature channels, and the number of feature channels is determined by the structure and configuration parameters of the neural network model. For example, if a convolutional layer outputs N feature maps, then the convolutional layer corresponds to N feature channels; the output vector length of the attention layer is equal to the number of feature channels; a fully connected layer outputs a 512-dimensional vector, which means that the fully connected layer contains 512 feature channels.

[0025] This is just an example and is not enough to limit specific solutions. Different network layers can be set with the same or different number of feature channels. The feature channels can be set adaptively according to the model type, task scenario and edge device resources.

[0026] In the embodiments of this application, each feature channel corresponds to a large number of activation values. For the same input sample, after a certain layer of the neural network completes the forward computation, each feature channel will output multiple activation values ​​based on its spatial location, sequence location, or computation node. For example, in a convolutional layer, the same feature channel will correspond to multiple activation values ​​at different spatial locations on the feature map; in attention layers and fully connected layers, the same feature channel will correspond to multiple activation values ​​at different neuron nodes or different sequence locations.

[0027] In this embodiment, the activation value can be understood as the feature response value output by each feature channel after weighted calculation and activation function processing during the forward propagation of input data in each layer of the neural network. It is used to characterize the extraction strength, activation degree and response state of the corresponding feature channel to the input data, etc., and there is no limitation on it.

[0028] In some embodiments, during the forward propagation of a neural network model during training or inference, feature acquisition nodes can be inserted at the output of each layer to capture the feature response data of each feature channel after computation in real time. The output of each layer can be intercepted and saved through a forward propagation callback interface to obtain the activation values ​​of each feature channel at different spatial locations, sequence positions, or neuron nodes. By traversing each convolutional layer, attention layer, fully connected layer, and feature extraction layer of the neural network, the output tensor of each layer can be read and split according to the feature channel dimension to obtain all activation values ​​corresponding to each feature channel. Alternatively, the intermediate feature storage area during model computation can be directly accessed through memory mapping or a tensor reading interface to efficiently obtain the activation values ​​of each feature channel.

[0029] This is provided as an example only and does not constitute a limitation on any specific solution. The above implementation methods can be used individually or in combination to obtain the activation values ​​generated by all feature channels in each layer.

[0030] Specifically, as input data propagates forward through the feature extraction layer, attention layer, convolutional layer, and fully connected layer of the neural network model, each feature channel outputs an activation value at its corresponding network layer. During the neural network model training phase, by collecting and statistically analyzing the activation values ​​of each layer and each feature channel in real time, we can obtain characteristics such as the numerical distribution, fluctuation amplitude, and central tendency of the activation values ​​of each feature channel. Based on the statistical features, we calculate the activation feature parameters of each channel, and then structurally integrate and represent the activation feature parameters of each layer and each feature channel to obtain the activation distribution map.

[0031] In some embodiments of this application, a distribution feature reflecting the activation state of each feature channel is determined based on the representative value of the activation value of each feature channel; the activation distribution map is constructed based on the distribution feature of each feature channel.

[0032] Specifically, for each feature channel in each layer of the neural network model, statistical calculations can be performed on all activation values ​​generated by each feature channel under the current input data to obtain representative values ​​of each feature channel activation value.

[0033] The representative value may include at least one of the following: mean, variance, standard deviation, maximum value, minimum value, median, and mode of the activation values; or it may include the peak value, distribution center, and probability density interval obtained based on kernel density estimation, or based on K... Cluster centers, cluster dispersion, and cluster percentage obtained from clustering algorithms such as Means are statistical quantities that can characterize the overall distribution characteristics of activation values, and no specific limitations are imposed on them.

[0034] The representative value comprehensively reflects the overall level and dispersion of the activation response of each feature channel. Based on the representative value, the distribution characteristics of the activation state of the corresponding feature channel are further determined. The distribution characteristics are used to characterize the activation intensity, response stability, numerical central tendency, and fluctuation amplitude of the feature channel during the feature extraction process, and can directly reflect the accuracy sensitivity of the feature channel in the subsequent quantization process. According to the hierarchical order and channel order of the neural network, the distribution characteristics corresponding to all layers and all feature channels are organized, structured, and uniformly represented to form a complete dataset that can comprehensively reflect the activation rules of each layer of the model, i.e., the activation distribution map.

[0035] For example, the K-Means clustering algorithm can be used to construct an activation distribution map. First, for each feature channel of each layer in the neural network model, the activation values ​​corresponding to all positions under each feature channel are collected, and the one-dimensional activation value dataset is used as the clustering input. The number of clusters is set and the cluster centers are initialized. By iteratively calculating the distance between samples and each cluster center, the cluster centers are continuously updated until the algorithm converges, completing the grouping of all activation values ​​for each feature channel. The cluster centers, the number of samples within each cluster, and the numerical dispersion within each cluster are extracted as the distribution features of the feature channel, thereby characterizing the concentration, distribution ratio, and fluctuation of activation values. Following the order of network layers and feature channels, the above clustering operation is performed sequentially on all layers and all channels. The clustering results and distribution features corresponding to each feature channel are summarized, and all data are regularized, stored, and visualized to finally form a complete activation distribution map.

[0036] In a preferred embodiment, the activation score of each feature channel is obtained based on the mean and standard deviation, wherein the higher the activation score, the better the activation state of the corresponding feature channel.

[0037] Specifically, such as Figure 2 As shown, in step S201, training data is input, for example, a batch of labeled training samples (such as images with category labels) are input into the neural network model. In step S202, forward propagation is performed, and forward computation is executed on the input training samples, with each convolutional layer, attention layer, and fully connected layer sequentially outputting feature vectors. In step S203, activation values ​​are collected in real time. Specifically, during the forward propagation process, for each feature channel of each layer of the model, the activation values ​​at all positions (spatial positions or sequence positions) of each feature channel are collected in real time, forming a set of activation values ​​for each feature channel. In step S204, the mean of the activation values ​​is calculated. In step S205, the standard deviation of the activation values ​​is calculated, and an activation distribution map is constructed based on the mean and standard deviation (step S206).

[0038] For example, the activation score of each feature channel c can be calculated according to formula (1).Q c : Formula (1) In formula (1), µ c The mean value of the activation values ​​of the characteristic channels reflects the average activation intensity. (Parameter) σ c The standard deviation of the activation values ​​of the feature channels represents the degree of dispersion of the activation distribution. (Parameter) γ It is a preset regularization constant used to ensure that the denominator is non-zero and to balance the effect of variance.

[0039] Formula (1) quantifies the stability and central tendency of the activation value distribution by combining the mean and standard deviation, and the activation scores. Q c A higher activation score indicates more stable activation and higher signal strength in the corresponding feature channel, making it easier to maintain accuracy during subsequent quantization. Q c A lower activation score indicates potential abnormal fluctuations or noise in the corresponding feature channel, requiring additional attention. Thus, based on the activation score... Q c It can construct fine-grained activation distribution maps.

[0040] Returning to the embodiments of this application, in step S102, key feature channels that significantly affect the model prediction accuracy are identified based on the activation distribution map.

[0041] If the activation state of a feature channel is highly correlated with the final prediction result of the model and contributes significantly, and if the feature information of that feature channel is compressed, lost, or the quantization error is too large, it can easily lead to a significant decrease in the overall prediction accuracy of the neural network model. This indicates that the feature channel is a key feature channel and that the key feature channel has a significant impact on the prediction accuracy of the model.

[0042] Specifically, based on the activation scores output in the activation distribution map, feature channels can be sorted from high to low and selected as key feature channels according to a predetermined proportion. Alternatively, identification can be based on the standard deviation of the activation values ​​of each feature channel; for example, feature channels with a standard deviation greater than a predetermined standard deviation are considered key feature channels.

[0043] In some embodiments, a correlation score is determined based on the degree of association between the weights of each feature channel and the corresponding activation values, and the key feature channels are identified based on the correlation score.

[0044] Specifically, such as Figure 3As shown, in step S301, the activation distribution map is input, that is, the constructed activation distribution map is input to the key channel identification module. In step S302, the weight norm of the feature channel is obtained, specifically including reading the weights of each feature channel of the neural network model, and calculating the norm of its corresponding weight (such as L1 norm or L2 norm) for each feature channel to characterize the overall size and importance of the feature channel weight. In step S303, the correlation score between weight and activation is calculated, that is, the activation value of the feature channel in the activation distribution map is fused with the weight norm of the feature channel, and the degree of correlation between the two is calculated to obtain the correlation score. For example, the weight norm can be combined with activation strength and stability through weighted product, cosine similarity or linear combination. The higher the correlation score, the better the matching degree between the weight of the feature channel and the activation response, and the more critical its contribution to the model prediction.

[0045] For example, a saliency analysis algorithm can be used to identify key feature channels that significantly affect the model's prediction accuracy. These key feature channels are considered for protection during the quantization process. Specifically, the correlation score between the weight of each feature channel c and its activation value is calculated based on the correlation evaluation formula (2). C c : Formula (2) In formula (2), the parameter The L2 norm of the feature channel c represents the overall magnitude of the weights. (Parameter) Q c The activation score, as described above, represents the stability and concentration of activation values. Formula (2) quantifies the correlation between feature channels and activation values ​​by combining the weight norm and the activation value, resulting in a correlation score. C c The higher the value, the greater the contribution of the corresponding feature channel to the model inference and the more sensitive it is to quantization error. Therefore, it is necessary to prioritize its protection during the quantization process to ensure that the model performance does not decrease significantly.

[0046] In step S304, the relevance scores are sorted; in step S305, key feature channels are filtered, and a list of key feature channels is output (step S306). In other words, the relevance scores are sorted... C c You can filter by relevance score C c A certain number of the leading feature channels are designated as key feature channels.

[0047] A joint optimization method based on activation distribution mapping and saliency analysis was adopted to achieve accurate identification and protection of key feature channels. Under the constraint of limited storage resources on edge devices, the quantization process can prioritize the numerical accuracy of parameters that have the greatest impact on model accuracy. This allows the lightweight model to be deployed on memory-constrained embedded platforms, making more efficient use of the limited fixed-point representation range, significantly suppressing the accuracy decay introduced by the quantization process, and ensuring the inference accuracy of the model in practical application scenarios such as visual recognition.

[0048] In some embodiments, the gradient contribution of the weights of each feature channel to the model prediction loss is obtained; the key feature channels are identified based on the gradient contribution.

[0049] Specifically, during the backpropagation process of model training, for each feature channel in each layer of the neural network, the gradient of the corresponding weight parameter with respect to the model's prediction loss function is obtained, and the absolute value of this gradient is calculated as the gradient contribution of that feature channel. A higher gradient contribution indicates a greater impact of changes in the weight parameters of that feature channel on the model's prediction loss, and a more significant impact of the corresponding feature channel on the model's prediction accuracy. Subsequently, the gradient contributions of all feature channels are sorted from high to low, and feature channels with the highest gradient contributions are selected according to a preset ratio. These feature channels, which are most sensitive to the model's prediction loss, are marked as key feature channels.

[0050] Returning to the embodiment of this application, in step S103, the scaling factor of the key feature channel is determined by the gradient descent method, and the scaling factor is used to adjust the dynamic range of the weights of each feature channel.

[0051] The gradient descent method is an optimization algorithm that finds the minimum value of an objective function through iterative optimization. It uses the gradient information of the objective function at the current point to determine the descent direction, and gradually updates the parameters along the opposite direction of the gradient, so that the value of the objective function continuously decreases until it converges to a local or global minimum value.

[0052] Specifically, the scaling factor can be set as a trainable parameter and initialized. Then, an objective function is constructed with the core objective of minimizing the prediction accuracy loss of the quantized model, while also including a dynamic range constraint term for the weights. Then, during the backpropagation process of model training, the gradient of the objective function with respect to each scaling factor is calculated using the chain rule. The scaling factor is then iteratively updated according to the gradient descent update rule until the accuracy loss of the model on the validation set reaches a threshold or the number of iterations reaches an upper limit. This yields the optimal scaling factor that can adjust the dynamic range of the key feature channel weights and reduce quantization error while ensuring the model's prediction accuracy.

[0053] In some embodiments, a loss function is constructed based on the correlation score and weight norm of each feature channel, with the deviation between the scaled weight norm and a preset target constant as the optimization objective; the scaling factor is iteratively updated along the gradient descent direction until the loss function converges, and the scaling factor is determined.

[0054] In other words, the loss function uses the deviation between the scaled weight norm and the preset target constant as the optimization objective. The deviation between the scaled weight norm and the preset target constant refers to the difference between the original weight norm of the feature channel after adjustment by the scaling factor and the preset target constant. The magnitude of the deviation represents the degree to which the current weight amplitude deviates from the ideal value range. A larger deviation indicates a more significant deviation of the overall weight amplitude from the standard, and the rounding error generated during quantization will also increase accordingly. During the iteration process, the scaling factor is continuously adjusted so that the overall amplitude of the scaled weights of each feature channel continuously approaches the preset target constant, causing the weight range of the key feature channels to converge to the dynamic range suitable for model quantization, thereby reducing the rounding error in the quantization process.

[0055] The preset target constant can be a value pre-set according to the model quantization accuracy, weight value range and hardware storage specifications. It can be used as a benchmark value of the weight norm to uniformly constrain the overall numerical range of the weights after scaling of all feature channels.

[0056] Iterative optimization is carried out using the gradient descent algorithm based on the loss function. The gradient of the loss function with respect to the scaling factor is calculated, and the value of the scaling factor is updated step by step in the opposite direction of the gradient. The iterative process of forward calculation of loss, backward calculation of gradient, and parameter update is repeated until the value of loss function tends to stabilize and convergence condition is met. At this time, the optimal scaling factor corresponding to each feature channel can be determined. This scaling factor can accurately adjust the dynamic range of weights and effectively reduce rounding errors in the quantization process.

[0057] For example, an automatic optimization search algorithm can be used to dynamically search for and determine the optimal scaling factor. This scaling factor is used to adjust the dynamic range of the weight values, effectively reducing the rounding error generated in the quantization process, so that the weights of the key feature channels fall precisely within the fitting range after quantization, minimizing the accuracy loss caused by rounding.

[0058] For example, the loss function can be derived from formula (3). L s To determine: Formula (3) Among them, parameters C c The relevance score for feature channel c reflects the importance of key feature channels; The L2 norm represents the weights of the key feature channel c, characterizing the magnitude and strength of the weights.S c It is the scaling factor, parameter β It is a preset target constant used to control the matching target of the dynamic range of weights.

[0059] Based on the loss function, the scaling factor is determined using the gradient descent method. S c This loss function uses relevance scores. C c Weighting is applied to make the scaling factor... S c with weight norm The weights of each feature channel are inversely proportional, guiding them towards a preset target constant after scaling. β Convergence is achieved, thus maintaining the relative stability of the weights during the quantization process. During optimization, the scaling factor is adjusted more precisely for key feature channels with high correlation scores. S c This significantly reduces rounding errors caused by quantization; while for non-critical feature channels, relatively large quantization errors are allowed, thereby maximizing overall compression efficiency while ensuring model accuracy.

[0060] After determining the scaling factor, it is applied to all weight parameters of the key feature channels and the attention layer to complete the weight value range calibration. Then, the features output by the calibrated network are sent as input features to the attention layer.

[0061] Returning to the embodiment of this application, in step S104, the correlation score matrix between the query and the key is sparsified in the attention layer to obtain the attention weight distribution.

[0062] In other words, the sparsification process of the attention layer selectively retains and filters the relevance score matrix of the query (Q) and key (K) based on the weights after scaling factor calibration, thereby obtaining a lightweight attention weight distribution.

[0063] Specifically, the network output features, optimized by weight scaling factors and calibrated for key feature channels, are used as the input features for the attention layer. In the attention layer, these input features are linearly transformed to generate query vectors and key vectors. A matrix dot product is then performed on the two sets of vectors, and the matrix is ​​scaled according to the feature dimension to obtain a correlation score matrix representing the strength of association at each feature location. This score matrix is ​​then sparsified by setting elements with low correlation and small contributions to zero, retaining only the valid correlation scores. The sparsified matrix is ​​then normalized to convert the values ​​into weight coefficients conforming to a probability distribution, resulting in a sparse attention weight distribution. This effectively reduces the computational and storage overhead of the attention layer.

[0064] In some embodiments of this application, an attention score between a query and a key is determined based on the correlation between the query and the key and a saliency bias term; an attention threshold for the query is determined based on the attention scores of each key corresponding to the same query; the difference between the two when the attention score is greater than the attention threshold is used as the weight of the corresponding key, and the weight of the corresponding key when the attention score is less than or equal to the attention threshold is reset to zero, thus obtaining a sparse attention weight distribution.

[0065] Specifically, such as Figure 4 In step S401, the query vector Q and key vector K are input. A matrix dot product operation is performed on the query vector Q and key vector K, followed by numerical scaling by dividing by the square root of the feature dimension. Then, a saliency bias term is introduced to correct the result, resulting in a relevance score matrix. Each row of this relevance score matrix corresponds to a query vector Q, and each column corresponds to a key vector K. Each element in the relevance score matrix represents the strength of the association between the corresponding query and key. In step S402, based on the relevance between the query and key, and with the introduction of a saliency bias term, the original attention score is calculated.

[0066] For example, in the attention layer, the Sparsemax function can be used to completely replace the traditional Softmax function to calculate the attention weight distribution between the query and the key. Specifically, the query is calculated according to formula (4). i s and keys j Attention scores between Z ij : Formula (4) Among them, parameters q i Indicates the first i A query vector, parameters k j Indicates the first j A key vector, parameters d The parameters represent the feature dimensions of the query and the key. θ For trainable scaling factors, parameters S i This represents the significance score of the i-th query, which is derived from the relevance score of the feature channels. C c The mapping yields the result that reflects the first i The importance of each query. For example, the relevance score of a feature channel can be directly used. C c Assign the value to the significance score of the corresponding query.

[0067] Thus, by adding a saliency bias term based on query saliency θS iIt can enhance the attention score of important queries, improve the overall attention score of important queries, make it easier to retain attention to key information during sparsification, and reduce the interference of irrelevant information.

[0068] The attention weights were precisely sparsely distributed by employing the Sparsemax function. On edge devices, a large number of zero weights means that corresponding memory read operations and multiplication operations can be skipped or optimized, which significantly reduces the power consumption of dynamic random access memory (DRAM) and the computational load of the arithmetic logic unit (ALU). During inference on edge devices, the sparse attention weight matrix significantly reduces the number of elements that need to be stored and participate in multiply-accumulate operations, directly reducing the device's memory bandwidth usage and the number of activations of computing units (such as the MAC unit in the NPU), thereby reducing the energy consumption and computational latency of the inference process and creating better hardware execution conditions for subsequent quantization operations.

[0069] In other embodiments, a Hardmax function or gated attention mechanism can be used to directly mask low-relevance inputs using a trainable binary mask, or a local attention mechanism can be combined to limit the computational range of queries and keys, achieving weight sparsity in a simpler way. For example, a gated attention mechanism can be used to generate a binary mask that can directly control the data path switching of the edge device's computing unit, or serve as a conditional instruction to skip the computation of zero-weight regions, thereby saving energy at the hardware level.

[0070] In step S403, the attention threshold for the query is calculated and determined based on the attention scores of each key corresponding to the same query.

[0071] Specifically, each query can be calculated according to formula (5). i attention threshold τ i : Formula (5) In formula (5), parameter N represents the number of key vectors. Based on formula (5), the attention threshold can be determined by mean adjustment, thereby performing preliminary pruning of the attention score and providing a basis for subsequent sparsity processing.

[0072] In step S404, a sparsified output function is applied, in step S405 a sparsified attention weight distribution is generated, and in step S406 it is passed to the subsequent quantization step.

[0073] Specifically, sparsification output formulas can be used. To calculate attention weights p ijWhen the attention score corresponding to a key is greater than the attention threshold, the difference between the two is used as the weight of the key; when the attention score is less than or equal to the attention threshold, the weight of the corresponding key is reset to zero.

[0074] Attention scores are sparsified by using attention threshold filtering, which removes invalid weights with low relevance and retains only valid weights with high contribution. Inputs with low relevance or irrelevant information are assigned zero weights, thereby avoiding abnormal activation values ​​caused by the model forcibly assigning weights when processing irrelevant information. This reduces computational and storage overhead while ensuring the accuracy of key information.

[0075] In this embodiment of the application, in step S105, based on the scaling factor and attention weight distribution, the target data of each feature channel is quantized with low bit depth to obtain a lightweight model.

[0076] Low-bit quantization is used to convert high-precision floating-point parameters in the model into low-bit-width integers, thereby reducing model storage and computational overhead by decreasing numerical precision. The target data includes weights and / or activation values, preferably the weights of feature channels calibrated by a scaling factor.

[0077] Specifically, the feature channel weights calibrated by the scaling factor and the sparsified attention weight distribution are used together as quantization objects. The scaling factor is used to calibrate the numerical distribution of each feature channel weight to the range that is suitable for low-bit representation. Combined with the differences in the importance of key channels in the attention weight distribution, a differentiated quantization strategy is implemented for different feature channels.

[0078] For example, higher bit precision is used for critical feature channels to control rounding errors, while lower bit precision is used for less important feature channels to improve compression ratio. In this way, low-bit quantization of the entire model is completed, resulting in a lightweight model.

[0079] The scaling factor pre-corrects the numerical distribution of the weights, reducing truncation errors and accuracy loss during quantization. The differentiated quantization strategy guided by the attention weight distribution achieves a balance between accuracy and compression rate, ensuring the model's ability to process key information while significantly reducing the model's storage footprint and inference latency, providing reliable support for the efficient deployment of edge devices.

[0080] In some embodiments of this application, the floating-point values ​​of the target data of each feature channel are mapped to integers with a preset bit width; the quantization range of each feature channel is dynamically calibrated based on the scaling factor and attention weight distribution; and the mapped integer values ​​are adjusted according to the calibrated quantization range to obtain the target data after low-bit quantization.

[0081] Specifically, the original weights of the model are mostly high-precision floating-point numbers (such as 32-bit), while low-bit quantization is to convert these continuous floating-point numbers into discrete integers through linear or non-linear mapping. For example, suppose there is a set of floating-point numbers [1.2, 3.5, -2.1, 0.8], which can be mapped to [0, 3, 1, 2].

[0082] The preset bit width refers to the number of binary bits in the quantized integer, such as an 8-bit integer or a 4-bit integer. The fewer bits, the fewer discrete values ​​are represented. The smaller the bit width, the less storage space each value occupies, and the overall size of the model will be significantly reduced. For example, replacing a 32-bit floating-point number with a 4-bit integer can theoretically achieve a compression ratio of 8 times.

[0083] Specifically, the scaling factor is used to initially correct the numerical distribution of the weights of each feature channel, constraining the overall data range to a quantization interval more suitable for low-bit quantization. Then, combined with the channel saliency information reflected in the attention weight distribution, key and non-key feature channels are identified, and differentiated quantization range settings are applied to different feature channels. Specifically, for key feature channels, the quantization range is appropriately narrowed and the precision of the quantization step is increased to reduce the impact of rounding errors on key information; for non-key feature channels, the quantization range is widened and a coarser quantization step is used to improve overall compression efficiency.

[0084] For example, floating-point values ​​are mapped to hardware-friendly 8-bit integer representations, and the 8-bit integer weights are quantized according to the adaptive quantization formula of formula (6). W q : Formula (6) in, W f Represents the original floating-point weights. S c This is the scaling factor. p ij α represents the sparsification attention weight, and α represents a small constant to prevent division by zero.

[0085] By combining scaling factor adjustment and attention sparsity information, the quantization range of each feature channel is dynamically calibrated. A finer quantization interval is used for the effective weighted regions after sparsification to preserve the accuracy of key information; while the quantization accuracy requirements are relaxed for zero-weighted regions. This maximizes the overall compression efficiency while ensuring that the loss of key information is controllable.

[0086] This quantization process integrates both symmetric and asymmetric quantization strategies. When the channel weight distribution is symmetrical, a zero-centered symmetric quantization method is used; when the weight distribution is offset, an offset-based asymmetric quantization method is used, thereby effectively reducing model storage requirements and computational complexity. The 8-bit integer data generated after quantization will be directly used for edge device inference in subsequent steps, and will also provide a compressed data foundation for real-time dequantization operations in model deployment.

[0087] In some embodiments of this application, when the lightweight model is deployed to an edge device: the low-bit quantized target data corresponding to the key feature channel is stored in the cache of the edge device, while the low-bit quantized target data corresponding to the other feature channels is stored in main memory.

[0088] Specifically, during the edge device deployment phase, the quantized model can be scored according to its relevance. C c Tiered storage loads the 8-bit integer quantization weights corresponding to key feature channels that exceed the threshold into the high-speed cache of the edge device, while the quantization weights corresponding to the remaining feature channels are loaded into main memory. The low latency of the high-speed cache is used to improve the access efficiency of key weights.

[0089] In some embodiments of this application, dequantization is triggered on demand during inference, and dequantization is initiated when a critical operation is triggered; dynamic precision compensation is performed on the target data after low-bit quantization to restore it to a floating-point value.

[0090] In this inference process, the on-demand dequantization operation does not perform a full dequantization of all low-bit data at once. Instead, dequantization is dynamically initiated only when high-precision floating-point data is required in critical computational steps. This restores the low-bit quantized target data to its original floating-point value, achieving dynamic precision compensation. For example, when the model enters a critical operation requiring high-precision computation (such as attention score calculation, gradient correlation, or nonlinear activation processing), the system invokes the dequantization process according to preset trigger conditions. Using the scaling factor and offset recorded during the quantization phase, it maps discrete low-bit integers back to continuous floating-point values. Other non-critical steps directly use low-bit data for computation. This approach avoids the memory consumption and computational overhead of maintaining high-precision floating-point data throughout the entire process, while dynamically restoring data precision at critical nodes to compensate for quantization errors, thus balancing the efficiency and accuracy of model inference.

[0091] Specifically, the original floating-point weights can be achieved through the inverse quantization calibration formula (7). W f Accurate restoration: Formula (7) In formula (7), W 4 represents the 8-bit integer quantization weight based on the output of formula (6), min( W f ) and max( W f ) represents the minimum and maximum values ​​of the dynamic range of the original floating-point weights. S c It is a scaling factor. Represents attention weights p ij The summation value, where α represents a small constant to prevent division by zero.

[0092] Thus, through correlation scores C c and The synergistic effect of these factors leads to high correlation scores. C c And high The weight regions corresponding to the key feature channels receive stronger accuracy compensation, resulting in lower correlation scores. C c Maintaining high computational efficiency: After calibration, significant weights stored in the cache can reduce cumulative error by more than 30%; non-significant weights stored in main memory ensure computational efficiency while meeting model accuracy requirements. Dequantization operations are performed on demand; frequently used dequantization results in the cache are permanently stored using a cache locking mechanism, while dequantization data in main memory is pre-loaded using a prefetch algorithm.

[0093] The multi-core processors of the edge device complete the numerical conversion work in parallel. The system synchronously collects performance data such as access latency and inverse quantization time of each storage area. This data, together with the model output results, forms a complete set of real-time monitoring indicators.

[0094] Based on the embodiments of this application, adaptive runtime management on edge devices can be realized. Continuous monitoring can dynamically adjust model operating parameters according to the physical state of the edge device, such as real-time remaining power, chip temperature, and current computing load. For example, when the device overheats, the computational load can be reduced by increasing the sparsity threshold, preventing the chip from throttling due to overheating and causing performance instability, thereby ensuring the long-term reliable operation of the system in complex physical environments.

[0095] Figure 5 The diagram illustrates an adaptive optimization process during model inference, which dynamically adjusts quantization parameters and sparsity thresholds based on real-time feedback from performance and output data.

[0096] In step S501, after inference monitoring begins, steps S502 (collecting resource usage data) and S507 (collecting model output data) are executed simultaneously. Both types of data are aggregated in step S503 to analyze performance metrics and generate two optimization instructions. One instruction proceeds to step S504 to adjust quantization parameters, while the other proceeds to step S508 to adjust the sparsity threshold. The two sets of new parameters are integrated and applied to the model in step S505. In step S506, the balance between accuracy and efficiency after adjustment is evaluated, and the evaluation result is fed back to step S502, forming a continuous iterative closed-loop optimization.

[0097] Throughout the inference process, model output and resource usage are continuously monitored, and the quantized weight parameters and attention threshold are dynamically adjusted based on real-time performance feedback. This involves collecting data on computational latency, memory usage, and output confidence variance, and using this data to construct a dynamic balancing controller. This controller employs a bidirectional adjustment strategy: when resource usage exceeds a preset limit, it automatically reduces the number of quantized bits and increases the attention threshold to prioritize operational efficiency; conversely, when output quality deteriorates, it increases quantization accuracy and decreases the attention threshold to restore the model's expressive power. The adjustment process is gradual to avoid sudden parameter changes impacting the system. Performance metrics are re-evaluated after each adjustment until an optimal balance between compression ratio and model accuracy is achieved. This ensures the long-term stability and adaptability of the lightweight model in the complex operating environment of edge devices, providing performance assurance for the entire attention-based edge device AI model lightweight compression and inference method.

[0098] The model processing methods provided in the various embodiments of this application systematically solve the technical connection problem from server training environment to resource-constrained edge device deployment environment, enabling the model to adapt to the acceleration of common integer instruction sets of edge processors. At the same time, by dynamically balancing the computational load and model accuracy through runtime monitoring, it ultimately achieves low-latency, high-precision, and low-memory-consumption AI inference capabilities that meet the requirements of practical applications on specific edge hardware such as ARM Cortex-A series CPUs or edge AI acceleration chips, significantly improving the practical value of industry.

[0099] In some embodiments of this application, a neural network model is provided, such as Figure 6The neural network model includes an input data layer, a neural network feature extraction layer, an activation value collection and analysis module, a weight channel saliency analysis module, a scaling factor optimization module, an attention mechanism optimization layer, an adaptive quantization module, an edge device storage layering module, an inference execution module, a performance monitoring and feedback module, and a dynamic balance controller. The input data layer is configured to receive text or image data. After the neural network feature extraction layer extracts features, the activation value collection and analysis module calculates the mean and variance of the activation values ​​and generates activation scores (which are also quantization adaptation scores), constructing an activation distribution map. The weight channel saliency analysis module calculates the correlation score of each feature channel based on the activation distribution map, generating a list of key feature channels (i.e., a list of saliency weight channels). The scaling factor optimization module solves for the optimal scaling factor for each feature channel using the gradient descent algorithm. The attention mechanism optimization layer combines the sparsemax function and the sparsity formula to generate sparse attention weights. The adaptive quantization module combines the optimal scaling factor and the sparse attention weights, employing symmetric or asymmetric quantization strategies to convert the model into an 8-bit integer quantization model. The edge device storage tiering module stores significant weights in a high-speed cache and insignificant weights in main memory, forming a hierarchical storage model. The inference execution module uses an instantaneous dequantization calibration formula during critical operations to restore low-bit data to floating-point values ​​for high-precision computation. The performance monitoring and feedback module collects latency, memory usage, and output confidence metrics from the inference process. The dynamic balancing controller dynamically adjusts quantization parameters and sparsity thresholds based on feedback, achieving closed-loop optimization of compression ratio and accuracy.

[0100] This neural network model is suitable for edge intelligence scenarios that process visual data (images, video frames, and real-time camera feeds), meaning it performs AI inference directly on the terminal device without relying on cloud computing. The model's input consists of images or video frames, such as real-time feeds from smartphone cameras, security cameras, or vehicle vision sensors, and it needs to perform tasks such as object detection, image classification, or semantic segmentation. These edge devices face strict physical limitations in terms of memory capacity, computing power (CPU / GPU / NPU), battery consumption, and real-time response latency. Therefore, lightweight compression of large models with image input and efficient inference are urgent technical challenges that need to be addressed.

[0101] In some embodiments of this application, a model processing apparatus is provided, the model processing apparatus including a processor configured to execute the steps of the model processing methods described in various embodiments of this application.

[0102] The model processing methods of various embodiments of this application can all be incorporated herein, and will not be described in detail here.

[0103] The processor can be a processing device that includes one or more general-purpose processing devices, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that runs other instruction sets, or a processor that runs a combination of instruction sets. The processor can also be one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), system-on-a-chip (SoCs), etc.

[0104] In some embodiments of this application, a computer program product is provided, the computer program product comprising computer-executable instructions, which, when executed by a processor, implement the steps of the model processing method according to various embodiments of this application.

[0105] This application describes various operations or functions that can be implemented as software code or instructions, or defined as software code or instructions. Such content can be directly executable source code or differential code (“incremental” or “patch” code) (“object” or “executable” form). The software code or instructions can be stored in a computer-readable storage medium and, when executed, can cause a machine to perform the described functions or operations, and include any mechanism for storing information in a machine-accessible form, such as recordable or non-recordable media (e.g., read-only memory (ROM), random access memory (RAM), disk storage media, optical storage media, flash memory devices, etc.).

[0106] Implementations of such methods may include software code, such as microcode, assembly language code, high-level language code, etc. Various software programming techniques can be used to create various programs or program modules. For example, program parts or program modules can be designed using or with the aid of Java, Python, C, C++, assembly language, or any known programming language. One or more of such software parts or modules can be integrated into a computer system and / or a computer-readable medium. Such software code may include computer-readable instructions for performing various methods. This software code can form part of a computer program product or a computer program module. Furthermore, in the example, the software code may be tangibly stored on one or more volatile, non-transitory, or non-volatile tangible computer-readable media, for example, during execution or at other times. Examples of such tangible computer-readable media may include, but are not limited to, hard disks, removable disks, removable optical discs (e.g., optical discs and digital video discs), magnetic tape cassettes, memory cards or memory sticks, random access memory (RAM), read-only memory (ROM), etc.

[0107] Furthermore, although exemplary embodiments have been described herein, their scope includes any and all embodiments based on this application that have equivalent elements, modifications, omissions, combinations (e.g., schemes involving intersections of various embodiments), adaptations, or alterations. Elements in the claims will be interpreted broadly based on the language used in the claims and are not limited to the examples described in this specification or during the implementation of this application, which will be interpreted as non-exclusive. Therefore, this specification and examples are intended to be considered illustrative only, and the true scope and spirit are indicated by the following claims and the full scope of their equivalents.

[0108] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. Other embodiments can be used by those skilled in the art when reading the above description. Furthermore, in the above detailed description, various features may be grouped together to simplify the application. This should not be construed as an intention that a disclosed feature not claimed is necessary for any claim. Rather, the subject matter of the application may be less than all the features of a particular disclosed embodiment. Thus, the claims are incorporated herein by reference as examples or embodiments, wherein each claim is an independent, separate embodiment, and these embodiments are contemplated as being able to be combined with each other in various combinations or arrangements. The scope of this application should be determined by reference to the appended claims and the full scope of their equivalents.

[0109] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.

Claims

1. A model processing method, characterized in that, The model processing method includes: The activation values ​​generated by the feature channels of each layer of the neural network model are obtained, and the activation distribution map is determined based on the activation values. Based on the activation distribution map, key feature channels that significantly affect the model's prediction accuracy were identified. The scaling factor of the key feature channel is determined by the gradient descent method, and the scaling factor is used to adjust the dynamic range of the weights of each feature channel. The attention layer performs sparsification on the correlation score matrix between the query and the key to obtain the attention weight distribution; Based on the scaling factor and attention weight distribution, the target data of each feature channel is quantized with low bit depth to obtain a lightweight model.

2. The model processing method according to claim 1, characterized in that, The activation distribution map is determined based on the activation values, including: Based on the representative values ​​of the activation values ​​of each feature channel, the distribution characteristics that reflect the activation state of each feature channel are determined. The activation distribution map is constructed based on the distribution characteristics of each feature channel.

3. The model processing method according to claim 2, characterized in that, The representative values ​​include the mean and standard deviation of the activation values. The activation score of each feature channel is obtained based on the mean and standard deviation. The higher the activation score, the better the activation state of the corresponding feature channel.

4. The model processing method according to claim 1, characterized in that, Based on the activation distribution map, key feature channels that significantly affect the model's prediction accuracy were identified, including: The correlation score is determined based on the degree of correlation between the weights of each feature channel and the corresponding activation values, or the gradient contribution of the weights of each feature channel to the model's prediction loss is obtained. The key feature channels are identified based on the correlation score or gradient contribution.

5. The model processing method according to claim 1, characterized in that, The scaling factor for the key feature channels is determined using a gradient descent method, including: Based on the correlation scores and weight norms of each feature channel, a loss function is constructed with the deviation between the scaled weight norm and the preset target constant as the optimization objective. The scaling factor is iteratively updated along the gradient descent direction until the loss function converges, and the scaling factor is then determined.

6. The model processing method according to claim 1, characterized in that, The attention layer performs sparsification on the query-key correlation score matrix to obtain the attention weight distribution, including: Based on the relevance between the query and the key, as well as the salience bias term, determine the attention score between the query and the key; The attention threshold of the query is determined based on the attention scores of each key corresponding to the same query. When the attention score is greater than the attention threshold, the difference between the two is used as the weight of the corresponding key. When the attention score is less than or equal to the attention threshold, the weight of the corresponding key is reset to zero, resulting in a sparse attention weight distribution.

7. The model processing method according to claim 1, characterized in that, The target data includes weights and / or activation values; Based on the scaling factor and attention weight distribution, low-bit quantization is performed on the target data of each feature channel, including: Map the floating-point values ​​of the target data in each feature channel to integers with a preset bit width; The quantization range of each feature channel is dynamically calibrated based on the scaling factor and attention weight distribution. Based on the calibrated quantization range, the mapped integer values ​​are adjusted to obtain the target data after low-bit quantization.

8. The model processing method according to claim 1, characterized in that, The model processing method further includes: in the case of deploying the lightweight model to an edge device: The target data corresponding to the key feature channels, after low-bit quantization, is stored in the cache of the edge device, while the target data corresponding to the other feature channels, after low-bit quantization, is stored in the main memory.

9. The model processing method according to claim 1, characterized in that, The model processing method further includes: Dequantization is triggered as needed during the inference process, and dequantization processing is initiated when a critical operation is triggered. Dynamic precision compensation is performed on the target data after low-bit quantization to restore it to a floating-point value.

10. A model processing device, characterized in that, The model processing apparatus includes a processor configured to perform the steps of the model processing method according to any one of claims 1-9.