Pruning method of neural network model, data processing method and related equipment

By acquiring calibration data to calculate weight importance scores, utilizing noise perturbation sensitivity analysis and change-aware assessment, and combining unstructured pruning strategies, sparsity rates are adaptively allocated to generate sparse neural network models. This solves the problems of insufficient weight evaluation and model stability in existing pruning methods, and achieves efficient and stable model compression and deployment.

CN122047352APending Publication Date: 2026-05-15BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2025-12-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing pruning methods rely excessively on input activation or weight magnitude to assess the importance of weights, resulting in insufficient generalization ability of pruning decisions. Furthermore, the lack of systematic quantification of the impact of weights on the stability of model output during the pruning process can easily lead to performance degradation or unstable accuracy.

Method used

By acquiring calibration data, calculating the importance score of the weights, utilizing noise perturbation sensitivity analysis and change-aware importance assessment, and combining unstructured pruning strategies, an adaptive sparsity rate is allocated to generate a sparse neural network model.

Benefits of technology

It achieves high-precision model compression and deployment without additional fine-tuning, significantly reducing parameter size and computational overhead, and improving the stability and deployment adaptability of the pruning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122047352A_ABST
    Figure CN122047352A_ABST
Patent Text Reader

Abstract

The invention provides a pruning method of a neural network model, a data processing method and related equipment. The method comprises the following steps: acquiring a neural network model to be pruned and calibration data; calculating importance scores of weight parameters in the neural network model to be pruned based on the calibration data; and pruning the weight parameters according to the importance score to generate a sparse neural network model. According to the method, the precision stability of the neural network model is ensured, the pruning efficiency and the actual deployment adaptability are improved, and the method is suitable for efficient compression, low-power-consumption deployment and model reasoning optimization of a large-scale deep neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of deep learning technology, and in particular to a pruning method, data processing method and related equipment for a neural network model. Background Technology

[0002] With the widespread application of deep learning in fields such as computer vision, natural language processing, and intelligent control, the scale and number of parameters of deep neural networks have exploded. Although large-scale neural networks have achieved excellent performance in various tasks, their inference process is accompanied by high computational overhead, storage pressure, and energy consumption, which seriously restricts their deployment capabilities in resource-constrained devices or real-time application scenarios.

[0003] In recent years, neural network pruning has gradually become a research hotspot as an effective model compression method, accelerating inference by reducing redundant parameters. However, existing pruning methods still face several technical challenges in practical applications: on the one hand, weight importance evaluation often relies excessively on input activation or weight magnitude, resulting in insufficient generalization ability of pruning decisions; on the other hand, during the pruning process, the inability to systematically quantify the impact of weights on model output stability can easily lead to performance degradation or unstable accuracy after pruning. Furthermore, some methods require complex fine-tuning to restore accuracy, increasing the cost and complexity of the pruning process.

[0004] Therefore, there is an urgent need for a method that can achieve high-precision, fast, and reliable model compression and deployment without additional fine-tuning.

[0005] It should be noted that the above description of the technical background is only for the purpose of providing a clear and complete explanation of the technical solutions of the present invention and facilitating understanding by those skilled in the art. It should not be assumed that the above technical solutions are known to those skilled in the art simply because they have been described in the background section of this invention. Summary of the Invention

[0006] In view of the above, the purpose of one or more embodiments of this disclosure is to provide a pruning method, data processing method and related device for neural network models, so as to solve or partially solve the problems raised in the background art.

[0007] In a first aspect, this disclosure provides a pruning method for a neural network model, the method comprising: Obtain the neural network model to be pruned and calibration data; wherein, the calibration data includes text sequences randomly extracted from an unsupervised corpus; Based on the calibration data, the importance score of the weights in the neural network model to be pruned is calculated; The weights are pruned based on the importance score to generate a sparse neural network model.

[0008] A second aspect of this disclosure provides a data processing method, the method comprising: Acquire the target data and the sparse neural network model trained according to the method described in the first aspect; The target data is processed based on the sparse neural network model.

[0009] A third aspect of this disclosure provides a pruning device for a neural network model, the device comprising: The acquisition module is configured to acquire the neural network model to be pruned and calibration data; wherein the calibration data includes text sequences randomly extracted from an unsupervised corpus. The calculation module is configured to: calculate the importance score of the weights in the neural network model to be pruned based on the calibration data; The pruning module is configured to prune the weights based on the importance score to generate a sparse neural network model.

[0010] A fourth aspect of this disclosure provides a data processing apparatus, the apparatus comprising: The acquisition module is configured to: acquire target data and a sparse neural network model trained according to the method described in the first aspect; The processing module is configured to process the target data based on the sparse neural network model.

[0011] A fifth aspect of this disclosure provides a computer device including one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, the programs including instructions for performing the method according to the first or second aspect.

[0012] A sixth aspect of this disclosure provides a non-volatile computer-readable storage medium comprising a computer program that, when executed by one or more processors, causes the processors to perform the method described in the first or second aspect.

[0013] A seventh aspect of this disclosure provides a computer program product including computer program instructions that, when executed on a computer, cause the computer to perform the method described in the first or second aspect.

[0014] The pruning method, data processing method, and related equipment for neural network models disclosed herein determine the importance scores of weights in the neural network model using calibration data, and then prune the weights according to the importance scores to achieve sparsity of the neural network model. This can significantly reduce the parameter scale and computational overhead while ensuring model accuracy, and has good engineering deployment value and broad application prospects. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in one or more embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only one or more embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart illustrating an exemplary method provided in an embodiment of this disclosure; Figure 2 A schematic diagram of an exemplary device provided in accordance with the embodiments of this disclosure; Figure 3 A flowchart illustrating another exemplary method provided in an embodiment of this disclosure; Figure 4 A schematic diagram of another exemplary device provided in the embodiments of this disclosure; Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0018] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar words used in one or more embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0019] As described in the background section, the relevant technologies have the following problems: First, existing pruning methods rely excessively on single indicators such as input activation or weight magnitude to evaluate the importance of weights, resulting in insufficient generalization ability of pruning decisions; Second, the pruning process lacks a systematic quantitative mechanism for assessing the impact of weights on the stability of model output, which can easily lead to model performance degradation or accuracy fluctuations.

[0020] Therefore, there is an urgent need to develop a technical solution for pruning neural network models to meet the engineering requirements of rapidly and robustly evaluating the importance of weights under multiple calibration data, achieving adaptive sparse allocation, and achieving high-precision model compression and reliable deployment without additional fine-tuning.

[0021] Based on some implementations of this disclosure, a pruning method for neural network models is provided. In this method, efficient sparsification of the neural network is achieved by adaptively allocating weight perturbation sensitivity analysis, change-aware importance assessment, and layer sparsity rate, combined with an unstructured pruning strategy. The entire pruning process requires no adjustment to the model structure or additional fine-tuning, ensuring both model accuracy and stability while improving pruning efficiency and practical deployment adaptability. It is suitable for efficient compression, low-power deployment, and model inference optimization of large-scale deep neural networks.

[0022] refer to Figure 1 The present invention discloses a pruning method for a neural network model, comprising the following steps: S101. Obtain the neural network model to be pruned and calibration data; wherein, the calibration data includes text sequences randomly extracted from an unsupervised corpus.

[0023] In this embodiment, the present disclosure obtains the neural network model to be pruned and calibration data, providing a stable and reliable model foundation and input context for subsequent noise perturbation injection and weight importance assessment, which is a prerequisite for achieving fine-tuning-free pruning.

[0024] In this embodiment, obtaining the neural network model to be pruned and the calibration data specifically includes: S1011. Obtain the neural network model to be pruned and the calibration data.

[0025] Obtain the neural network model to be pruned; wherein, the neural network model to be pruned is a pre-trained and stable neural network model, which includes... The nth linear layer, the th The weight matrix of the layer is denoted as Its parameters have been pre-trained on a large-scale corpus, and the neural network model has good convergence and generalization ability on the target task, with overall performance in a stable state.

[0026] In this embodiment, the linear layer can refer to a network layer that receives input and applies a linear transformation. It multiplies each input by its corresponding weight, sums the results, and then adds a bias to obtain the output. Linear layers are commonly used to map input data to the feature space of the next layer. In this embodiment, the linear layer can refer to a linear layer in an attention-based neural network model, or a linear layer in a feedforward neural network (FFN).

[0027] S1012. Randomly select from general unsupervised corpora. A representative text sequence was used as calibration data; wherein, the general unsupervised corpus refers to The data in the dataset; the representative text sequence refers to The natural language text for each data point in the dataset; typical values ​​can be... .

[0028] In this embodiment, the The dataset is a large-scale, high-quality English text dataset, officially named Colossal Clean Crawled Corpus.

[0029] S102. Based on the calibration data, calculate the importance score of the weights in the neural network model to be pruned.

[0030] In this embodiment, controlled noise is injected to quantize the sensitivity of the weights to the model output, providing a dynamic response basis for subsequent importance assessment.

[0031] In this embodiment, calculating the importance score of the weights in the neural network model to be pruned based on the calibration data specifically includes: S1021. Input the calibration data into the neural network model to be pruned to obtain the input activation matrix corresponding to each linear layer in the neural network model to be pruned. The neural network model to be pruned contains at least one linear layer.

[0032] In this embodiment, the input activation matrix reflects the internal representation of the input data by the neural network model in a real inference scenario, and is used to quantize the output changes caused by subsequent perturbations.

[0033] In this embodiment, the present disclosure introduces a small amount of unlabeled calibration data to provide a data foundation for evaluating the impact of weight perturbations on the output of the neural network model. The calibration data is used only for importance score calculation during the pruning process and does not participate in any model parameter updates or fine-tuning operations. This ensures that the entire pruning process is a post-training pruning method, significantly reducing computational overhead and deployment complexity, while guaranteeing the universality and reproducibility of the pruning strategy.

[0034] S1022. Apply noise perturbation to the original weight matrix of each linear layer to generate the perturbed weight matrix.

[0035] For the neural network model to be pruned, the first... Layer weight matrix Apply structured Gaussian noise perturbation to generate the perturbed weight matrix. :

[0036] in, For the neural network model to be pruned, the first... Layer weight matrix; To and A standard normal random matrix of the same dimension, i.e. ; This is a global scaling factor used to precisely control the overall disturbance intensity; For element-wise multiplication; It is an absolute value.

[0037] In this embodiment, the perturbation design of this disclosure has two key characteristics: First, since the noise is introduced in absolute value form, the perturbed weights maintain the same sign as the original weights, effectively avoiding the cancellation effect caused by weight sign reversal in subsequent calculations and ensuring the accuracy of sensitivity assessment; second, the perturbation amplitude is proportional to the original weight size, so that weights with larger values ​​will be subject to greater perturbation, thereby more realistically reflecting the functional influence of the weight in the model and embodying the scale-aware sensitivity assessment principle. Therefore, it ensures that the weight sign remains unchanged after perturbation, and that the perturbation amplitude is proportional to the original weight size, reflecting scale-aware characteristics.

[0038] S1023. Based on the input activation matrix, the original weight matrix, and the perturbed weight matrix, the output deviation is calculated, specifically including: The original output is calculated using the original weight matrix:

[0039] The perturbation output is calculated using the perturbated weights:

[0040] Based on the original output and the perturbed output, the output deviation is calculated as follows:

[0041] in, For the neural network model to be pruned, the first... The original output matrix of the linear layer, For the neural network model to be pruned, the first... The perturbation output matrix of the linear layer. For the neural network model to be pruned, the first... The input activation matrix of a linear layer, For the neural network model to be pruned, the first... The transpose of the original weight matrix of the linear layer. For the neural network model to be pruned, the first... The transpose of the perturbation weight matrix of the linear layer. For the neural network model to be pruned, the first... The output bias matrix of the linear layer.

[0042] In this embodiment, the present disclosure applies controlled, small noise perturbations to the weights of a trained neural network model. While keeping the input data constant, forward computation is performed on multiple calibration data points, recording the magnitude and direction of output changes before and after the perturbation. This quantifies the local sensitivity of each weight to the neural network model's output. This method accurately characterizes the contribution of weights to the neural network model's performance and improves the reliability of pruning decisions.

[0043] S1024. Based on the output deviation, obtain the noise sensitivity score matrix for each of the weights.

[0044] In this embodiment, the output deviation is attributed inversely to the weight space to construct a weight importance score, providing a highly reliable ranking basis for subsequent pruning decisions.

[0045] Utilizing the differentiability of the linear layer, the output bias is projected onto the weight space:

[0046] in, For the neural network model to be pruned, the first... Weight change estimation matrix of the linear layer For the neural network model to be pruned, the first... The output bias matrix of the linear layer, For the neural network model to be pruned, the first... The input activation matrix of a linear layer, The number of calibration data.

[0047] In this embodiment, the projection operation is a linear backpropagation, which distributes the output changes of each linear layer to each weight according to their contribution, thereby quantifying the functional sensitivity of each weight.

[0048] Based on the projection results, the noise sensitivity score matrix for each of the weights is obtained:

[0049] in, For the neural network model to be pruned, the first... Weight change estimation matrix of the linear layer For the neural network model to be pruned, the first... The noise sensitivity score matrix of the linear layer reflects the degree of influence of the corresponding weight on the output stability of the neural network model under controlled noise perturbation.

[0050] S1025. Based on the input activation matrix, calculate the average activation intensity vector of each output neuron in each of the linear layers, and generate an activation scaling factor.

[0051] To further enhance the robustness and generalization ability of importance assessment, activation statistics are fused to characterize the activity of the output features of the weights. Specifically, based on the input activation matrix, the average activation intensity vector of each output neuron in each of the linear layers is calculated:

[0052] in, For the neural network model to be pruned, the first... In the linear layer, the first The average activation intensity vector of each output neuron; For the first The sample data is input into the first... A vector of output neurons; The number of calibration data; For the neural network model to be pruned, the first... The output dimension of a linear layer. This statistic measures the activity of each output channel in actual inference; channels used frequently should be given higher retention priority.

[0053] The average activation intensity vector Broadcast along the input dimension to generate the weight matrix. Same-dimensional activation scaling factor .

[0054] S1026. The weight matrices, the activation scaling factor, and the noise sensitivity score matrix are fused to obtain the importance score of each weight, as shown below:

[0055] in, For the neural network model to be pruned, the first... The weight matrix of the linear layer; For the neural network model to be pruned, the first... The average activation intensity vector of each output neuron in a linear layer; For the neural network model to be pruned, the first... Noise sensitivity scoring matrix for linear layers; It is a smoothing constant. It is a very small constant to prevent the importance from being distorted when the sensitivity score is zero.

[0056] In this embodiment, the comprehensive score integrates three pieces of information: weight magnitude, activation intensity, and noise sensitivity. This allows for a more comprehensive and accurate characterization of the actual contribution of weights in the neural network model, significantly outperforming existing methods that rely on only a single indicator. This provides a solid and reliable ranking foundation for subsequent high-precision, fine-tuning-free pruning operations.

[0057] In this embodiment, the present disclosure discloses a change-aware weight importance evaluation method. Based on multiple calibration data and independent noise perturbations, it statistically aggregates the output changes corresponding to each weight to comprehensively reflect the impact of weights on the overall stability of the neural network model, thereby obtaining a robust weight importance score. This method effectively improves the generalization ability and accuracy stability of the pruning strategy.

[0058] S103. Prune the weights according to the importance score to generate a sparse neural network model.

[0059] In this embodiment, sparsity is dynamically allocated according to the sensitivity of each network layer to output changes, thereby achieving differentiated pruning intensity control and maximizing the preservation of the model's key capabilities under the global compression objective.

[0060] In this embodiment, the step of pruning the weights based on the importance score to generate a sparse neural network model specifically includes: S1031. Based on the distribution of the importance score in each of the linear layers, determine the global sparsity rate and adaptively allocate the sparsity rate of each linear layer in the network to be pruned, specifically as follows: Calculate the average sensitivity score of each linear layer in the neural network model to be pruned:

[0061] in, For the neural network model to be pruned, the first... In the linear layer, the first Noise sensitivity score with individual weights; For the neural network model to be pruned, the first... The output dimension of a linear layer; For the neural network model to be pruned, the first... The input dimension of a linear layer; For the neural network model to be pruned, the first... The total number of weights in a linear layer. The average sensitivity score comprehensively reflects the overall response strength of the layer to perturbations; a higher score indicates that the layer is more sensitive to the output of the neural network model and has lower redundancy.

[0062] To ensure comparability of sensitivities between different linear layers, the average sensitivity scores of each linear layer are min-max normalized and mapped to... Interval:

[0063] in, For the neural network model to be pruned, the first... Average sensitivity score of the linear layer; The minimum average sensitivity score among all linear layers; The highest average sensitivity score among all linear layers; For the neural network model to be pruned, the first... Sensitivity scores after normalization of the linear layers. This normalization operation enhances the distinction between sensitive and redundant layers, providing a clear control signal for subsequent sparsity rate allocation.

[0064] Given a global target sparsity Based on this, an initial sparsity rate is assigned to each of the linear layers:

[0065] in, For the neural network model to be pruned, the first... The coefficient rate of the initial assignment of the linear layer; For the neural network model to be pruned, the first... Normalized sensitivity scores for the linear layer; The global target sparsity; This is the sparsity adjustment coefficient, controlling the degree to which the sparsity of each layer deviates from the global average. Sensitive layers will receive a lower sparsity (i.e., retain more parameters), while redundant layers will be assigned a higher sparsity (i.e., more aggressive pruning).

[0066] However, the initial allocation described above may not strictly satisfy the global parameter constraint. Therefore, a rescaling mechanism is introduced to ensure that the final sparsity accurately achieves the global target. Thus, the initial sparsity of each of the linear layers is rescaled to obtain the global sparsity:

[0067] in, The preset global target coefficient rate, For the neural network model to be pruned, the first... The coefficient rate of the initial assignment of the linear layer. For the neural network model to be pruned, the first... The total number of weights in a linear layer. The total number of parameters in the model. For the first The sparsity of the layer after rescaling calibration. This rescaling operation maintains the relative differences in sparsity among linear layers while forcing the overall compression ratio to be met, thus balancing the flexibility of pruning with the consistency of the objective.

[0068] In this embodiment, the present disclosure effectively overcomes the problems of excessive compression in sensitive layers and insufficient compression in redundant layers in traditional uniform pruning by adopting an adaptive allocation strategy, thus achieving the pruning goal of "on-demand allocation and precise compression". Without changing the original network topology, fine-grained weight removal operations are performed based on importance scores to generate a sparse model that can be directly deployed, thereby achieving efficient acceleration and storage compression in the inference stage.

[0069] S1032. Based on the global sparsity rate, the weights are pruned to generate a sparse neural network model.

[0070] Based on the global sparsity rate, the weights of each linear layer are sorted in descending order, retaining the top weights. Set the rest to zero and generate a binary mask. ;in, , For the first The sparsity of the layer after rescaling and calibration. For the first Total number of layer weights.

[0071] Based on the binary mask, the weights are pruned to generate a sparse neural network model. The weights of the sparse neural network model are:

[0072] in, For the neural network model to be pruned, the first... The weight matrix of the layer, For the neural network model to be pruned, the first... The mask matrix of the layer, For the neural network model to be pruned, the first... The weight matrix after pruning of the layer.

[0073] In this embodiment, unstructured pruning is employed to preserve the network computation graph structure, with only some weights set to zero. The resulting sparsed neural network model can be directly used with sparse tensor libraries or dedicated hardware to accelerate inference without fine-tuning, significantly reducing deployment costs.

[0074] In this embodiment, this disclosure analyzes the importance distribution and sensitivity of weights at each layer, and differentiates the pruning intensity for different network layers under the constraint of a global sparsity objective, thereby optimizing the overall compression effect without relying on additional fine-tuning. Based on this, unstructured pruning is used to perform model sparsification, removing redundant weights according to the determined importance ranking and layer sparsity rate, while maintaining the original network topology. This disclosure ensures that the compressed model retains its complete executability without modifying the computation graph or making fine-tuning.

[0075] This disclosure presents a high-efficiency pruning strategy combining noise perturbation analysis and change-aware mechanisms. First, controlled noise perturbation is applied to the weights of the neural network model, and the magnitude of this perturbation is monitored to characterize the sensitivity of the weights to model performance. Then, an importance score for each layer's weights is calculated based on the change-aware analysis module to quantify the structural redundancy between different layers. Next, a layer-adaptive sparsity allocation algorithm dynamically optimizes the sparsity ratio. Finally, model compression and parameter sparsification are completed without additional fine-tuning, achieving efficient operation during the inference phase. This disclosure significantly reduces parameter size and computational overhead while maintaining model accuracy, improving the stability and robustness of the pruning process, and possesses significant engineering deployment value and broad application prospects.

[0076] It is understandable that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities.

[0077] It should be noted that the methods of one or more embodiments of this disclosure can be executed by a single device, such as a computer or server. The methods of this embodiment can also be applied in a distributed scenario, where multiple devices cooperate to complete the process. In such a distributed scenario, one of these devices may execute only one or more steps of the methods of one or more embodiments of this disclosure, and the multiple devices will interact with each other to complete the method described.

[0078] It should be noted that the above description pertains to specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than those shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0079] Based on the same inventive concept, corresponding to the pruning method of a neural network model in any of the above embodiments, this disclosure also provides a pruning device for a neural network model. For example... Figure 2 As shown, the above-mentioned device includes: The acquisition module 201 is configured to: acquire the neural network model to be pruned and calibration data; wherein, the calibration data includes text sequences randomly extracted from an unsupervised corpus; The calculation module 203 is configured to: calculate the importance score of the weights in the neural network model to be pruned based on the calibration data; The pruning module 204 is configured to prune the weights according to the importance score to generate a sparse neural network model.

[0080] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, when implementing one or more embodiments of this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0081] The apparatus described above is used to implement a pruning method for a corresponding neural network model in the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0082] This disclosure also provides a data processing method that enables the pruned model to be stored and computed in a sparse format, thereby significantly reducing memory usage and bandwidth pressure, and achieving a significant improvement in inference speed while ensuring accuracy.

[0083] Figure 3 A flowchart illustrating an exemplary method provided in an embodiment of this disclosure is shown, such as... Figure 3 As shown, the method may further include the following steps.

[0084] Acquire the target data and the sparse neural network model; The target data is processed based on the sparse neural network model.

[0085] In this embodiment, the weights of the sparse neural network model are obtained through training using any of the aforementioned neural network model pruning methods. Specifically, based on the neural network model to be pruned and calibration data, an importance score is calculated for the weight parameters in the neural network model to be pruned, and the weight parameters are pruned according to the importance score to obtain the sparse neural network model. The weights of the sparse neural network model are stored in a standard sparse format (which may be Compressed Sparse Line Format, CSR), which significantly reduces memory usage by retaining only non-zero weights and their position indices.

[0086] When performing inference calculations, sparse computing libraries (such as Intel MKL, OpenBLAS-Sparse, etc.) only need to read and process the non-zero values ​​of the weights of the sparse neural network model, completely skipping the access and multiplication-addition operations of zero values, thereby greatly reducing the memory bandwidth pressure on the CPU and improving computing efficiency.

[0087] Furthermore, the hardware features of modern CPUs (such as memory prefetching, out-of-order execution, and the AVX-512 vectorized instruction set) work efficiently with sparse computing libraries to effectively mitigate the overhead caused by irregular memory accesses. Moreover, when the model sparsity exceeds 50%, the saved computational and memory bandwidth far outweigh the additional access overhead, ultimately achieving an end-to-end inference speedup of 1.3 to 1.6 times.

[0088] In some embodiments, the sparse neural network model can be trained using any embodiment or arrangement and combination of the aforementioned neural network model pruning method, and can have the technical effects of the corresponding embodiments, which will not be elaborated here.

[0089] This disclosure also provides a data processing apparatus. For example... Figure 4 As shown, the device includes: The acquisition module 401 is configured to acquire target data and a sparse neural network model. The processing module 402 is configured to process the target data based on the sparse neural network model.

[0090] In this embodiment, the weights of the sparse neural network model are obtained through training using any of the aforementioned neural network model pruning methods. Specifically, based on the neural network model to be pruned and calibration data, an importance score is calculated for the weight parameters in the neural network model to be pruned, and the weight parameters are pruned according to the importance score to obtain the sparse neural network model. The weights of the sparse neural network model are stored in a standard sparse format (which may be Compressed Sparse Line Format, CSR), which significantly reduces memory usage by retaining only non-zero weights and their position indices.

[0091] When performing inference calculations, sparse computing libraries (such as Intel MKL, OpenBLAS-Sparse, etc.) only need to read and process these non-zero values, completely skipping the access and multiplication-addition operations of zero values, thereby greatly reducing the memory bandwidth pressure on the CPU and improving computing efficiency.

[0092] Furthermore, the hardware features of modern CPUs (such as memory prefetching, out-of-order execution, and the AVX-512 vectorized instruction set) work efficiently with sparse computing libraries to effectively mitigate the overhead caused by irregular memory accesses. Moreover, when the model sparsity exceeds 50%, the saved computational and memory bandwidth far outweigh the additional access overhead, ultimately achieving an end-to-end inference speedup of 1.3 to 1.6 times.

[0093] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, when implementing one or more embodiments of this disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0094] The apparatus described above is used to implement a corresponding data processing method in the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0095] Figure 5 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0096] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure.

[0097] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this disclosure are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0098] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0099] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0100] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0101] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this disclosure, and not necessarily all the components shown in the figures.

[0102] The electronic devices described above are used to implement the corresponding methods in the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0103] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0104] Based on the same inventive concept, corresponding to the pruning method and data processing method of the neural network model in any of the above embodiments, this disclosure also provides a computer program product, which includes one or more computer programs. In some embodiments, the one or more computer programs are executable by one or more processors to cause the one or more processors to execute the pruning method and data processing method of the neural network model. Corresponding to the execution entity for each step in each embodiment of the pruning method and data processing method of the neural network model, the processor executing the corresponding step may belong to the corresponding execution entity. The computer program product of the above embodiments is used to cause the processor to execute the pruning method and data processing method of the neural network model as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0105] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.

[0106] Additionally, to simplify the description and discussion, and to avoid obscuring one or more embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring one or more embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which one or more embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that one or more embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0107] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0108] This disclosure includes one or more embodiments intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A pruning method for a neural network model, characterized in that, The method includes: Obtain the neural network model to be pruned and calibration data; wherein, the calibration data includes text sequences randomly extracted from an unsupervised corpus; Based on the calibration data, the importance score of the weights in the neural network model to be pruned is calculated; The weights are pruned based on the importance score to generate a sparse neural network model.

2. The method according to claim 1, characterized in that, Based on the calibration data, the importance scores of the weights in the neural network model to be pruned are calculated, including: The calibration data is input into the neural network model to be pruned to obtain the input activation matrix corresponding to each linear layer in the neural network model to be pruned; wherein, the neural network model to be pruned contains at least one linear layer; Noise perturbation is applied to the original weight matrix of each linear layer to generate a perturbed weight matrix; The output deviation is calculated based on the input activation matrix, the original weight matrix, and the perturbed weight matrix. Based on the output deviation, a noise sensitivity score matrix for each of the weights is obtained; Based on the input activation matrix, the average activation intensity vector of each output neuron in each of the linear layers is calculated, and an activation scaling factor is generated. The importance score of each weight is obtained by fusing the weights, the activation scaling factor, and the noise sensitivity score matrix.

3. The method according to claim 2, characterized in that, Based on the importance score, the weights are pruned to generate a sparse neural network model, including: Based on the distribution of the importance scores in each of the linear layers, the global sparsity rate is determined; Based on the global sparsity rate, the weights are pruned to generate a sparse neural network model.

4. The method according to claim 2, characterized in that, The output deviation is calculated based on the input activation matrix, the original weight matrix, and the perturbed weight matrix, including: The original output is calculated using the original weight matrix: The perturbation output is calculated using the perturbated weights: Based on the original output and the disturbance output The output deviation was calculated. : in, For the neural network model to be pruned, the first... The original output matrix of the linear layer, For the neural network model to be pruned, the first... The perturbation output matrix of the linear layer. For the neural network model to be pruned, the first... The input activation matrix of a linear layer, For the neural network model to be pruned, the first... The transpose of the original weight matrix of the linear layer. For the neural network model to be pruned, the first... The transpose of the perturbation weight matrix of the linear layer. For the neural network model to be pruned, the first... The output bias matrix of the linear layer.

5. The method according to claim 3, characterized in that, Determining the global sparsity rate based on the distribution of the importance scores across each of the linear layers includes: Calculate the average sensitivity score of each linear layer in the neural network model to be pruned: in, For the neural network model to be pruned, the first... In the linear layer, the first Noise sensitivity score with individual weights; For the neural network model to be pruned, the first... The output dimension of a linear layer; For the neural network model to be pruned, the first... The input dimension of a linear layer; For the neural network model to be pruned, the first... The total number of weights in a linear layer; The average sensitivity scores of each of the linear layers are normalized: in, For the neural network model to be pruned, the first... Average sensitivity score of the linear layer; The minimum average sensitivity score among all linear layers; The highest average sensitivity score among all linear layers; For the neural network model to be pruned, the first... Sensitivity scores after normalization of the linear layer; Assign an initial sparsity to each of the linear layers: in, For the neural network model to be pruned, the first... The coefficient rate of the initial assignment of the linear layer; For the neural network model to be pruned, the first... Normalized sensitivity scores for the linear layer; The preset global target sparsity; This is the sparsity adjustment coefficient; The initial sparsity of each of the linear layers is rescaled to obtain the global sparsity: in, The preset global target coefficient rate, For the neural network model to be pruned, the first... The coefficient rate of the initial assignment of the linear layer. For the neural network model to be pruned, the first... The total number of weights in a linear layer. The total number of parameters in the model. For the neural network model to be pruned, the first... The sparsity of the layer after rescaling and calibration.

6. The method according to claim 3, characterized in that, The step of pruning the weights according to the global sparsity rate to generate a sparse neural network model includes: Based on the global sparsity, the weights of each linear layer are sorted in descending order, and a binary mask is generated. Based on the binary mask, the weights are pruned to generate a sparse neural network model.

7. The method according to claim 4, characterized in that, The step of obtaining the noise sensitivity score matrix for each weight based on the output deviation includes: Project the output bias onto the weight space: in, For the neural network model to be pruned, the first... Weight change estimation matrix of the linear layer For the neural network model to be pruned, the first... The output bias matrix of the linear layer, For the neural network model to be pruned, the first... The input activation matrix of a linear layer, The number of calibration data; Based on the projection results, the noise sensitivity score matrix for each of the weights is obtained: in, For the neural network model to be pruned, the first... Weight change estimation matrix of the linear layer For the neural network model to be pruned, the first... Noise sensitivity scoring matrix for linear layers; Based on the input activation matrix, the average activation intensity vector of each output neuron in each of the linear layers is calculated, and an activation scaling factor is generated, including: in, For the neural network model to be pruned, the first... In the linear layer, the first The average activation intensity vector of each output neuron; For the first The sample data is input into the first... A vector of output neurons; The number of calibration data; For the neural network model to be pruned, the first... The output dimension of a linear layer; The average activation intensity vector Broadcast along the input dimension to generate the weight matrix. Same-dimensional activation scaling factor ; The importance score of each weight is obtained by fusing the weights, the activation scaling factor, and the noise sensitivity score matrix, including: in, For the neural network model to be pruned, the first... The weight matrix of the linear layer; For the neural network model to be pruned, the first... The average activation intensity vector of each output neuron in a linear layer; For the neural network model to be pruned, the first... Noise sensitivity scoring matrix for linear layers; It is a smoothing constant. It is a constant.

8. A data processing method, characterized in that, The method includes: Acquire target data and a sparse neural network model trained according to any one of claims 1-7; The target data is processed based on the sparse neural network model.

9. A pruning device for a neural network model, characterized in that, The device includes: The acquisition module is configured to acquire the neural network model to be pruned and calibration data; wherein the calibration data includes text sequences randomly extracted from an unsupervised corpus. The calculation module is configured to: calculate the importance score of the weights in the neural network model to be pruned based on the calibration data; The pruning module is configured to prune the weights based on the importance score to generate a sparse neural network model.

10. A data processing apparatus, characterized in that, The device includes: The acquisition module is configured to: acquire target data and a sparse neural network model trained according to any one of claims 1-7; The processing module is configured to process the target data based on the sparse neural network model.