Neural network sparsification method and device, electronic equipment and storage medium
By using structured pruning algorithms and encoding-decoding modules in neural networks, the weight parameters are optimized, which solves the problems of insufficient resources and unbalanced computing load on edge devices, and achieves efficient hardware implementation and computing efficiency.
Patent Information
- Application Number
- CN202510014391.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
AI Technical Summary
When deploying on edge devices with limited resources, existing neural networks face insufficient storage and computing resources caused by huge weight parameters and unbalanced computing load caused by unstructured pruning algorithms.
A structured pruning algorithm is proposed, which divides the weight parameter matrix of neural network weight parameter, deletes the weight parameters with the absolute value sorting, and uses the encoder and decoder composed of linear layers to compress and reconstruct the sparse weights, and optimizes hardware storage and calculation.
It realizes improving the operation efficiency of neural networks on resource-constrained devices, reducing hardware storage and computing overhead, overcoming the problem of unbalanced computing load, and maintaining computing accuracy.
Smart Images

Figure CN119940451A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning technology, and in particular to a neural network sparsification method, device, electronic device and storage medium. Background Art
[0002] In recent years, with the rapid development of deep learning technology, neural networks have shown excellent potential in various computer vision tasks. However, the huge weight parameters of neural networks make it challenging to deploy on resource-constrained edge devices. To address this problem, some existing neural network hardware acceleration algorithms use model sparsification technology to reduce frequent hardware off-chip data reading by pruning redundant weight parameters while maintaining relatively high accuracy (only reduced by 1% to 2%). Although this pruning algorithm can usually significantly improve the sparsity of the model (sparseness can reach 95%), thereby greatly reducing the floating-point calculation amount of the model, the spatial uncertainty of non-zero weights caused by irregular pruning requires additional storage and addressing overhead to handle this random distribution. In addition, unstructured pruning technology also brings new challenges to computing - in parallel computing, the sparse weights that have been randomly pruned are assigned to different hardware processing units for matrix calculations. Due to the random distribution of non-zero elements, the number of weights assigned to different hardware processing units is different, which inevitably leads to an imbalance in the overall workload, thereby reducing resource utilization and parallel efficiency of computing. Summary of the invention
[0003] The main purpose of the embodiments of the present application is to propose a neural network sparsification method, device, electronic device and storage medium to improve the operating efficiency of the neural network on resource-constrained devices.
[0004] To achieve the above object, an embodiment of the present application provides a neural network sparsification method, which includes the following steps:
[0005] Performing structured pruning on the weight parameters of each layer in the weight parameter matrix of the neural network to delete a number of weight parameters ranked low in absolute value in each column of the weight parameter matrix to obtain an original pruned weight matrix; wherein each column of the original pruned weight matrix includes a number of non-zero weight parameters of the same number;
[0006] Configuring corresponding row indices for each of the non-zero weight parameters in the original pruning weight matrix and then compressing the original pruning weight matrix using an encoder composed of linear layers;
[0007] Storing the original pruning weight matrix compressed by the encoder in an external memory of an accelerator; wherein the accelerator is used to run the neural network;
[0008] Reading the compressed original pruning weight matrix from the external memory to the internal memory of the accelerator, and then decoding the compressed original pruning weight matrix using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix;
[0009] The accelerator is used to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix.
[0010] In some embodiments, before performing structured pruning on the weight parameters of each layer in the weight parameter matrix of the neural network to delete a number of weight parameters with the lowest absolute value in each column of the weight parameter matrix to obtain the original pruned weight matrix, the method further comprises the following steps:
[0011] Training the neural network until the neural network converges;
[0012] The weight parameters when the neural network converges are set as the weight parameter matrix.
[0013] In some embodiments, the structured pruning of the weight parameters of each layer in the weight parameter matrix of the neural network to delete several weight parameters with the lowest absolute value in each column of the weight parameter matrix to obtain the original pruned weight matrix includes the following steps:
[0014] Sort the weight parameters in each column of the weight parameter matrix of the neural network by absolute value;
[0015] Determining a binary weight mask matrix; wherein each parameter in the binary weight mask matrix is used to determine whether a corresponding weight parameter in the weight parameter matrix is retained or deleted;
[0016] According to the binary weight mask matrix, a plurality of weight parameters whose absolute values are ranked at the bottom in each column of the weight parameter matrix are deleted to retain a plurality of non-zero weight parameters whose absolute values are ranked at the top in each column of the weight parameter matrix and whose number is the same, to obtain a candidate pruning weight matrix;
[0017] Training the neural network provided with the candidate pruning weight matrix until the neural network converges;
[0018] Return to the step of sorting the weight parameters in each column of the weight parameter matrix of the neural network according to the absolute value, until the neural network is trained a set number of times, and use the current candidate pruning weight matrix as the original pruning weight matrix.
[0019] In some embodiments, configuring the corresponding row index for each of the non-zero weight parameters in the original pruning weight matrix comprises the following steps:
[0020] A row index in the COO format is configured for each of the non-zero weight parameters in the original pruning weight matrix.
[0021] In some embodiments, decoding the compressed original pruning weight matrix using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix comprises the following steps:
[0022] Decoding the compressed original pruning weight matrix using the decoder composed of linear layers;
[0023] A dot product operation is performed on the non-zero weight parameters in each column of the decoded original pruning weight matrix and the corresponding row index with the corresponding input activation value to obtain the reconstructed pruning weight matrix.
[0024] In some embodiments, decoding the compressed original pruning weight matrix using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix comprises the following steps:
[0025] The decoder composed of linear layers and reconstruction loss are used to decode the compressed original pruning weight matrix to reconstruct the reconstructed pruning weight matrix.
[0026] In some embodiments, the accelerator is disposed in at least one of a gateway, a camera, or a drone;
[0027] The step of using the accelerator to read the reconstructed pruning weight matrix from the internal memory and then running the neural network provided with the reconstructed pruning weight matrix comprises at least one of the following steps:
[0028] Using the accelerator disposed in the gateway to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix to process and transfer information sent by a terminal; wherein the terminal is a device that establishes a communication connection with the gateway;
[0029] Alternatively, the accelerator disposed in the camera is used to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix to process the image captured by the camera;
[0030] Alternatively, the accelerator disposed in the drone is used to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix to process the sensor information acquired by the drone.
[0031] To achieve the above-mentioned purpose, another aspect of the embodiment of the present application provides a neural network sparsification device, the device comprising:
[0032] A weight pruning unit is used to perform structured pruning on the weight parameters of each layer in the weight parameter matrix of the neural network, so as to delete a number of weight parameters with the lowest absolute value in each column of the weight parameter matrix, and obtain an original pruned weight matrix; wherein each column of the original pruned weight matrix includes a number of non-zero weight parameters of the same number;
[0033] A weight compression unit, configured to configure a corresponding row index for each of the non-zero weight parameters in the original pruning weight matrix and then compress the original pruning weight matrix using an encoder composed of a linear layer;
[0034] A weight storage unit, used to store the original pruned weight matrix compressed by the encoder into an external memory of an accelerator; wherein the accelerator is used to run the neural network;
[0035] A weight reconstruction unit, configured to read the compressed original pruning weight matrix from the external memory to the internal memory of the accelerator, and then decode the compressed original pruning weight matrix using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix;
[0036] A neural network running unit is used to use the accelerator to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix.
[0037] To achieve the above objective, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program and data, and the processor implementing the above method when executing the computer program.
[0038] To achieve the above objective, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.
[0039] The embodiments of the present application include at least the following beneficial effects:
[0040] The present application can perform structured pruning on the weight parameters of each layer in the weight parameter matrix of a neural network to delete several weight parameters whose absolute values are ranked low in each column of the weight parameter matrix to obtain an original pruned weight matrix; wherein each column of the original pruned weight matrix includes several non-zero weight parameters of the same number; each non-zero weight parameter in the original pruned weight matrix is configured with a corresponding row index and then the original pruned weight matrix is compressed using an encoder composed of a linear layer; the original pruned weight matrix compressed by the encoder is stored in an external memory of an accelerator; wherein the accelerator is used to run the neural network; the compressed original pruned weight matrix is read from the external memory to the internal memory of the accelerator, and then the compressed original pruned weight matrix is decoded using a decoder composed of a linear layer to reconstruct a reconstructed pruned weight matrix; the reconstructed pruned weight matrix is read from the internal memory using the accelerator and then the neural network provided with the reconstructed pruned weight matrix is run. The present application can reduce unimportant weight connections within the neural network by performing structured pruning on weight parameters, thereby reducing hardware storage overhead and overall calculation, and further reducing power consumption; and the structured pruning of the present application overcomes the characteristics of unbalanced computational load caused by traditional irregular pruning algorithms, greatly improves the resource utilization of each module, and enables sparse weights to be processed in parallel more effectively on hardware accelerators. On the other hand, the pruned weight parameters are further compressed by an encoder composed of linear layers, and the weight parameters are stored in a more compact form. When matrix calculations are required, the compressed weight parameters are restored and reconstructed using a decoder composed of linear layers to improve calculation accuracy; through the synergy of the encoder and the decoder, the present application achieves a significant reduction in the amount of calculation while reducing the storage of model weights, thereby reducing hardware power consumption, but does not cause a loss of accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0042] Figure 1 A schematic diagram of a process flow of a neural network sparsification method provided in an embodiment of the present application;
[0043] Figure 2 An example flow chart of an irregular pruning algorithm provided in an embodiment of the present application;
[0044] Figure 3 Pseudo code diagram of the structured pruning algorithm provided in the embodiment of the present application;
[0045] Figure 4 An example flow chart of a structured pruning algorithm provided in an embodiment of the present application;
[0046] Figure 5 A calculation logic flow chart of a hardware unit after structured pruning provided in an embodiment of the present application;
[0047] Figure 6 A calculation flow chart of the encoding-decoding module provided in an embodiment of the present application;
[0048] Figure 7 A schematic diagram of the structure of a neural network sparsification device provided in an embodiment of the present application;
[0049] Figure 8 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0051] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0052] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0054] Before describing the embodiments of the present application in detail, some related technologies that may be involved in the embodiments of the present application are first described as follows:
[0055] Related technology 1 involves a neural network weight pruning algorithm. The difference between this application and related technology 1 is that the weight pruning algorithm in related technology 1 is an unstructured pruning algorithm. Although it can achieve extremely high weight sparsity, the non-zero elements after pruning are unevenly distributed, which cannot play its advantages in hardware implementation and affects the efficiency of parallel computing. This method adopts a structured pruning algorithm, which can achieve high sparsity while ensuring that the workload of the hardware computing unit is the same, thereby improving the overall computing efficiency.
[0056] Related technology 2 relates to a structured pruning algorithm for neural network weights. The difference between this application and related technology 2 is that related technology 2 proposes a cycle-oriented pruning method to improve the resource utilization of each computing unit. However, the method of related technology 2 ideally divides an equal number of filter weights between different computing units, without considering the actual logic of hardware calculations, and also ignores the additional overhead caused by weight addressing. This method uses column dimension division to divide data in the same column into the same computing unit, optimizes the weight data storage format, and thus simplifies the addressing and calculation logic of each computing unit.
[0057] Related technology 3 proposes a coding and decoding module. The difference between this application and related technology 3 is that related technology 3 proposes a coding module for the Transformer model, which is used to compress different data, and the module is not globally shared. Each data has a separate coding-decoding module, and the decoding module needs to be frequently loaded from outside the chip for corresponding calculations. The present application adopts a global sharing design concept, which can provide the same decoding matrix data for different compression weights throughout the calculation cycle. Therefore, the decoding module can realize data reuse after loading the decoding matrix for the first time, without frequently accessing the off-chip memory in each clock cycle, thereby effectively reducing the frequency of off-chip data transmission.
[0058] This application aims to solve several key technical problems faced in the hardware implementation of neural networks. First, traditional neural network models usually face the challenge of limited storage and computing resources when deployed to edge devices. Due to the large model parameters, frequent off-chip data reading will increase power consumption and latency, limiting efficient deployment on resource-constrained devices. This application introduces a sparse algorithm to reduce the storage requirements and computational complexity of the model, thereby adapting to the resource limitations of edge devices. Secondly, although the traditional unstructured pruning algorithm can significantly improve the sparsity of the model, its irregular pruning method makes it difficult to process sparse weights in parallel on the hardware accelerator. This will lead to an unbalanced workload between computing units, reducing resource utilization and computing efficiency. This application effectively alleviates this problem by designing a hardware-friendly structured pruning algorithm, so that sparse weights can be more evenly distributed to hardware processing units for parallel computing, thereby improving computing efficiency. In addition, another technical problem is the high power consumption caused by frequent hardware access to off-chip storage. This application integrates a learnable, lightweight and globally shared weight encoding-decoding module, which further compresses and stores the pruned weights, reducing the need for the hardware to frequently access external storage, thereby reducing power consumption. The introduction of this module effectively reduces the burden of hardware accelerators in computing and storage, and improves the energy efficiency of the overall system.
[0059] In order to achieve efficient and reliable deployment of neural networks on resource-constrained edge devices, it is necessary to explore a hardware-friendly neural network sparsification solution.
[0060] The embodiments of the present application provide a neural network sparsification method, device, electronic device and storage medium. The technical solution of the present application includes: performing structured pruning on the weight parameters of each layer in the weight parameter matrix of the neural network to delete several weight parameters whose absolute values are ranked later in each column of the weight parameter matrix to obtain the original pruned weight matrix; wherein each column of the original pruned weight matrix includes several non-zero weight parameters of the same number; configuring corresponding row indexes for each non-zero weight parameter in the original pruned weight matrix and then compressing the original pruned weight matrix using an encoder composed of linear layers; storing the original pruned weight matrix compressed by the encoder to the external memory of the accelerator; wherein the accelerator is used to run the neural network; reading the compressed original pruned weight matrix from the external memory to the internal memory of the accelerator, and then decoding the compressed original pruned weight matrix using a decoder composed of linear layers to reconstruct the reconstructed pruned weight matrix; using the accelerator to read the reconstructed pruned weight matrix from the internal memory and then run the neural network provided with the reconstructed pruned weight matrix. The present application can reduce unimportant weight connections within the neural network by performing structured pruning on weight parameters, thereby reducing hardware storage overhead and overall calculation, and further reducing power consumption; and the structured pruning of the present application overcomes the characteristics of unbalanced computational load caused by traditional irregular pruning algorithms, greatly improves the resource utilization of each module, and enables sparse weights to be processed in parallel more effectively on hardware accelerators. On the other hand, the pruned weight parameters are further compressed by an encoder composed of linear layers, and the weight parameters are stored in a more compact form. When matrix calculations are required, the compressed weight parameters are restored and reconstructed using a decoder composed of linear layers to improve calculation accuracy; through the synergy of the encoder and the decoder, the present application achieves a significant reduction in the amount of calculation while reducing the storage of model weights, thereby reducing hardware power consumption, but does not cause a loss of accuracy.
[0061] The embodiments of the present application provide a neural network sparsification method, device, electronic device and storage medium, and relate to the field of deep learning technology. The neural network sparsification method, device, electronic device and storage medium provided in the embodiments of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured to provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and cloud servers for basic cloud computing services such as big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a knowledge extraction method, etc., but is not limited to the above forms.
[0062] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0063] Reference Figure 1 , the embodiment of the present application provides a neural network sparsification method, which may include but is not limited to S110 to S150, as follows:
[0064] S110: Perform structured pruning on the weight parameters of each layer in the weight parameter matrix of the neural network to delete several weight parameters with lower absolute value order in each column of the weight parameter matrix to obtain an original pruned weight matrix; wherein each column of the original pruned weight matrix includes a same number of non-zero weight parameters.
[0065] Further, S110 may include the following steps S111 to S115:
[0066] S111: sorting the weight parameters in each column of the weight parameter matrix of the neural network according to their absolute values;
[0067] S112: Determine a binary weight mask matrix; wherein each parameter in the binary weight mask matrix is used to determine whether a corresponding weight parameter in the weight parameter matrix is retained or deleted;
[0068] S113: deleting a number of weight parameters whose absolute values are ranked at the bottom in each column of the weight parameter matrix according to the binary weight mask matrix, so as to retain a number of non-zero weight parameters whose absolute values are ranked at the top in each column of the weight parameter matrix and whose number is the same, to obtain a candidate pruning weight matrix;
[0069] S114: training the neural network provided with the candidate pruning weight matrix until the neural network converges;
[0070] S115: Return to the step of sorting the weight parameters in each column of the weight parameter matrix of the neural network according to the absolute value, until the neural network is trained a set number of times, and the current candidate pruning weight matrix is used as the pruning weight matrix.
[0071] In some embodiments, before step S110, the embodiment of the present application may further include the following steps S101-S102:
[0072] S101: training the neural network until the neural network converges;
[0073] S102: setting each weight parameter when the neural network converges as the weight parameter matrix.
[0074] S120: configuring corresponding row indices for each of the non-zero weight parameters in the original pruning weight matrix and then compressing the original pruning weight matrix using an encoder composed of linear layers.
[0075] Further, configuring the row index corresponding to each of the non-zero weight parameters in the original pruning weight matrix in S120 includes the following steps:
[0076] S121: configuring a row index in a COO format corresponding to each of the non-zero weight parameters in the original pruning weight matrix.
[0077] S130: Storing the original pruning weight matrix compressed by the encoder into an external memory of an accelerator; wherein the accelerator is used to run the neural network.
[0078] S140: Reading the compressed original pruning weight matrix from the external memory to the internal memory of the accelerator, and then decoding the compressed original pruning weight matrix using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix.
[0079] Furthermore, in S140, decoding the compressed original pruning weight matrix using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix includes the following steps:
[0080] S141: Decoding the compressed original pruning weight matrix using the decoder composed of linear layers;
[0081] S142: Perform a dot product operation on the non-zero weight parameters in each column of the decoded original pruning weight matrix and the corresponding row index with the corresponding input activation value to obtain the reconstructed pruning weight matrix.
[0082] As another optional implementation, the method of decoding the compressed original pruning weight matrix by using a decoder composed of linear layers in S140 to reconstruct a reconstructed pruning weight matrix includes the following steps:
[0083] S143: Utilize the decoder composed of linear layers and reconstruction loss decoding to decode the compressed original pruning weight matrix to reconstruct the reconstructed pruning weight matrix.
[0084] S150: Using the accelerator to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix.
[0085] Furthermore, the accelerator of this embodiment may be provided in at least one of a gateway, a camera or a drone;
[0086] Then S150 includes at least one of the following steps:
[0087] S151: using the accelerator disposed in the gateway to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix to process and transfer information sent by a terminal; wherein the terminal is a device that establishes a communication connection with the gateway;
[0088] S152: using the accelerator disposed in the camera to read the reconstructed pruning weight matrix from the internal memory and then running the neural network provided with the reconstructed pruning weight matrix to process the image captured by the camera;
[0089] S153: Using the accelerator disposed in the drone to read the reconstructed pruning weight matrix from the internal memory and then running the neural network provided with the reconstructed pruning weight matrix to process the sensor information acquired by the drone.
[0090] Next, the solution of the embodiment of the present application will be introduced and explained in detail with reference to specific application examples.
[0091] First, the scheme of this embodiment is summarized: This embodiment proposes a neural network sparsification method suitable for hardware implementation. Different from the traditional scheme, this embodiment integrates two modules, the efficient pruning algorithm for hardware computing load balancing and the encoding-decoding module for low-power hardware implementation. Among them, the pruning algorithm is mainly used to reduce the unimportant weight connections between the neural network, thereby reducing the hardware storage overhead and the overall calculation, and then reducing the power consumption. In addition, the pruning algorithm of this embodiment overcomes the characteristics of the unbalanced computing load caused by the traditional irregular pruning algorithm, greatly improves the resource utilization of each module, and enables the sparse weights to be processed in parallel more effectively on the hardware accelerator. On the other hand, the weight encoding-decoding module further compresses the pruned weights and stores the weights in a more compact form. When matrix calculation is required, the decoding module is used to restore the compressed weights to ensure the calculation accuracy. Through the synergy of these two modules, it is achieved that the amount of calculation is significantly reduced while reducing the model weight storage, thereby reducing the hardware power consumption, but it does not cause the loss of accuracy. This embodiment adopts a multi-stage training strategy, so that each module can be fine-tuned during the optimization process, ensuring the best inference speed and performance in hardware deployment, and providing a new solution for achieving efficient and low-power hardware inference.
[0092] Specifically, this embodiment may include the following solutions:
[0093] 1. Implement a structured pruning algorithm for hardware workload balancing.
[0094] Traditional hardware accelerators efficiently improve performance through pipelining and parallelization when processing dense matrices. However, some irregular pruning methods currently lead to the distribution of non-zero elements in sparse models often lacking regularity, which brings unevenness to data access and calculation. This unstructured sparsity increases the complexity of memory access and leads to a decrease in cache hit rate. In addition, the hardware needs to process a large number of branch judgments to skip zero elements, further reducing computing efficiency. Figure 2 As shown in Figure 1, irregular sparse matrices make the workload of the processing units unbalanced, and cannot fully utilize the parallel computing advantages of hardware accelerators. Solving these challenges usually requires designing dedicated sparse matrix accelerators or using complex scheduling algorithms to balance the load, but this increases the complexity of hardware design and resource consumption.
[0095] Considering the limitations of unstructured pruning in hardware implementation, this embodiment proposes a hardware-friendly pruning algorithm to reduce storage and computing overhead. The algorithm uses a column-oriented structured pruning method to make full use of the sparsity of the model to improve the computing speed and efficiency of the hardware accelerator.
[0096] The pruning algorithm proposed in this embodiment mainly adopts a multi-round iterative pruning strategy to gradually improve the sparsity of the model. Figure 3 In this embodiment, the proposed pruning algorithm is described in detail. For a randomly initialized neural network f(x; θ), this embodiment does not perform pruning in the first round, but trains the model until convergence. Starting from the second round, this embodiment sorts the trained weights by absolute value and prunes the weight matrix column by column. In order to achieve a preset pruning sparsity, this embodiment only retains the part of the data with the largest absolute value in each column of the weight matrix, thereby achieving workload balancing for hardware computing. In addition, during the iterative pruning process, this embodiment maintains a binary weight mask M∈{0,1} k , which is used to select whether the corresponding weight is to be kept or discarded. After each adjustment, this embodiment obtains a subnetwork f(x; M⊙θ), and then reinitializes and trains it until convergence. After K iterations, the final neural network retains d in each column dimension of the weight matrix k (e.g. 8) weights to achieve model compression. Figure 4 As shown, the pruning method proposed in this embodiment ensures that the number of non-zero weights in each column remains consistent, thereby achieving workload balancing of the hardware computing unit and improving parallel processing efficiency.
[0097] In addition, thanks to the pruning algorithm, the present embodiment achieves extremely high sparsity weight data (≥90%) in the neural network, significantly reducing data density. In order to fully utilize the advantages of sparse computing, the present embodiment further optimizes the sparse operation logic of each computing unit in the accelerator, including the storage format. The conventional coordinate storage format (COO) requires the storage of complete coordinate information for each non-zero element, including its value and the corresponding row and column indexes. In contrast, the design of the present embodiment simplifies the process by reducing storage overhead by retaining only non-zero values and their associated row indices. The present embodiment slices the weight matrix by columns, thereby grouping all non-zero elements. During the calculation process, each hardware computing unit dynamically maps input data with row indices to perform effective multiplication of related elements. As Figure 5As shown, this embodiment eliminates the zero elements after pruning during the storage process, and aggregates the remaining non-zero elements in each column. Next, this embodiment stores the corresponding row coordinate information using the optimized COO format. During the processing, each computing unit can quickly locate the relevant input and perform multiplication calculations, avoiding the power consumption overhead generated by the search.
[0098] 2. Weight parameter encoding-decoding.
[0099] Although this embodiment significantly improves the sparsity of the model through the pruning algorithm and achieves high efficiency of parallel computing of hardware units, considering that there are still redundant parameters in different row data, the data handling cost is high. This is because there are still a large number of unpruned parameters in the row dimension, which causes frequent data transmission between different computing units and unbalanced resource scheduling, further affecting the overall computing speed and system performance.
[0100] In order to solve the bottleneck problem of neural network weight storage and transmission in high-performance computing, this embodiment proposes a lightweight, learnable and globally shared encoding-decoding module, which can compress the pruned weights so that the compressed weights can be stored in a smaller volume in the chip external memory. Figure 6 The specific process of the weight encoding-decoding module is shown in detail, where Figure 6 The original weights are the uncompressed original pruned weight matrix. Specifically, first, this embodiment uses an encoder composed of linear layers to compress the weights that have been pruned, compressing the data dimension to a more compact dimension (such as 256->8), thereby achieving data compression. Subsequently, this embodiment saves the compressed and quantized weight data to the external storage (DRAM) of the accelerator for easy reading by the hardware accelerator. When the accelerator needs to use weights for matrix calculations, the hardware system dynamically loads the compressed weight matrix from the external storage to the accelerator. Afterwards, through a decoder module that is also composed of linear layers, the compressed data is reconstructed into the original weight matrix and saved to the internal memory (SRAM) of the accelerator for subsequent calculations. This method designed in this embodiment makes full use of the redundancy in the weight data, significantly reduces the power consumption generated by frequently moving data from the outside of the accelerator to the inside of the accelerator, replaces expensive data movement with low-cost data calculation, and improves the utilization rate of the internal storage resources of the accelerator, thereby providing an efficient solution for deep learning models in low-power hardware implementation.
[0101] In addition, in order to ensure the similarity between the reconstructed weight data after the decoder and the original weight data, this embodiment introduces reconstruction loss during the training process to minimize the difference between the original weight and the reconstructed weight (for example, ||WW′||0). This enables the encoding module to be jointly optimized with the model weights, thereby reducing the data transfer cost between the accelerator's external storage and internal storage while maintaining model accuracy. After integrating the encoder module into the pre-trained model, this embodiment fine-tunes the entire system to achieve joint training of the encoder and model weights.
[0102] The key technical points of this embodiment are summarized as follows:
[0103] This embodiment proposes a jointly designed digital circuit-friendly and computationally efficient neural network sparsification method. This method fully explores the spatial sparsity of the neural network, and through the synergy of the structured sparse algorithm and the weight encoding-decoding module, it significantly reduces the amount of calculation while reducing the model weight storage, thereby reducing hardware power consumption. In addition, the algorithm is simple, efficient, and widely applicable. It can be applied to a variety of mainstream neural network architectures (CNN, SNN, and Transformer, etc.), providing new possibilities for digital circuit design and computational efficiency.
[0104] The method consists of the following two parts:
[0105] (1) A hardware-friendly structured pruning algorithm is used to prune each column of the weight matrix to the same number, thereby ensuring workload balance of the hardware computing unit and improving the parallel efficiency of the hardware computing module.
[0106] (2) A lightweight and efficient encoding-decoding module is used to compress the pruned weights into a more compact data dimension (90% of the original size) for storage. When calculation is required, the decoding module is used to restore the data dimension, thereby saving data transmission and converting expensive data movement into low-cost data calculation.
[0107] The key technical points of this embodiment have the following beneficial effects:
[0108] 1. The pruning technology of this embodiment is particularly suitable for the implementation of hardware accelerators. Traditional unstructured pruning algorithms often lead to irregular data access and unbalanced workload of processing units in hardware. The pruning algorithm of this embodiment greatly simplifies computing scheduling and data allocation by ensuring the consistency of the number of non-zero elements in each column, reduces cache failure rate and memory access latency, and improves the consistency of the computing process. At the same time, since the distribution of non-zero elements is more concentrated and regular, hardware accelerators can achieve parallel computing with less control overhead. This method reduces storage requirements while alleviating storage bandwidth pressure, so that the overall system achieves higher performance while saving energy, bringing significant hardware efficiency improvements to neural network reasoning in high-performance computing.
[0109] 2. The encoding-decoding module of this embodiment can achieve the exchange of expensive data movement for low-cost calculation, because it compresses the pruned weight matrix into a compact representation to reduce data movement, at the cost of slightly increased computational effort in order to restore them back for matrix calculation. During the encoding process, this embodiment compresses the data to 10% of the original matrix, so that the data can be stored in an off-chip memory, thereby saving valuable chip internal storage space when storage resources are limited. In addition, during the decoding process, since this embodiment uses a globally shared decoding module, it is only necessary to load the weight matrix of the decoding module from the external memory to the internal storage in advance, so that data reuse can be achieved during calculation, and frequent data access is not required, thereby significantly reducing the energy consumption generated by data reading.
[0110] Reference Figure 7 The embodiment of the present application further provides a neural network sparsification device, which can implement the above-mentioned neural network sparsification method, and the device includes:
[0111] A weight pruning unit is used to perform structured pruning on the weight parameters of each layer in the weight parameter matrix of the neural network, so as to delete a number of weight parameters with the lowest absolute value in each column of the weight parameter matrix, and obtain an original pruned weight matrix; wherein each column of the original pruned weight matrix includes a number of non-zero weight parameters of the same number;
[0112] A weight compression unit, configured to configure a corresponding row index for each of the non-zero weight parameters in the original pruning weight matrix and then compress the original pruning weight matrix using an encoder composed of a linear layer;
[0113] A weight storage unit, used to store the original pruned weight matrix compressed by the encoder into an external memory of an accelerator; wherein the accelerator is used to run the neural network;
[0114] A weight reconstruction unit, configured to read the compressed original pruning weight matrix from the external memory to the internal memory of the accelerator, and then decode the compressed original pruning weight matrix using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix;
[0115] A neural network running unit is used to use the accelerator to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix.
[0116] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0117] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores computer programs and data, and the processor implements the above-mentioned neural network sparsification method when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0118] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0119] See also Figure 8 , Figure 8 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0120] The processor 801 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0121] The memory 802 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 802, and the processor 801 calls and executes the neural network sparsification method of the embodiment of this application;
[0122] Input / output interface 803, used to implement information input and output;
[0123] The communication interface 804 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0124] A bus 805 that transmits information between the various components of the device (e.g., the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);
[0125] The processor 801 , the memory 802 , the input / output interface 803 and the communication interface 804 are connected to each other in communication within the device via a bus 805 .
[0126] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned neural network sparsification method is implemented.
[0127] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0128] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0129] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0130] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0131] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0132] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0133] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0134] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0135] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0136] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0137] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0138] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0139] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A neural network sparsification method, characterized in that: The method comprises the following steps: Performing structured pruning on the weight parameters of each layer in the weight parameter matrix of the neural network to delete a number of weight parameters ranked low in absolute value in each column of the weight parameter matrix to obtain an original pruned weight matrix; wherein each column of the original pruned weight matrix includes a number of non-zero weight parameters of the same number; Configuring corresponding row indices for each of the non-zero weight parameters in the original pruning weight matrix and then compressing the original pruning weight matrix using an encoder composed of linear layers; Storing the original pruning weight matrix compressed by the encoder in an external memory of an accelerator; wherein the accelerator is used to run the neural network; Reading the compressed original pruning weight matrix from the external memory to the internal memory of the accelerator, and then decoding the compressed original pruning weight matrix using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix; The accelerator is used to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix.
2. The neural network sparsification method according to claim 1, characterized in that: Before performing structured pruning on the weight parameters of each layer in the weight parameter matrix of the neural network to delete a number of weight parameters with the lowest absolute value in each column of the weight parameter matrix to obtain the original pruned weight matrix, the method further comprises the following steps: Training the neural network until the neural network converges; The weight parameters when the neural network converges are set as the weight parameter matrix.
3. The neural network sparsification method according to claim 1, characterized in that: The structured pruning of the weight parameters of each layer in the weight parameter matrix of the neural network to delete several weight parameters with the lowest absolute value in each column of the weight parameter matrix to obtain the original pruned weight matrix includes the following steps: Sort the weight parameters in each column of the weight parameter matrix of the neural network by absolute value; Determining a binary weight mask matrix; wherein each parameter in the binary weight mask matrix is used to determine whether a corresponding weight parameter in the weight parameter matrix is retained or deleted; According to the binary weight mask matrix, a plurality of weight parameters whose absolute values are ranked at the bottom in each column of the weight parameter matrix are deleted to retain a plurality of non-zero weight parameters whose absolute values are ranked at the top in each column of the weight parameter matrix and whose number is the same, to obtain a candidate pruning weight matrix; Training the neural network provided with the candidate pruning weight matrix until the neural network converges; Return to the step of sorting the weight parameters in each column of the weight parameter matrix of the neural network according to the absolute value, until the neural network is trained a set number of times, and use the current candidate pruning weight matrix as the original pruning weight matrix.
4. The neural network sparsification method according to claim 1, characterized in that: The configuring a corresponding row index for each of the non-zero weight parameters in the original pruning weight matrix comprises the following steps: A row index in the COO format is configured for each of the non-zero weight parameters in the original pruning weight matrix.
5. The neural network sparsification method according to claim 1, characterized in that: The method of decoding the compressed original pruning weight matrix by using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix comprises the following steps: Decoding the compressed original pruning weight matrix using the decoder composed of linear layers; A dot product operation is performed on the non-zero weight parameters in each column of the decoded original pruning weight matrix and the corresponding row index with the corresponding input activation value to obtain the reconstructed pruning weight matrix.
6. The neural network sparsification method according to claim 1, characterized in that: The method of decoding the compressed original pruning weight matrix by using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix comprises the following steps: The decoder composed of linear layers and reconstruction loss are used to decode the compressed original pruning weight matrix to reconstruct the reconstructed pruning weight matrix.
7. The neural network sparsification method according to any one of claims 1 to 6, characterized in that: The accelerator is disposed in at least one of a gateway, a camera or a drone; The step of using the accelerator to read the reconstructed pruning weight matrix from the internal memory and then running the neural network provided with the reconstructed pruning weight matrix comprises at least one of the following steps: Using the accelerator disposed in the gateway to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix to process and transfer information sent by a terminal; wherein the terminal is a device that establishes a communication connection with the gateway; Alternatively, the accelerator disposed in the camera is used to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix to process the image captured by the camera; Alternatively, the accelerator disposed in the drone is used to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix to process the sensor information acquired by the drone.
8. A neural network sparsification device, characterized in that: The device comprises: A weight pruning unit is used to perform structured pruning on the weight parameters of each layer in the weight parameter matrix of the neural network, so as to delete a number of weight parameters with the lowest absolute value in each column of the weight parameter matrix, and obtain an original pruned weight matrix; wherein each column of the original pruned weight matrix includes a number of non-zero weight parameters of the same number; A weight compression unit, configured to configure a corresponding row index for each of the non-zero weight parameters in the original pruning weight matrix and then compress the original pruning weight matrix using an encoder composed of a linear layer; A weight storage unit, used to store the original pruned weight matrix compressed by the encoder into an external memory of an accelerator; wherein the accelerator is used to run the neural network; A weight reconstruction unit, configured to read the compressed original pruning weight matrix from the external memory to the internal memory of the accelerator, and then decode the compressed original pruning weight matrix using a decoder composed of linear layers to reconstruct a reconstructed pruning weight matrix; A neural network running unit is used to use the accelerator to read the reconstructed pruning weight matrix from the internal memory and then run the neural network provided with the reconstructed pruning weight matrix.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program and data, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.