Large language model (LLM) pruning using extended Kronecker approximation

By estimating the local curvature and dynamic weight updates of the neural network, and using the Kronecker approximation to prune large language models, the deployment challenge of the converter architecture under resource-constrained conditions is solved, achieving efficient model compression and performance preservation.

CN121889808APending Publication Date: 2026-04-17QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2024-07-26
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing large language model (LLM) transformer architectures are difficult to deploy under computational, environmental, or device-specific constraints due to their massive size, especially under conditions of limited memory and computing resources.

Method used

By estimating the local curvature of the lost landscape in the neural network, the parameters to be removed from the neural network are dynamically assigned and the remaining weights are updated. Unstructured, semi-structured, and structured compressions are performed using local curvature and weight correlations, and the Kronecker approximation is used to optimize the pruning process.

Benefits of technology

This approach achieves the compression of large language models into smaller models while maintaining computational efficiency, meeting deployment constraints, and maintaining high performance, thus avoiding potential performance losses in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121889808A_ABST
    Figure CN121889808A_ABST
Patent Text Reader

Abstract

An apparatus has one or more memories and one or more processors coupled to the memories. The processor is configured to estimate a local curvature of a lost landscape of the neural network. The processor is further configured to dynamically allocate parameters to be removed from the neural network based on the local curvature. The processor is further configured to update a remaining weight of the neural network based on the parameter to be removed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Patent Application No. 18 / 476,729, filed September 28, 2023, entitled “LARGE LANGUAGE MODEL (LLM) PRUNING USING EXTENDED KRONECKER APPROXIMATIONS”, the entire disclosure of which is expressly incorporated herein by reference. Technical Field

[0003] All aspects of this disclosure relate to pruning of large language models (LLMs) using the extended Kronecker approximation. Background Technology

[0004] Artificial neural networks can comprise interconnected groups of artificial neurons (e.g., neuron models). An artificial neural network (ANN) can be a computing device or represented as a method to be performed by a computing device. A convolutional neural network (CNN) is a type of feedforward ANN. A CNN can comprise an ensemble of neurons, each with a receptive field and collectively constructing the input space. CNNs such as deep convolutional neural networks (DCNs) have numerous applications. Specifically, these neural network architectures are used in a variety of technologies, such as image recognition, speech recognition, acoustic scene classification, keyword retrieval, autonomous driving, and other classification tasks. State-of-the-art language models are becoming increasingly large in an effort to achieve peak performance on large corpora of available text data. However, the sheer size of the transformer architectures used makes deploying models within computationally, environmentally, or device-specific constraints increasingly difficult. Summary of the Invention

[0005] Various aspects of this disclosure relate to an apparatus. The apparatus has one or more memories and one or more processors coupled to the one or more memories. The processors are configured to estimate the local curvature of a loss landscape of a neural network. The processors are also configured to dynamically assign parameters to be removed from the neural network based on the local curvature. The processors are further configured to update the remaining weights of the neural network based on the parameters to be removed.

[0006] In other aspects of this disclosure, a processor-implemented method includes estimating the local curvature of a loss landscape of a neural network. The method also includes dynamically assigning parameters to be removed from the neural network based on the local curvature. Furthermore, the method includes updating the remaining weights of the neural network based on the parameters to be removed.

[0007] In other aspects of this disclosure, a non-transitory computer-readable medium having program code recorded thereon is disclosed. The program code is executed by a processor and includes program code for estimating the local curvature of the loss landscape of a neural network. The program code also includes program code for dynamically assigning parameters to be removed from the neural network based on the local curvature. The program code further includes program code for updating the remaining weights of the neural network based on the parameters to be removed.

[0008] Other aspects of this disclosure relate to an apparatus. The apparatus includes components for estimating the local curvature of a loss landscape of a neural network. The apparatus also includes components for dynamically assigning parameters to be removed from the neural network based on the local curvature. Furthermore, the apparatus includes components for updating the remaining weights of the neural network based on the parameters to be removed.

[0009] Additional features and advantages of this disclosure will be described below. Those skilled in the art will understand that this disclosure can be readily used as the basis for modifying or designing other structures for implementing the same purposes as this disclosure. Those skilled in the art will also recognize that such equivalent constructions do not depart from the teachings of this disclosure as set forth in the appended claims. Novel features considered characteristic of this disclosure, in both their organization and manner of operation, along with further objects and advantages, will be better understood when considered in conjunction with the accompanying drawings. However, it is to be clearly understood that each drawing is provided for illustrative and descriptive purposes only and is not intended to be a definition of a limitation of this disclosure. Attached Figure Description

[0010] The features, substance, and advantages of this disclosure will become more apparent when understood in conjunction with the accompanying drawings, in which the same reference numerals are consistently used for identification.

[0011] Figure 1 Example implementations of neural networks using a system-on-a-chip (SoC) (including a general-purpose processor) according to certain aspects of this disclosure are illustrated.

[0012] Figure 2A , Figure 2B and Figure 2C This is a diagram illustrating various aspects of a neural network according to this disclosure.

[0013] Figure 2D This is a diagram illustrating an exemplary deep convolutional network (DCN) according to various aspects of this disclosure.

[0014] Figure 3 This is a block diagram illustrating an exemplary deep convolutional network (DCN) according to various aspects of this disclosure.

[0015] Figure 4This is a block diagram illustrating an exemplary software architecture that enables modularization of artificial intelligence (AI) functions according to various aspects of this disclosure.

[0016] Figure 5 This is an illustration of data-based compression according to various aspects of this disclosure.

[0017] Figure 6 These are illustrations of structured, semi-structured, and unstructured pruning according to various aspects of this disclosure.

[0018] Figure 7 This is a flowchart illustrating examples of processes for operating a neural network according to various aspects of this disclosure. Detailed Implementation

[0019] The detailed description that follows, taken in conjunction with the accompanying drawings, is intended as a description of various configurations and not as representing only configurations in which the described concepts can be practiced. To provide a comprehensive understanding of the various concepts, the detailed description includes specific details. However, it will be apparent to those skilled in the art that these concepts can be practiced without these specific details. In some instances, to avoid obscuring such concepts, well-known structures and components are shown in block diagram form.

[0020] Based on the teachings, those skilled in the art will recognize that the scope of this disclosure is intended to cover any aspect of this disclosure, whether implemented independently of or in combination with any other aspect of this disclosure. For example, an apparatus or method may be implemented using any number of the aspects described. Furthermore, the scope of this disclosure is intended to cover such apparatuses or methods practiced using other structures, functionalities, or structures and functionalities that complement or differ from the various aspects of this disclosure described. It should be understood that any aspect of this disclosure may be embodied by one or more elements of the claims.

[0021] The word “exemplary” is used to mean “serving as an example, instance, or illustration.” Any aspect described as “exemplary” need not be interpreted as superior to or better than other aspects.

[0022] While specific aspects are described, numerous variations and substitutions of these aspects fall within the scope of this disclosure. Although some benefits and advantages of preferred aspects are mentioned, the scope of this disclosure is not intended to be limited to a particular benefit, use, or purpose. Rather, aspects of this disclosure are intended to be widely applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the accompanying drawings and the following description of preferred aspects. The detailed description and drawings are merely illustrative and not limiting of this disclosure, the scope of which is defined by the appended claims and their equivalents.

[0023] This disclosure introduces large language model (LLM) pruning techniques (referred to as surgical techniques) for unstructured, semi-structured, and structured compression of neural networks (e.g., LLMs). While large language models are described, the techniques of this disclosure are also applicable to multilayer perceptrons (MLPs) and neural networks. These techniques aim to find optimal pruning by unfolding the curvature of the model's loss landscape. These techniques leverage modern Fisher approximations to extend accurate pruning to the domain of large language models with millions or billions of parameters, while remaining practical for memory and computational resources. These techniques use weight strengths and activations from the forward pass, as well as gradient information from the backward pass, to correlate the expected cost of weight removal with the global final goal. Typically, model compression occurs only once in practice, and can then be deployed multiple times with the achieved compressed performance. This inspired this approach, which, compared to baseline techniques, requires a longer compression time but achieves the most favorable performance / compression balance.

[0024] Figure 1 An example implementation of a system-on-a-chip (SOC) 100 is illustrated, which may include a central processing unit (CPU) 102 or a multi-core CPU configured for pruning artificial neural networks. Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., a weighted neural network), latency, frequency interval information, and task information may be stored in a memory block associated with a neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from the program memory associated with the CPU 102 or from memory block 118.

[0025] The SOC 100 may also include additional processing blocks tailored for specific functions, such as a GPU 104, a DSP 106, a connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation LTE (4G) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 capable of, for example, detecting and recognizing gestures. In one specific implementation, an NPU 108 is implemented within a CPU 102, a DSP 106, and / or a GPU 104. The SOC 100 may also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120, which may include a global positioning system.

[0026] The SOC 100 may be based on the ARM instruction set. In various aspects of this disclosure, instructions loaded into the general-purpose processor 102 may include code for estimating the local curvature of the loss landscape of the neural network. The general-purpose processor 102 may also include: code for dynamically allocating parameters to be removed from the neural network based on the local curvature; and code for updating the remaining weights of the neural network based on the parameters to be removed.

[0027] Deep learning architectures perform object recognition tasks by learning to represent inputs at progressively higher levels of abstraction in each layer, thereby constructing useful feature representations of the input data. In this way, deep learning addresses a major bottleneck of traditional machine learning. Before deep learning, machine learning methods for object recognition problems often relied heavily on human-designed features, possibly in conjunction with shallow classifiers. Shallow classifiers could be two-class linear classifiers, where a weighted sum of feature vector components is compared to a threshold to predict which class the input belongs to. Human-designed features could be templates or kernels customized for a specific problem domain by engineers with domain expertise. In contrast, while deep learning architectures can learn to represent features similar to those that human engineers might design, this requires training. Furthermore, deep networks can learn to represent and recognize novel types of features that humans might not have considered.

[0028] Deep learning architectures can learn hierarchical structures of features. For example, if presented with visual data, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to recognize spectral power at specific frequencies. The second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.

[0029] Deep learning architectures perform particularly well when applied to problems with a natural hierarchical structure. For example, the classification of motorized vehicles can benefit from first learning to identify features such as wheels, windshields, and others. These features can then be combined in different ways at higher levels to identify cars, trucks, and airplanes.

[0030] Neural networks can be designed to have multiple connectivity patterns. In feedforward networks, information is passed from lower layers to higher layers, where each neuron in a given layer communicates with neurons in higher layers. As described above, hierarchical representations can be built in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be passed to another neuron in the same layer. Recurrent architectures can help identify patterns across more than one block of input data that is sequentially delivered to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when the recognition of higher-level concepts can aid in discerning specific lower-level features of the input.

[0031] The connections between layers in a neural network can be fully connected or locally connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In the fully connected neural network 202, neurons in the first layer can transmit their outputs to each neuron in the second layer, so that each neuron in the second layer will receive inputs from each neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is illustrated. In the locally connected neural network 204, neurons in the first layer can connect to a finite number of neurons in the second layer. More generally, the locally connected layers of the locally connected neural network 204 can be configured such that each neuron in the layer will have the same or similar connectivity pattern, but the connection strength can have different values ​​(e.g., 210, 212, 214, and 216). The connectivity pattern of locally connected networks can produce spatially different receptive fields in higher layers because neurons in higher layers in a given region can receive inputs that are tuned to the characteristics of a restricted portion of the total input to the network through training.

[0032] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. The convolutional neural network 206 can be configured such that the connection strength associated with the input for each neuron in the second layer is shared (e.g., 208). Convolutional neural networks may be well-suited for problems where the spatial location of the input is meaningful.

[0033] One type of convolutional neural network is the deep convolutional network (DCN). Figure 2DA detailed example of a DCN 200 designed to recognize visual features from an image 226 input from an image capture device 230 (such as an in-vehicle camera) is illustrated. The DCN 200 in this example can be trained to identify traffic signs and the numbers provided on them. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or traffic lights.

[0034] Supervised learning can be used to train the DCN 200. During training, an image (such as image 226 of a speed limit sign) can be presented to the DCN 200, and forward passes can then be computed to produce output 222. The DCN 200 may include a feature extraction part and a classification part. Upon receiving image 226, convolutional layer 232 may apply a convolutional kernel (not shown) to image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel used for convolutional layer 232 may be a 5x5 kernel that generates 28x28 feature maps. In this example, because four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to image 226 at convolutional layer 232. Convolutional kernels may also be referred to as filters or convolutional filters.

[0035] The first set of feature maps 218 can be subsampled by a max-pooling layer (not shown) to generate a second set of feature maps 220. The max-pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (e.g., 14x14) is smaller than the size of the first set of feature maps 218 (e.g., 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).

[0036] exist Figure 2D In the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. Furthermore, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 may include a number corresponding to a possible feature of image 226, such as "sign", "60", and "100". A softmax function (not shown) can convert the numbers in the second feature vector 228 into probabilities. Therefore, the output 222 of DCN 200 can be the probability that image 226 includes one or more features.

[0037] In this example, the probabilities for "sign" and "60" in output 222 are higher than the probabilities for other numbers in output 222 (such as "30", "40", "50", "70", "80", "90", and "100"). Before training, output 222 generated by DCN 200 may be incorrect. Therefore, the error between output 222 and the target output can be calculated. The target output is the ground truth value of image 226 (e.g., "sign" and "60"). The weights of DCN 200 can then be adjusted so that output 222 of DCN 200 is more closely aligned with the target output.

[0038] To adjust the weights, the learning algorithm computes the gradient vector of the weights. The gradient indicates by how much the error will increase or decrease as the weights are adjusted. At the top layers, the gradient corresponds directly to the values ​​of the weights connecting the activated neurons in the penultimate layer to the neurons in the output layer. In lower layers, the gradient depends on the values ​​of the weights and the error gradient computed in the higher layers. The weights can then be adjusted to reduce the error. This method of adjusting weights is called "backpropagation" because it involves "passing backward" through the neural network.

[0039] In practice, the error gradient of the weights can be calculated using a small number of examples to make the calculated gradient approximate the true error gradient. This approximation method is called stochastic gradient descent. Stochastic gradient descent can be repeated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level. After learning, a new image (e.g., a speed limit sign in image 226) can be presented to DCN 200, and output 222 can be generated through the forward pass of DCN 200, which can be considered as the inference or prediction of DCN 200.

[0040] Deep Belief Networks (DBNs) are probabilistic models that include multiple layers of hidden nodes. DBNs can be used to extract hierarchical representations of training datasets. DBNs are obtained by stacking layers of Restricted Boltzmann Machines (RBMs). An RBM is a type of artificial neural network that learns a probability distribution from a set of inputs. Because RBMs can learn a probability distribution without information about the class each input should be classified into, they are often used for unsupervised learning. Using a hybrid paradigm of supervised and unsupervised learning, the bottom RBM of a DBN can be trained unsupervised and used as a feature extractor, while the top RBM can be trained supervisedly (on the joint distribution of inputs from the previous layer and the target class) and used as a classifier.

[0041] DCN is a network of convolutional networks configured with additional pooling and normalization layers. DCN has achieved state-of-the-art performance on many tasks. DCN can be trained using supervised learning, where both the input and output targets are known for many paradigms and are used to modify the network's weights using gradient descent.

[0042] DCNs can be feedforward networks. Furthermore, as described above, connections from neurons in the first layer of a DCN to a set of neurons in the next higher layer are shared across neurons in the first layer. The feedforward and shared connections of a DCN can be used for fast processing. For example, the computational cost of a DCN may be much smaller than that of a similarly sized neural network that includes recurrent or feedback connections.

[0043] The processing at each layer of a convolutional network can be thought of as a spatially invariant template or base projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then a convolutional network trained on that input can be thought of as three-dimensional, with two spatial dimensions along the axes of the image and a third dimension capturing color information. The output of the convolutional connections can be thought of as forming a feature map in the next layer, where each element in the feature map (e.g., 220) receives input from a range of neurons in the previous layer (e.g., feature map 218) and from each of the multiple channels. The values ​​in the feature map can be further processed with nonlinearities (such as correction, max(0,x)). Values ​​from neighboring neurons can be further pooled, which corresponds to downsampling and provides additional local invariance and dimensionality reduction. Normalization corresponding to whitening can also be applied through lateral inhibition between neurons in the feature map.

[0044] Figure 3 This is a block diagram illustrating a DCN 350. A DCN 350 can include multiple layers of different types based on connectivity and weight sharing. For example... Figure 3 As shown, DCN 350 includes convolutional blocks 354A and 354B. Each convolutional block in convolutional blocks 354A and 354B can be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max pooling layer (MAX POOL) 360.

[0045] Although only two of the convolutional blocks 354A and 354B are shown, this disclosure is not limited thereto, and any number of convolutional blocks 354A and 354B may be included in the DCN 350 according to design preferences.

[0046] Convolutional layer 356 may include one or more convolutional filters that can be applied to the input data to generate feature maps. Normalization layer 358 can normalize the output of the convolutional filters. For example, normalization layer 358 can provide whitening or lateral suppression. Max pooling layer 360 can provide spatial downsampling aggregation to achieve local invariance and dimensionality reduction.

[0047] For example, the parallel filter bank of a deep convolutional network can be loaded onto a SOC 100 (e.g., Figure 1 The CPU 102 or GPU 104 of the SOC 100 can be used to achieve high performance and low power consumption. In an alternative implementation, a parallel filter bank can be loaded onto the DSP 106 or ISP 116 of the SOC 100. Additionally, the DCN 350 can access other processing blocks that may exist on the SOC 100, such as the sensor processor 114 and navigation module 120, which are dedicated to sensors and navigation, respectively.

[0048] The DCN 350 may also include one or more fully connected layers 362 (FC1 and FC2). The DCN 350 may also include logistic regression (LR) layers 364. Weights (not shown) to be updated are located between each of the layers 356, 358, 360, 362, and 364 of the DCN 350. The output of each layer (e.g., 356, 358, 360, 362, and 364) can be used as input to the next layer in the DCN 350 to learn hierarchical feature representations from the input data 352 (e.g., images, audio, video, sensor data, and / or other input data) provided at the first convolutional block in convolutional block 354A. The output of the DCN 350 is a classification score 366 of the input data 352. The classification score 366 may be a set of probabilities, where each probability is the probability that the input data includes a feature from the feature set.

[0049] Figure 4 This is a block diagram illustrating an exemplary software architecture 400 with modular artificial intelligence (AI) functionality. According to various aspects of this disclosure, using architecture 400, a system-on-a-chip (SoC) 420 (which may be similar to...) can be designed... Figure 1 Various processing blocks of the SoC 400 (e.g., CPU 422, DSP 424, GPU 426, and / or NPU 428) support pruned neural networks for AI applications 402. Architecture 400 can be included, for example, in a computing device such as a smartphone.

[0050] AI application 402 can be configured to invoke functions defined in user space 404, which may, for example, provide the detection and recognition of a scene indicating the current location of the computing device (including architecture 400). For example, AI application 402 may configure microphones and cameras differently depending on whether the recognized scene is an office, lecture hall, restaurant, or an outdoor environment such as a lake. AI application 402 may make requests to compiled program code associated with libraries defined in the AI ​​Function Application Programming Interface (API) 406. This request may ultimately rely on the output of a deep neural network configured to provide inferred responses based on, for example, video and location data.

[0051] Runtime engine 408 (which may be compiled code of a runtime framework) may be further accessible to AI application 402. AI application 402 may cause runtime engine 408 to request inference, for example, at specific time intervals or when triggered by events detected by the user interface of AI application 402. When runtime engine 408 provides an inference response, it may then signal to the operating system (OS) space 410 running on SOC 420, such as kernel 412. In some examples, kernel 412 may be a LINUX kernel. The operating system may then enable sequential relaxation of quantization to be performed on CPU 422, DSP 424, GPU 426, NPU 428, or some combination thereof. CPU 422 may be directly accessible by the operating system, while other processing blocks may be accessed via drivers such as drivers 414, 416, or 418 for DSP 424, GPU 426, or NPU 428, respectively. In an exemplary example, the deep neural network may be configured to run on a combination of processing blocks such as CPU 422, DSP 424 and GPU 426, or on NPU 428.

[0052] State-of-the-art language models are becoming increasingly large in an effort to achieve peak performance on large corpora of available text data. However, the sheer size of transformer architectures makes deploying models within computationally, environmentally, or device-specific constraints increasingly difficult. Aspects of this disclosure introduce data-driven compression of existing pre-trained models as an alternative to training small models from scratch. To this end, a scalable Kronecker factor approximation of the local loss landscape curvature is improved, allowing for dynamic allocation of removable structures and updating of remaining weights. A general framework for unstructured, semi-structured, and structured compression is provided to improve existing estimates of the loss landscape to better capture correlations between weights while maintaining computational efficiency.

[0053] Recent advances in language modeling have enabled the fitting of large language models (LLMs) with millions or billions of parameters to large text corpora, resulting in high performance. Unfortunately, the size of these LLMs can make them difficult to deploy within practical constraints, such as limited memory for on-device inference or reduced computational or memory requirements.

[0054] The aim is to compress existing large language models into smaller models based on data to meet deployment constraints, as an alternative to training new models from scratch. Some benefits of this learning paradigm are: (i) leveraging the performance of existing pre-trained LLMs without requiring large datasets or expensive training, while (ii) obtaining models that accurately meet deployment requirements (including environmental or hardware constraints).

[0055] Compared to conventional LLM pruning, aspects of this disclosure use a more accurate approximation of the loss landscape curvature and consider more weight correlations when updating the remaining weights. Unlike most existing data-based compression methods in LLM, aspects of this disclosure use weight strengths and activations from the forward pass, as well as gradient information from the backward pass, to correlate the expected cost of weight removal with the final global objective, allowing for global thresholding. Furthermore, multiple processes and first-order weight corrections are introduced to further improve compression performance.

[0056] Figure 5 These are illustrations illustrating various aspects of data compression according to this disclosure. Figure 5 In the leftmost part 502, model 504 is trained from scratch on a small amount of target data 506, resulting in low performance. (As in...) Figure 5 As seen in the central part 510, the target data 506 can be used to train a large language model 512. Data-based surgery (also known as neural network pruning) on ​​the pre-trained large language model produces a compressed model 514. (As in...) Figure 5 As seen in the rightmost part 520, the compressed model 514 can be deployed on the mobile device 522.

[0057] Neural network pruning aims to remove parameters from the model while minimizing the negative impact on final performance. More formally, we will refer to LLM... The parameters are represented as vectors This is achieved by flattening attention and fully connected blocks. The weight matrix is ​​obtained from the weight matrix, where The data has been fitted. To minimize negative likelihood loss The pruned vector can be calculated. In the form of a compression model: , satisfying based on The pruning constraints are: (1) The constraints chosen may affect (e.g., determine) the compressed weights (and the pruned vectors). The structure.

[0058] Figure 6 This is an illustration of structured, semi-structured, and unstructured pruning according to various aspects of this disclosure. The original weight matrix W (shown in the leftmost image 602) can be pruned in various ways. In structured pruning 604, entire rows and columns are set to zero. In M:N semi-structured pruning 606, M weights out of every N consecutive weights are set to zero. Figure 6 In the example, one of every two weights is set to zero. In unstructured pruning 608, the score of the total weight elements is set to zero. Structured pruning results in the most direct gains in terms of memory and computation because it directly reduces the dimension of the matrix that needs to be explicitly represented, but is generally considered a more difficult compression task. In other schemes, maintaining high performance is often easier, but requires specialized algorithms that leverage sparse structures to benefit from deployment. The LLM operation considers all the pruning types discussed above, with a focus on structured pruning in LLM.

[0059] Equation 1 typically cannot be solved directly because the space of possible pruning configurations exceeds the space that can be evaluated in practice. For example, searching for all possible unstructured pruning masks for an LLM containing 125 million parameters would require... An evaluation. Therefore, the idea is to introduce a more tractable proxy model for the loss landscape. : (2) Using proxy model To find the pruned solution to equation 1 If it is an agent By choosing a specific Gaussian form, the solutions to unstructured, semi-structured, and structured pruning constraints can be derived in closed-form, as detailed in Appendix A below.

[0060] Now describe a good proxy for Taylor expansion to obtain losses. One approach is to work around pre-trained weights. Using the second-order Taylor expansion to locally expand the logarithmic loss, we obtain: (3) in It represents Jacobian, and This represents the Hessian. The first-order term vanishes when it reaches its optimal value. This is an oversimplification because (i) the neural network may not be optimized to its minimum, (ii) we may use a different loss than the one used to train the network for compression, and (iii) we perform pruning multiple times, inevitably causing the weights to deviate from their optimal values. However, we initially follow this simplification assumption and consider staggered first-order corrections to mitigate this problem. The quadratic expansion of Equation 3 forms the basis of the optimal brain injury and optimal brain surgery pruning method, as described below in Appendix A. From a probabilistic perspective, the quadratic approximation of the log-likelihood implies a Gaussian approximation of the likelihood. This is the well-known Laplace approximation. ), where the pre-trained weights are used as the average value, and the local inverse Hessian is the covariance matrix that captures the correlation between the weights.

[0061] One way to approximate Hessian is through the Fisher information matrix: (4) It has the advantage of always being semidefinite, in which the inverse is thus formed. The appropriate covariance matrix, and can be used We approximate the Monte Carlo samples. For most LLMs, we simply treat the network's softmax output as a classification distribution. And sample from it. The complete Fisher In parameters The quantity expands quadratically. To overcome this, Fisher typically uses layer-by-layer blocks. This is represented, and approximated by considering only the block diagonal components of each layer: , (5) in This represents the Kronecker product. Because cross-layer interactions can be ignored, it can be used... replace Used with the weight matrix The associated Fisher block, the weight matrix from the input Generate output ,in Presentation layer, and This represents a data point. Therefore, the Fisher block can be activated using input. This indicates that these input activations are achieved through forward pass data. and the associated output gradient from backpropagation Obtained.

[0062] We now discuss pruning as a constrained optimization. This is based on the log-likelihood loss obtained through quadratic expansion. Gaussian approximation obtained Optimal update (Therefore also) The problem becomes a constrained quadratic optimization problem with the following equations: (6) in It is positive semidefinite, and It is pruned (e.g., set to zero). A set of indices.

[0063] we will Represented as a matrix, its row vectors are the canonical basis vectors of the elements to be pruned. One of the most standard methods for solving equation 6 is to use Lagrange multipliers, which yields a loss. Expected increase and optimal weight update The general closed-form solution: (7) (8) It can be used to derive all unstructured, semi-structured, and structured pruning by combining with the modern Fisher approximation. The same general formulas of equations 7 and 8 appear in existing work on LLM pruning, but like SparseGPT, it uses a simpler layer-by-layer Fisher approximation that ignores gradient information and does not consider structured pruning.

[0064]

[0065] LLM surgery

[0066] The components of this technique (called LLM surgery) are summarized in the pseudocode shown in Procedure 1. First, the calculation of the approximate curvature is described. Even considering only Fisher's equation 5... When approximating block by block, each block is still Big The sum of the matrices, which is too large to actually fit into memory, Representing rows and columns. Instead, we adapt the classical Kronecker solution (e.g., the Kronecker factor approximation curvature (KFAC)) that assumes independence of activation and derivative (IAD) to Fisher blocks. Approximately a single Kronecker product: ,in (9) Activation from forward pass and gradient from backpropagation structure.

[0067] Another advantage of approximating the Fisher block as the Kronecker product is that the inverse is particularly easy to compute. Only the inverse factor needs to be calculated. This fact allows us to avoid explicitly constructing the formula in memory. and The big one Instead of using multiple matrices, we directly use much smaller matrices. and .

[0068] Now we describe the computational cost (cost per row / column) in the final loss. The number of possible combinations of weights that can be removed grows exponentially with the parameter count, making it difficult to estimate the individual cost of each such removal. This is not feasible. Therefore, one strategy is to calculate the removal cost. Weights are processed independently. However, this does not necessarily mean updating the weights after selecting the weights to remove. We make the same strong independence assumption. Unlike most conventional systems, we present relevant weight updates by considering the off-diagonal elements of the Fisher approximation.

[0069] For semi-structured and unstructured pruning, various weight elements can be used. The independent cost, and for structured pruning, all rows can be used. and column The independent cost. The appropriate cost from the general cost formula equation 7 can be determined by setting... To derive, where the canonical basis vectors index The weight to be removed is selected for each individual one-hot element at a given location. For structured pruning, the row... and column It can be set or Choose from, among them Substituting into equation 7: , , (10) The complete derivation can be found in Appendix A below. Single element The cost is equivalent to those found in optimal brain surgery, and and This is very similar to structured brain surgery, but in our case, it's derived for matrix rows and columns (see Appendix A - Removing a Single Row or Column). Given a curvature estimate, the cost of elements or rows and columns can be computed in parallel. Furthermore, the Kronecker factor approximation is used. The more general cost of and can be derived in Appendix C through eigenvalue decomposition.

[0070] Dynamic weight assignment with a global threshold involves calculating the threshold and selecting the rows and columns to remove in process 1. Unlike existing work on layer-by-layer compression, we use a global threshold. This allows for dynamic allocation of sparsity across layers, pruning the most where the impact is minimal. The proposed compression method can compress the model to a specific target size. This is defined as the fraction of weights that should be retained, for example, as non-zero, after compression. In all structured, semi-structured, and unstructured pruning, we select as many weights as possible to remove, such that the target size... With the lowest possible cost To achieve this, the minimum possible cost is calculated based on the computational cost in the final loss, as described above. For unstructured pruning, this includes pruning all weights in the network. Sort by cost and set a global threshold This makes the weights The score falls on the threshold For M:N semi-structured pruning, the M costs of every N consecutive weights can be sorted, and the M weights with the lowest costs can be selected. In the case of multiple planning (see below), the cost of each block can be found by summing the costs of these M lowest costs in each block, sorting the costs of each block across the entire network, and setting a global threshold similar to the unstructured case. This makes the weights The score falls within the threshold. Finally, for structured pruning, sorting can be performed, which can be appropriately weighted by the number of elements constituting a row or column, and a global threshold can be set. This makes all weights The score falls within the threshold. Then, scores that fall within the threshold can be removed. , All rows and columns within.

[0071] The relevant weight update involves calculating the weight update in process 1. For the weight update... Unlike most existing work, we did not perform cost calculations. The same simplified independence assumption applies throughout. That is, when calculating the expected cost, we assume that rows, columns, or weights are independent to avoid the need to calculate the cost for every possible combination in which weights can be removed, which would quickly become impractical. However, once we have chosen the set of weights for pruning, we can typically compute a single relevant weight update associated with the joint removal of multiple weights, rather than simply summing the weight updates associated with each individual removal. We derive this type of relevant weight update.

[0072] For rapid updates of unstructured and semi-structured related weights, mathematically, we represent the pruned weights as follows: ,in This is the one-hot canonical basis vector from which the weights to be removed are selected. Because each element... With a unique associated row and column Indexes, so we can also index these corresponding rows. and column Using canonical basis vectors (for example, we have) For all (All are valid).

[0073] We consider the Fisher approximation Characteristic decomposition To derive the unstructured weight update (removing a single element) in Appendix A, it is derived from Equation 8: (11) For the sake of simplicity, bar symbols are used. as well as and Vectorized matrix diagonal.

[0074] In programming terms, we always avoid explicitly representing large matrices in memory. Instead, we calculate the relevant quantities based on their factors. Similarly, we never represent sparse matrices in memory. , or Instead, it directly uses a list of indices for one-hot elements. For example, and Vectors can be constructed by copying row vectors. It can be constructed by indexing all pruned weights.

[0075] For the maximum number of relevant weights, the main computational bottleneck is in Equation 11. Matrix inverse. To control the compression speed, we can divide the pruned weights into disjoint subsets. such that each subset No more than the maximum number set for the relevant weights And sum the associated independent updates. By setting a lower Using less correlation allows for a trade-off between compressed quality and speed.

[0076] For rapid, structured weight updates, it is necessary to find the inverse. Each relevant weight The general case of the matrix differs, and the Fisher approximation using the Kronecker factor is employed. The weight update only needs to be removed Inverse calculation for each row Matrix, or in removing Find the inverse when there are columns. The matrix. Updating uses fewer resources (e.g., cheaper) than expected based on the effective number of weights in those rows and columns, which would mean finding the inverse. This results in a matrix. In practice, this leads to a significant speedup in structured pruning and weight updates that consider the correlations between rows or columns. When removing... When traveling, , ,... or remove When listing, , ,... ,in and We will represent the one-hot vectors of all rows and columns to be removed as follows: ... and ... . and removal Row-related weight updates can be set To determine: Remove multiple OK: (12) Remove multiple List:

[0077] From here, remove individual rows under the Kronecker approximation. or a single column The special case involves inverting a 1x1 matrix, and therefore only scalar division is required: Remove a single row : or a single column : (13) This is based on the independent structured updates of the convolutional filters. Therefore, we have extended the existing structured weight updates to rows and columns, and derived update rules that also consider the correlations between structured groups (e.g., rows and columns).

[0078] Multiple pruning plans involve updating the remaining weights in process 1. Various aspects of this disclosure allow for compression to a specific target size. It is defined as the fraction of the weights remaining after pruning, for example, In the total model parameters Set to zero. Proxy loss landscape Relying on Taylor expansion (Equation 3), Taylor expansion holds only locally, and therefore for large jumps in the parameter space... It becomes unreliable. We have done so many times. (each with a small update) This problem can be mitigated by pruning and subsequently re-estimating the loss surface. Multiple pruning operations allow for more computation (to... To improve the final compression performance (linearly per unit). Since the model only needs to be compressed once and can then be deployed multiple times, the ability to spend additional computational resources to improve the final compression performance is a desirable property. Performance increases monotonically with increasing number of compression iterations, and higher sparsity typically requires more iterations.

[0079] We will now describe the optional steps in Process 1 in terms of the interleaved low-rank first-order correction. So far, we have assumed that the model is in an optimal state, finding a closed-form solution to the quadratic constraint problem. However, in practice, this assumption may not hold due to several reasons discussed above. To mitigate this problem, we consider adjusting the pruning order along with the weights... We use low-rank adaptive interleaving for first-order correction, which is commonly used in large-scale language modeling. We ingest the update after each pruning, which means the next loss curvature estimate... This will be closer to the optimum, and the assumptions behind the quadratic expansion are likely to hold more closely. Through low-rank adaptation (LoRA) of absorbing LLM updates between iterations, the sum of low-rank updates may have a higher rank than a single update. That is, rank... rank( For any updates Equality only occurs when updates lie entirely in the same subspace, which is unlikely to happen in practice. This insight can also be used during regular LoRA fine-tuning and therefore has applications beyond model compression to allow for more expressive low-rank model adaptation at negligible cost.

[0080] In this disclosure, we introduce an LLM surgery technique for unstructured, semi-structured, and structured compression of neural networks. This work aims to find optimal pruning by unfolding the curvature of the model's loss landscape. The method leverages a modern Fisher approximation to extend accurate pruning to the domain of large language models with millions or billions of parameters, while remaining practical in terms of memory and computational resources. Unlike most existing work on data-based LLM compression, we use not only weight strengths and activations from the forward pass but also gradient information from the backward pass to correlate the expected cost of weight removal with the global final goal. Compared to existing work, this method uses a more accurate approximation of the loss landscape curvature and considers more weight correlations when updating the remaining weights, while remaining efficient on modern hardware. Typically, model compression occurs only once in practice, and can then be deployed multiple times with the achieved compressed performance. This inspired our method, which, compared to baselines, requires a longer compression time but achieves the most favorable performance / compression balance.

[0081] Appendix A

[0082] Suppose we use log-likelihood We use a quadratic approximation to apply our loss. If we use the Gaussian approximation, then the optimal compression becomes the solution to the following constrained optimization problem: (14) in It was pruned A set of indices.

[0083] Now let's discuss the general solution.

[0084] In some examples, the pruned element can be represented as Furthermore, the cost is given by solving equation 6 using Lagrange multipliers. and weight update The general closed-form solution: (15) (16) Now let's discuss removing a single element.

[0085] Optimal Brain Surgery (OBS): To remove indexed brain tissue For a single element, we only need to set :

[0086] This corresponds perfectly to the loss and update of optimal brain surgery.

[0087] Optimal Brain Damage (OBD): We can also consider the elements to be independent, and the Fisher diagonal. Note that this means the diagonal element of the inverse Fisher is the Fisher diagonal. After inverting the scalar values ​​of the elements, the formula simplifies to: , (18) This corresponds perfectly to the loss and renewal of optimal brain injury.

[0088] Vectorization.

[0089] For specific implementation purposes, it may be convenient to use vectorized symbols. or Calculate all expected losses in parallel: For OBD: (19)

[0091] For OBS:

[0092] Remove a single row or column.

[0093] Structured OBS: If we consider OBS with known inverses approximation So, in order to remove the index For each row, we consider the correlation within the elements of that row. In other words, we write out the matrix. It contains lines The one-hot row vectors of all elements in the equation. Substituting into the general solution equation 7, we find: (20) Among them, we use express The first in Row vectors. Similarly, we obtain the associated weight updates: (twenty one) A similar structured pruning update for the convolutional filter is derived. By considering We can equivalently derive the expected loss and update for the column. If we do this, we will find the row. or list Structured updates: Remove rows :

[0094] Remove column : (twenty two)

[0095] Structured OBD: We can also assume that when rows are removed... At that time, the elements within a line are also independent, which means Similarly, when a column is removed... hour, Therefore, we can simplify it to: Remove rows :

[0096] Remove column : , (twenty three)

[0097] This is similar in form to the structured OBD loss and update of convolutional filters. Equation 23 involves removing a single row. The derivation is slightly different because we start from the general solution equation 8, avoiding the need to re-derive the Lagrange multiplier for each possible structure.

[0098] Prune multiple (related) rows and columns.

[0099] Let's consider removing OK row or The column, whose index is ,in and We represent the matrix containing the one-hot vectors of all rows and columns to be removed as follows: (twenty four) Then, a matrix containing the one-hot row vectors of all elements to be removed. It can be written as: Multiple lines: ,(in ) Multiple columns: ,(in (25) To remove both rows and columns simultaneously, we can stack matrices that have had duplicate row vectors removed: Multiple rows and multiple columns: (26) Due to the small number of rows and columns There are [number] overlapping elements, so duplicate rows need to be removed. After removal, the total number of rows is [number]. We use an identity matrix of appropriate size. and For simplicity, we will denote the vector or matrix of pruned weights as... .

[0100] First, we define the removal matrix as follows: To derive The removal of rows, and the definition The complete weight update for multi-line removal becomes: (27) Similarly, we define the removal matrix as follows: To derive The column was removed and defined. The complete weight update after removing multiple columns becomes: (28) The pseudocode for the other processes is shown below.

[0101]

[0102] In procedure 2, we describe how the LLM procedure works given a pruning target α and initial weights. Perform structured pruning under these conditions. Repeat the following steps. Next, calculate the approximate curvature based on the data. (and optionally also calculate) Then, use and (and optional use) Calculate the cost per row / column and Then, use and Calculate the global threshold This ensures that it satisfies the given target size for the current iteration. Then, based on τ, select which to remove. and The rows and columns. Next, based on (and optionally based on) To calculate weight updates Finally, update the remaining weights. + Optionally, a low-rank update can be performed at the end. Low-rank update ( Repeat these steps when the target compression ratio α is reached. Then return the compressed weights. .

[0103]

[0104] In procedure 3, we describe how the LLM procedure works given a pruning target α and initial weights. In this case, perform semi-structured and unstructured pruning. Repeat the following steps. Next, calculate the approximate curvature based on the data. (Optional calculation) Then, use and (and optional use) Calculate the cost of each element. After that, use Calculate the global threshold This ensures that it satisfies the given target size for the current iteration. Then, based on Select to remove The elements. Next, based on and (and optionally based on) ) Calculate weight update Finally, update the remaining weights. + Optionally, a low-rank update can be performed at the end. Low-rank update ( Repeat these steps when the target compression ratio α is reached. Then return the compressed weights. .

[0105] Appendix B - Damping

[0106] In practice, we add diagonal items. and Come to and The matrix is ​​used for damping. In some specific implementations, multiplying values ​​in the range [0.01, 0.1] by the average diagonal term often works well. To maintain consistency with existing work and allow for a fair comparison with the baseline, we can use... Furthermore, we can... Used for structured implementation, and will Used for semi-structured and unstructured implementations.

[0107] Appendix C - Extended Curvature Estimation .

[0108] We can improve the approximation by summing multiple Kronecker factors instead of using a single Kronecker product: (30) Appendix C addresses how to find such approximations computationally and how to utilize them within a neural network pruning framework.

[0109] The closest Kronecker product or the sum of Kronecker products.

[0110] We can find the closest Kronecker product approximation to Fisher in terms of the Frobenius norm. Instead of assuming the independence of activation and derivative according to the classic KFAC technique: (31) Finding the closest sum of Kronecker factors can be reformulated as the classical problem of finding the closest eigenvalue of a rank-1 matrix.

[0111] (32)

[0112] Power method and order reduction.

[0113] After considering the reshaping, we can use power iteration to solve equations 31 and 32 to find the initial G and A matrices and the closest Kronecker factor. .

[0114]

[0115] In process 4, we can see the exponentiation method. A more extensive description follows. At the beginning of process 4, we initialize the exponentiation iteration with a vector of all 1s, where 1 = [1 1 … 1]. After each iteration, we can initialize the vector with the final estimate obtained during the previous period.

[0116]

[0117] For the Kronecker factor approximation For larger , The larger it is, the closer it is to the real Fisher As measured by root mean square error (RMSE).

[0118] Extended curvature approximation.

[0119] For those with IAD or Form The classical KFAC, which is closest to the Kronecker approximation, is inversely simplified to Unfortunately, we cannot apply this inverse identity to the sum of Kronecker factors, which is why we rely on eigenvalue decomposition. and This allows us to decompose Fisher into: (33) because and It is orthogonal, and and Diagonal, reverse Fisher becomes: (34) This problem becomes even more difficult in the context of neural network training, because we want to understand the underlying mechanisms and their implications. Each sample Incremental tectonic estimation and This avoids the need to store more than one or a batch of input activations in memory simultaneously. Or output gradient Multiple The sum of the Kronecker factors will produce a closer approximation, but also... The increase in memory demand linearly increases with the increase in [something], and makes [something]... Finding the inverse becomes more difficult.

[0120] Formulas used to calculate costs and update weights.

[0121] For the sum of Kronecker factors, we find the cost in Equation 7. Weight updates in Equation 8 The constrained optimization solution becomes the following inner product and matrix-vector product: (35) (36) Its core is a matrix that captures the correlation between weights. : (37) in It is diagonal, and therefore the inverse can be calculated element-wise. For There are 1 relevant weights, and the size of the remaining inverse is 1. .

[0122] Explanation of the sum of Kronecker factors.

[0123] Appendix C is included to illustrate the applicability of the techniques disclosed herein to applications other than neural network pruning.

[0124] Figure 7 This is a flowchart illustrating an example of a process 700 for pruning a neural network according to various aspects of this disclosure. For example... Figure 7 As shown, in some aspects, process 700 may include estimating the local curvature of the loss landscape of the neural network (box 702). For example, the process may estimate the local curvature based on the weight values ​​or activation outer products obtained from the forward pass of the neural network and the gradient outer product from the backward pass of the neural network. The remaining weights may be updated based on the Kronecker factor approximation curvature (KFAC). The remaining weights may be updated by assuming the dependencies of the elements of the neural network. In some aspects, the process may update the remaining weights by computing a single relevant weight update associated with removing all assigned parameters together.

[0125] In some aspects, process 700 may include dynamically assigning parameters to be removed from the neural network based on local curvature (box 704). For example, the process may iteratively estimate local curvature and dynamically assign parameters to be removed.

[0126] In some aspects, process 700 may include updating the remaining weights of the neural network based on the parameters to be removed (box 706). For example, the process may iteratively and dynamically assign the parameters to be removed and utilize low-rank adaptation to update the remaining weights.

[0127] Example

[0128] Aspect 1: An apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: estimate the local curvature of a loss landscape of a neural network; dynamically allocate parameters to be removed from the neural network based on the local curvature; and update the remaining weights of the neural network based on the parameters to be removed.

[0129] Aspect 2: According to the apparatus of aspect 1, wherein the at least one processor is further configured to estimate the local curvature based on the weight values ​​or activation outer products obtained from the forward pass of the neural network and the gradient outer product from the backward pass of the neural network.

[0130] Aspect 3: The apparatus according to aspect 1 or 2, wherein the at least one processor is further configured to update the remaining weights based on the Kronecker factor approximation curvature (KFAC).

[0131] Aspect 4: The apparatus according to aspect 1, 2 or 3, wherein the at least one processor is further configured to update the remaining weights based on the assumed dependencies of the elements of the neural network.

[0132] Aspect 5: The apparatus according to aspects 1 to 4, wherein the at least one processor is further configured to update the remaining weights by computing a single relevant weight update associated with removing all assigned parameters together.

[0133] Aspect 6: The apparatus according to any one of the preceding aspects, wherein the at least one processor is further configured to iteratively estimate the local curvature and dynamically allocate the parameters to be removed.

[0134] Aspect 7: The apparatus according to any one of the preceding aspects, wherein the at least one processor is further configured to iteratively and dynamically allocate the parameters to be removed and to update the remaining weights using low-rank adaptation.

[0135] Aspect 8: A method for wireless communication, the method comprising: estimating the local curvature of a loss landscape of a neural network; dynamically allocating parameters to be removed from the neural network based on the local curvature; and updating the remaining weights of the neural network based on the parameters to be removed.

[0136] Aspect 9: According to the method of aspect 8, wherein the estimation of the local curvature is based on: the weight value or activation outer product obtained from the forward pass of the neural network, and the gradient outer product from the backward pass of the neural network.

[0137] Aspect 10: The method described in aspects 7 to 9, wherein the remaining weights are updated based on the Kronecker factor approximation curvature (KFAC).

[0138] Aspect 11: The method according to aspects 7 to 10, wherein the remaining weights are updated based on the assumption of the dependencies of the elements of the neural network.

[0139] Aspect 12: According to the method of aspects 7 to 11, updating the remaining weights further includes calculating a single relevant weight update associated with removing all assigned parameters together.

[0140] Aspect 13: The method according to any one of aspects 7 to 12, the method further comprising iteratively estimating the local curvature and dynamically assigning the parameter to be removed.

[0141] Aspect 14: The method according to any one of Aspects 7 to 13, the method further comprising iteratively and dynamically allocating the parameters to be removed and utilizing low-rank adaptation to update the remaining weights.

[0142] Aspect 15: A non-transitory computer-readable medium having program code recorded thereon, the program code being executed by a processor and comprising: program code for estimating local curvature of a loss landscape of a neural network; program code for dynamically allocating parameters to be removed from the neural network based on the local curvature; and program code for updating the remaining weights of the neural network based on the parameters to be removed.

[0143] Aspect 16: The non-transitory computer-readable medium according to aspect 15, wherein the program code for estimating local curvature is based on: weight values ​​or activation outer products obtained from the forward pass of the neural network, and gradient outer products from the backward pass of the neural network.

[0144] Aspect 17: A non-transitory computer-readable medium according to aspect 15 or 16, wherein the program code for updating the remaining weights is based on the Kronecker factor approximation curvature (KFAC).

[0145] Aspect 18: The non-transitory computer-readable medium according to aspects 15 to 17, wherein the program code for updating the remaining weights is based on the assumption of the dependency of the elements of the neural network.

[0146] Aspect 19: The non-transitory computer-readable medium according to aspects 15 to 18, wherein the program code for updating the remaining weights further includes program code for calculating a single relevant weight update associated with removing all assigned parameters together.

[0147] Aspect 20: A non-transitory computer-readable medium according to any one of aspects 15 to 19, wherein the program code further includes program code for iteratively estimating the local curvature and dynamically allocating the parameter to be removed.

[0148] Aspect 21: A non-transitory computer-readable medium according to any one of aspects 15 to 20, wherein the program code further includes program code for iteratively and dynamically allocating the parameters to be removed and updating the remaining weights using low-rank adaptation.

[0149] Aspect 22: An apparatus for wireless communication, the apparatus comprising: means for estimating local curvature of a loss landscape of a neural network; means for dynamically allocating parameters to be removed from the neural network based on the local curvature; and means for updating the remaining weights of the neural network based on the parameters to be removed.

[0150] Aspect 23: The apparatus for wireless communication according to aspect 22, wherein the component for estimating the local curvature is based on: weight values ​​or activation outer products obtained from the forward pass of the neural network, and gradient outer products from the backward pass of the neural network.

[0151] Aspect 24: The apparatus for wireless communication according to aspect 22 or 23, wherein the component for updating the remaining weights is based on the Kronecker factor approximation curvature (KFAC).

[0152] Aspect 25: The apparatus for wireless communication according to aspects 22 to 24, wherein the component for updating the remaining weights is based on the assumption of the dependency of the elements of the neural network.

[0153] Aspect 26: The apparatus for wireless communication according to aspects 22 to 25, wherein the component for updating the remaining weights further includes a component for calculating a single relevant weight update associated with removing all assigned parameters together.

[0154] Aspect 27: An apparatus for wireless communication according to any one of aspects 22 to 26, the apparatus further comprising means for iteratively estimating the local curvature and dynamically allocating the parameter to be removed.

[0155] Aspect 28: An apparatus for wireless communication according to any one of aspects 22 to 27, the apparatus further comprising means for iteratively and dynamically allocating the parameters to be removed and updating the remaining weights using low-rank adaptation.

[0156] The various operations of the methods described above can be performed by any suitable component capable of performing the corresponding function. These components may include various hardware and / or software components and / or modules, including but not limited to circuits, application-specific integrated circuits (ASICs), or processors. Generally, in the cases where operations are illustrated in the accompanying drawings, these operations may have corresponding paired components with similar numbering plus functional components.

[0157] As used, the term "determine" encompasses a wide variety of actions. For example, "determine" can include calculation, computation, processing, derivation, research, searching (e.g., looking in a table, database, or other data structure), assertion, etc. Additionally, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), etc. Furthermore, "determine" can include parsing, selecting, choosing, building, etc.

[0158] As used, the phrase "at least one of the items" in a list of items refers to any combination of these items, including a single member. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, ab, ac, bc, and abc.

[0159] The various exemplary logic blocks, modules, and circuits described in this disclosure can be implemented or executed using a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic component, discrete hardware component, or any combination thereof designed to perform the described functions. While the general-purpose processor may be a microprocessor, in alternative embodiments, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration.

[0160] The steps or algorithms of the methods described in this disclosure may be directly embodied in hardware, a software module executed by a processor, or a combination of both. The software module may reside in any form of storage medium known in the art. Some examples of usable storage media include random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, and the like. The software module may include a single instruction or multiple instructions and may be distributed across several different code segments, across different programs, and across multiple storage media. The storage medium may be coupled to the processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor.

[0161] The disclosed method includes one or more steps or actions for implementing the described method. The steps and / or actions of the method may be interchanged without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of a particular step and / or action may be modified without departing from the scope of the claims.

[0162] The described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system within the device. This processing system may utilize a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnect buses and bridges. The bus can link various circuits together, including processors, machine-readable media, and bus interfaces. The bus interface can be used to connect network adapters, etc., to the processing system via the bus. The network adapter can be used to implement signal processing functions. In some respects, user interfaces (e.g., keypads, displays, mice, joysticks, etc.) may also be connected to the bus. The bus may also link various other circuits, such as timing sources, peripherals, voltage regulators, power management circuits, etc., which are well known in the art and will not be described further.

[0163] A processor may be responsible for managing the bus and general-purpose processing, including executing software stored on a machine-readable medium. A processor may be implemented using one or more general-purpose processors and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuitry that can execute software. Software should be interpreted broadly as instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. By way of example, a machine-readable medium may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, disks, optical disks, hard disks, or any other suitable storage medium, or any combination thereof. A machine-readable medium may be embodied as a computer program product. A computer program product may include packaging material.

[0164] In a hardware implementation, machine-readable media can be part of a processing system separate from the processor. However, as those skilled in the art will readily understand, machine-readable media, or any portion thereof, can be external to the processing system. By way of example, machine-readable media may include transmit lines, carrier waves modulated by data, and / or computer components separate from the device, all accessible to the processor via a bus interface. Alternatively or additionally, machine-readable media, or any portion thereof, may be integrated into the processor, such as in the case of a cache and / or a general-purpose register file. Although the various components discussed may be described as having a specific location, such as local components, they may also be configured in various ways, such as certain components being configured as part of a distributed computing system.

[0165] The processing system may be configured as a general-purpose processing system having one or more microprocessors providing processor functionality and external memory providing at least a portion of machine-readable medium, all of which are linked together with other supporting circuitry via an external bus architecture. Alternatively, the processing system may include one or more neuromorphic processors for implementing the described neuron and nervous system models. As another alternative, the processing system may be implemented using an application-specific integrated circuit (ASIC) having a processor, bus interface, user interface, supporting circuitry, and at least a portion of machine-readable medium integrated on a single chip, or using one or more field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic components, discrete hardware components, or any other suitable circuitry, or any combination of circuitry capable of performing the various functionalities described throughout this disclosure. Those skilled in the art will recognize how best to implement the described functionality of the processing system depends on the specific application and the overall design constraints imposed on the system as a whole.

[0166] Machine-readable media may include multiple software modules. These software modules include instructions that, when executed by a processor, cause the processing system to perform various functions. Software modules may include send and receive modules. Each software module may reside in a single storage device or be distributed across multiple storage devices. For example, when a triggering event occurs, a software module may be loaded from a hard disk drive into RAM. During the execution of a software module, the processor may load some of the instructions into a cache to improve access speed. One or more cache lines may then be loaded into a general-purpose register file for processor execution. When the functionality of a software module is referred to below, it will be understood that such functionality is implemented by the processor when executing the instructions from that software module. Furthermore, it should be understood that aspects of this disclosure result in improvements to the functionality of a processor, computer, machine, or other system implementing such aspects.

[0167] If implemented in software, the functions may be stored as one or more instructions or codes on or transmitted through a computer-readable medium. A computer-readable medium includes both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one location to another. A storage medium can be any available medium accessible to a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and is accessible to a computer. Additionally, any connection is also appropriately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, optical fiber, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared (IR), radio, and microwave, then such coaxial cable, optical fiber, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. The disks and optical discs used include compact discs (CDs), laser discs, optical discs, digital multifunction discs (DVDs), floppy disks, and Blu-ray discs. ® Optical discs, where magnetic disks typically reproduce data magnetically, and optical discs reproduce data optically using lasers. Therefore, in some aspects, computer-readable media may include non-transitory computer-readable media (e.g., tangible media). Furthermore, in other aspects, computer-readable media may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.

[0168] Therefore, some aspects may include a computer program product for performing the presented operations. For example, such a computer program product may include a computer-readable medium on which instructions are stored (and / or encoded) that can be executed by one or more processors to perform the described operations. In some aspects, the computer program product may include packaging material.

[0169] Furthermore, it should be understood that modules and / or other suitable components for performing the described methods and techniques may be downloaded and / or otherwise obtained by the user terminal and / or base station where applicable. For example, such devices can be coupled to a server to facilitate the delivery of components for performing the described methods. Alternatively, the various methods described can be provided via storage components (e.g., RAM, ROM, physical storage media such as CDs or floppy disks) so that the user terminal and / or base station can obtain the various methods once the storage component is coupled to or provided to the device. In addition, any other suitable techniques suitable for providing the described methods and techniques to the device may be utilized.

[0170] It should be understood that the claims are not limited to the precise configurations and components illustrated above. Various modifications, variations, and alterations may be made to the arrangement, operation, and details of the methods and apparatus described above without departing from the scope of the claims.

Claims

1. An apparatus, the apparatus comprising: At least one memory; and At least one processor, coupled to the at least one memory, is configured to: Estimating the local curvature of the loss landscape using a neural network; The parameters to be removed from the neural network are dynamically assigned based on the local curvature; and The remaining weights of the neural network are updated based on the parameters to be removed.

2. The apparatus of claim 1, wherein the at least one processor is further configured to estimate the local curvature based on: The weight values ​​or activation outer products obtained from the forward pass of the neural network, and The outer product of gradients from the backpropagation of the neural network.

3. The apparatus of claim 1, wherein the at least one processor is further configured to update the remaining weights based on the Kronecker factor approximation curvature (KFAC).

4. The apparatus of claim 1, wherein the at least one processor is further configured to update the remaining weights based on the assumed dependencies of the elements of the neural network.

5. The apparatus of claim 1, wherein the at least one processor is further configured to update the remaining weights by computing a single relevant weight update associated with removing all assigned parameters together.

6. The apparatus of claim 1, wherein the at least one processor is further configured to iteratively estimate the local curvature and dynamically allocate the parameters to be removed.

7. The apparatus of claim 1, wherein the at least one processor is further configured to iteratively and dynamically allocate the parameters to be removed and to update the remaining weights using low-rank adaptation.

8. A processor-implemented method, the method comprising: Estimating the local curvature of the loss landscape using a neural network; The parameters to be removed from the neural network are dynamically assigned based on the local curvature. as well as The remaining weights of the neural network are updated based on the parameters to be removed.

9. The method of claim 8, wherein the estimation of the local curvature is based on: The weight values ​​or activation outer products obtained from the forward pass of the neural network, and The outer product of gradients from the backpropagation of the neural network.

10. The method of claim 8, wherein the remaining weights are updated based on the Kronecker factor approximation curvature (KFAC).

11. The method of claim 8, wherein the remaining weights are updated based on the assumption of the dependencies of the elements of the neural network.

12. The method of claim 8, wherein updating the remaining weights further comprises calculating a single relevant weight update associated with removing all assigned parameters together.

13. The method of claim 8, further comprising iteratively estimating the local curvature and dynamically assigning the parameters to be removed.

14. The method of claim 8, further comprising iteratively and dynamically allocating the parameters to be removed and using low-rank adaptation to update the remaining weights.

15. A non-transitory computer-readable medium having program code recorded thereon, the program code being executed by a processor and comprising: Program code used to estimate the local curvature of the lost landscape for a neural network; Program code for dynamically allocating parameters to be removed from the neural network based on the local curvature; as well as Program code for updating the remaining weights of the neural network based on the parameter to be removed.

16. The non-transitory computer-readable medium of claim 15, wherein the program code for estimating the local curvature is based on: The weight values ​​or activation outer products obtained from the forward pass of the neural network, and The outer product of gradients from the backpropagation of the neural network.

17. The non-transitory computer-readable medium of claim 15, wherein the program code for updating the remaining weights is based on the Kronecker factor approximation curvature (KFAC).

18. The non-transitory computer-readable medium of claim 15, wherein the program code for updating the remaining weights is based on an assumption of the element dependencies of the neural network.

19. The non-transitory computer-readable medium of claim 15, wherein the program code for updating the remaining weights further comprises program code for calculating a single relevant weight update associated with the removal of all assigned parameters together.

20. The non-transitory computer-readable medium of claim 15, wherein the program code further comprises program code for iteratively estimating the local curvature and dynamically allocating the parameter to be removed.

21. The non-transitory computer-readable medium of claim 15, wherein the program code further comprises program code for iteratively and dynamically allocating the parameters to be removed and updating the remaining weights using low-rank adaptation.

22. An apparatus for wireless communication, the apparatus comprising: Components used to estimate the local curvature of the lost landscape for neural networks; A component for dynamically allocating parameters to be removed from the neural network based on the local curvature; and A component used to update the remaining weights of the neural network based on the parameters to be removed.

23. The apparatus for wireless communication according to claim 22, wherein the component for estimating the local curvature is based on: The weight values ​​or activation outer products obtained from the forward pass of the neural network, and The outer product of gradients from the backpropagation of the neural network.

24. The apparatus for wireless communication according to claim 22, wherein the component for updating the remaining weights is based on the Kronecker factor approximation curvature (KFAC).

25. The apparatus for wireless communication according to claim 22, wherein the component for updating the remaining weights is based on an assumption of the dependencies of the elements of the neural network.

26. The apparatus for wireless communication according to claim 22, wherein the component for updating the remaining weights further includes a component for calculating a single relevant weight update associated with removing all assigned parameters together.

27. The apparatus for wireless communication according to claim 22, the apparatus further comprising means for iteratively estimating the local curvature and dynamically allocating parameters to be removed.

28. The apparatus for wireless communication according to claim 22, the apparatus further comprising means for iteratively and dynamically allocating parameters to be removed and updating the remaining weights using low-rank adaptation.