Adaptive accuracy for deep neural network models

By dynamically adjusting the accuracy of individual weights of the neural network model, and using heuristics to determine the storage accuracy of the weights based on the influence of weights on the model output and value fluctuations, it solves the problem that it is difficult to maintain the model and training quality while reducing the implementation cost, and achieves more efficient resource utilization and processing time.

CN120106153APending Publication Date: 2025-06-06HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410570677.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-04
Filing Date
2024-05-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to maintain model and training quality while reducing the cost of neural network model implementation, especially when processing large-scale deep neural network models, resource consumption and processing time are high.

Method used

By dynamically adjusting the accuracy of individual weights of the neural network model, heuristics are used to determine the storage accuracy of the weight based on the influence of the weight on the model output and the value fluctuation, and personalized accuracy adjustment of the weight is achieved.

Benefits of technology

This method can better balance the savings of memory hardware resources and the improvement of model/training quality, reduce the monetary and resource costs of implementing neural network models, and improve processing time efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106153A_ABST
    Figure CN120106153A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to adaptive accuracy for deep neural network models. Examples of the disclosed technology provide computerized systems and methods for dynamically adjusting the amount of precision (i.e., the number of bits) used to represent and store individual weights of a neural network model (e.g., a DNN model) in response to training. Examples may use various heuristic methods to intelligently determine dynamic levels of precision for these personalizations. For example, a heuristic method may include one or more of: (1) a measure that quantifies a magnitude of an influence of a respective weight on an output of a neural network model during a latest set of training iterations (weights with relatively higher influence may be represented using higher accuracy); and (2) measurements that quantify the magnitude of the value fluctuations for the respective weights during the last set of training iterations (weights with relatively smaller value fluctuations may be represented using higher accuracy (i.e., finer accuracy)).
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] A machine learning model may include an algorithm-based computer program that is trained to recognize patterns in data and make predictions and / or classifications based on such learned pattern recognition.

[0002] A neural network model (sometimes called an artificial neural network) is a machine learning model inspired by the structure of the human brain. For example, a neural network model may include a series of algorithms that are trained to recognize patterns in data in a way that mimics the way the human brain works. For conceptualization, a neural network model is sometimes described as consisting of interconnected "neurons" arranged in layers (much like the human brain). Here, each neuron can represent an algorithm that receives one or more inputs and produces an output. For example, a common neuron-based conceptualization describes a neural network model in terms of: (a) an input layer of neurons that receives input data and produces an output; (b) one or more hidden layers of neurons that receive weighted outputs from the input layer (or, in the case of multiple hidden layers, receive weighted outputs from previous hidden layers) and produce their own outputs; and (c) an output layer of neurons that receives weighted outputs from the last hidden layer and produces an output prediction / classification. In such a neuron-based conceptualization, a corresponding neuron is typically connected to one or more neurons in (multiple) other layers. For example, a corresponding neuron in the first hidden layer can be connected to one or more neurons in the input layer and receive weighted outputs from it. Similarly, one or more neurons in the second hidden layer can be connected to corresponding neurons in the first hidden layer and receive weighted outputs therefrom. As described above, each "connection" (sometimes referred to as a synapse) between two neurons is associated with a digital weight that is multiplied by the output of one neuron to produce an input received by another neuron. For example, a connection between a first neuron of an input layer and a corresponding neuron of a first hidden layer can have a first weight. This first weight is multiplied by the output of a first neuron of the input layer to produce an input received by a corresponding neuron of the first hidden layer. Relatedly, a connection between a corresponding neuron of the first hidden layer and a first neuron of a second hidden layer can have a second weight. This second weight is multiplied by the output of a corresponding neuron of the first input layer to produce an input received by a first neuron of the second hidden layer. In this neuron-based conceptualization, the connection weights (sometimes more simply referred to herein as weights) of a neural network model are modified / tuned in response to training. In other words, by dynamically modifying its connection weights during training, a neural network model can "learn" to produce more accurate predictions / classifications.

[0003] While the above neuron-based conceptualization is helpful in understanding the neural network model, another representation / conceptualization of the neural network model involves weight matrices and matrix multiplication. In this matrix-based representation, the weights of the neural network model correspond to the elements of the weight matrix. The weight matrix and / or the number of matrix multiplication operations of the neural network model can correspond to the number of layers of the neural network model. For example, a simple three-layer neural network model may include: (a) a first weight matrix, which is multiplied with the input vector to produce a first output vector (the first weight matrix is ​​similar to the connection weights between the input layer of neurons and the hidden layer of neurons); and (b) a second weight matrix, which is multiplied with the first output vector to produce a second output vector (the second weight matrix is ​​similar to the connection weights between the hidden layer of neurons and the output layer of neurons). The second output vector can embody (or otherwise be used to make) the final prediction / classification of the neural network model. Returning to the neuron-based conceptualization, the number of neurons in the input layer and the first hidden layer of the neural network model can correspond to the dimensions of the weight matrix representing its connection weights (i.e., the number of columns and rows, respectively). Likewise, the number of neurons in the final hidden layer and output layer of the neural network model may correspond to the dimensions of the weight matrix representing the weights of their connections. As described above, the weights / elements of the weight matrix of the neural network model are modified / tuned in response to training. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] According to one or more different examples, the present disclosure is described in detail with reference to the following figures. The figures provided are for illustration purposes only and show only examples.

[0005] Figure 1A-Figure 1C An exemplary neuron-based conceptualization of a neural network model according to one or more examples is shown;

[0006] Figure 2 An example computing component for storing weights of a neural network model using different amounts of precision according to one or more examples is shown;

[0007] Figure 3 An example computing system for storing weights of a neural network model using different amounts of precision according to one or more examples is shown;

[0008] Figure 4 An example computing system is shown that can be used to dynamically adjust the amount of precision used to represent individual weights of a neural network in response to training according to one or more examples; and

[0009] Figure 5 An example flow chart is shown that can be used to dynamically adjust the amount of precision used to represent individual weights of a neural network in response to training according to one or more examples.

[0010] The drawings are not exhaustive and do not limit the disclosure to the precise forms disclosed. DETAILED DESCRIPTION

[0011] A deep neural network (DNN) model (sometimes referred to as a deep learning neural network) may refer to a neural network model with more than one "hidden layer". In other words, in terms of matrix-based representation, a DNN model may refer to a neural network model with three or more weight matrices / matrix multiplications. DNN models are popular in a variety of applications, including generative artificial intelligence (AI), due to their sophisticated / advanced predictive capabilities. Generally speaking, as the number of layers in a neural network model increases, its learning capabilities also increase. As a result, DNN models can be huge, sometimes including billions of weights and relying on thousands / millions of matrix multiplications to produce predictions / classifications. However, this scale comes at a severe monetary and resource cost.

[0012] For example, the processing and memory hardware required to implement large DNN models may be considerable. Therefore, DNN models are typically implemented on many physical computing units (e.g., general processing units (GPUs)). The implementation of many physical computing units may significantly increase monetary and resource costs and processing time. For example, in the case where the weights of a DNN model are stored across many physical computing units (e.g., GPUs and / or hardware accelerators), implementing the DNN model typically requires a model serving system that requests weights from multiple physical computing units as needed during runtime (e.g., for matrix multiplication). This request / movement of weights across many physical computing units will increase the huge consumption of processing resources and processing time required to implement the DNN model.

[0013] In general, the monetary and resource costs associated with implementing a neural network model (e.g., a DNN model) can be reduced by using lower precision (i.e., a lower number of memory bits) to represent the weights of the neural network model. This is in part because, with lower precision, a greater number of weights can be stored in a given (physical or logical) memory hardware segment. In other words, the amount of memory hardware required to store the neural network model weights can be reduced by storing / representing the neural network model weights with lower precision. The reduction in the amount of memory hardware required to store the neural network model weights can significantly reduce the monetary and resource costs for implementing the neural network model. For example, the weights of the neural network model can be stored across fewer physical computing units, thereby reducing: (1) memory hardware material costs; and (2) data delays associated with requesting and moving weights to different physical computing units as needed during runtime (e.g., for matrix multiplication). Related to the above, matrix multiplication can generally be performed faster with lower precision weights, which in some cases helps to reduce the processing time of the neural network model.

[0014] However, representing the weights of a neural network model with lower precision often degrades the model and training quality. For example, with lower precision weights, the neural network model may not achieve weight convergence (as used herein, weight convergence may refer to the convergence of the individual weight values ​​of the neural network model, which typically indicates that training is complete). In contrast, a neural network model with relatively higher precision weights is more likely to converge, and also tends to converge faster (e.g., using fewer training iterations). However (as described above), higher precision weight representation / storage requires more and more memory hardware, which comes with all the monetary and resource costs mentioned above.

[0015] In summary, an innovative technique is needed that can intelligently balance the competing interests of memory hardware resource conservation (typically achieved by representing / storing neural network model weights at relatively lower precision) and improved model / training quality (typically achieved by representing / storing neural network model weights at relatively higher precision).

[0016] In this context, examples of the disclosed technology provide computerized systems and methods for dynamically adjusting the amount of precision (i.e., the number of bits) used to represent and store individual weights of a neural network model (e.g., a DNN model) in response to training. The examples may use various heuristics to intelligently determine these personalized, dynamic levels of precision. For example, the heuristics may include one or more of the following: (1) a measure that quantifies the magnitude of the effect of the corresponding weight on the output of the neural network model during a recent set of training iterations (weights with relatively high effects may be represented using higher precision); and (2) a measure that quantifies the magnitude of the value fluctuations of the corresponding weight during a recent set of training iterations (weights with relatively small value fluctuations may be represented using higher precision (i.e., finer precision)).

[0017] The examples improve upon potential alternative techniques that simply: (a) statically modify weight precision (e.g., prior to training); and / or (b) uniformly adjust the precision of all weights of a neural network model. In other words, by dynamically and individually adjusting the precision of neural network model weights, the examples can better balance the aforementioned competing interests of memory hardware resource conservation (typically achieved by representing / storing neural network model weights at relatively lower precision) and improved model / training quality (typically achieved by representing / storing neural network model weights at relatively higher precision). Relatedly, the examples improve the functionality of computer memory systems / techniques for implementing neural network models by providing methods for more efficiently storing neural network model weights in computer memory hardware.

[0018] This improvement exploits an intelligent insight that higher precision for certain weights of a neural network model (e.g., those weights that had a relatively large effect on the neural network model output during a recent set of training iterations and / or those weights that had relatively small value fluctuations during a recent set of training iterations) can improve / affect weight convergence more than higher precision for other weights of the neural network model (e.g., those weights that had a relatively small effect on the neural network model output during a recent set of training iterations and / or those weights that had relatively high value fluctuations during a recent set of training iterations). By exploiting this insight to define a heuristic method for dynamically determining the amount of precision used to represent / store individual weights of a neural network model, examples can better balance the aforementioned competing interests of memory hardware resource conservation (typically achieved by representing / storing neural network model weights with relatively lower precision) and improved model / training quality (typically achieved by representing / storing neural network model weights with relatively higher precision).

[0019] For example, a system of the disclosed technology may: (1) in response to completion of a first set of training iterations of a neural network model, first classify the weights of the neural network model into a first memory block and a second memory block according to a heuristic method; (2) use a first number of memory bits to first store the individual weights first classified into the first memory block; (3) use a second number of memory bits to first store the individual weights first classified into the second memory block, wherein the first number of memory bits is less than the second number of memory bits; (4) in response to completion of a second set of training iterations of the neural network model having the first stored weights, second classify the first stored weights of the neural network model into the first memory block and the second memory block according to a heuristic method; (5) use the first number of memory bits to secondly store the individual weights secondly classified into the first memory block; and (6) use the second number of memory bits to secondly store the individual weights secondly classified into the second memory block. As described above, the heuristic method can include one or more of the following: (1) a measure that quantifies the magnitude of the influence of the corresponding weight on the output of the neural network model during the most recent set of training iterations (weights with relatively higher influence can be represented using a second (i.e., higher) number of memory bits); and (2) a measure that quantifies the magnitude of the value fluctuations of the corresponding weight during the most recent set of training iterations (weights with relatively smaller value fluctuations can be represented using a second (i.e., higher / finer) number of memory bits).

[0020] In some examples, first storing individual weights that are first classified into the first memory block may include first storing the individual weights that are first classified into the first memory block in individual (physical or logical) sub-segments of the first memory hardware (physical or logical) segment. Here, the corresponding sub-segment of the first memory hardware segment may include a first number of memory bits. Relatedly, first storing individual weights that are first classified into the second memory block may include first storing the individual weights that are first classified into the second memory block in individual (physical or logical) sub-segments of the second memory hardware (physical or logical) segment. Here, the corresponding sub-segment of the second memory hardware segment may include a second number of memory bits. In some examples, the first memory hardware segment may be implemented on a first streaming multiprocessor (SM) of a general processing unit (GPU), and the second memory hardware segment may be implemented on a second SM of a GPU (or another GPU). By grouping weights of similar precision in the same (physical or logical) memory unit (i.e., corresponding memory hardware segment / corresponding SM), the system can reduce data latency associated with moving weights between different memory units at runtime. Relatedly, the system can improve the ease / efficiency of programming and processing by causing a given (physical or logical) computing unit to perform matrix operations involving weights with common precision. For example, the system can first perform intra-memory block matrix operations (i.e., matrix operations involving weights stored in the same memory block / segment of the memory hardware). Then, in response to the completion of the intra-memory block matrix operations, the system can perform inter-memory block matrix operations (i.e., matrix operations between weights classified into a first memory block / stored in a first memory hardware segment and weights classified into a second memory block / stored in a second memory hardware segment). This organized sequence of performing matrix operations can be more efficient (i.e., resulting in fewer weight movements across memory hardware units / segments) than alternative methods of performing matrix operations involving weights stored across different memory hardware segments and / or different physical computing units in a less organized manner. As described above, improving matrix operation efficiency can improve processing time and, in some cases, reduce power consumption and monetary costs.

[0021] Examples of the disclosed technology will now be described with reference to the accompanying drawings.

[0022] Figure 1A-Figure 1C An exemplary neuron-based conceptualization of a neural network model 100 according to various examples of the disclosed technology is shown.

[0023] As shown, the neuron-based conceptualization of the neural network model 100 includes four layers of "neurons". The input layer of the neural network model 100 includes neurons A1 and A2. The first hidden layer of the neural network model 100 includes neurons B1, B2, and B3. The second hidden layer of the neural network model 100 includes neurons C1, C2, and C3. The output layer of the neural network model 100 includes neurons D1 and D2.

[0024] As described above, each neuron of the neural network model 100 can represent an algorithm that receives one or more inputs and produces an output. For example, neurons A1 and A2 of the input layer can receive inputs to the neural network model 100 (e.g., digital vectors representing input images to be classified) and produce outputs. Neurons B1-B3 of the first hidden layer can receive weighted outputs from neurons A1 and A2 of the input layer and produce their own outputs. Relatedly, neurons C1-C3 of the second hidden layer can receive weighted outputs from neurons B1-B3 of the first hidden layer and produce their own outputs. Finally, neurons D1-D2 of the output layer can receive weighted outputs from neurons C1-C3 of the second hidden layer and produce their own outputs. The outputs (e.g., numerical vectors) of neurons D1-D2 can embody (or otherwise be used to make) the final prediction / classification of the neural network model 100.

[0025] In the neuron-based conceptualization of the neural network model 100, each neuron is connected to one or more neurons in the other layer(s). For example, neuron B1 of the first hidden layer is connected to and receives weighted outputs from neurons A1 and A2 of the input layer. Similarly, neurons C1-C3 are connected to and receive weighted outputs from neuron B1. As described above, each "connection" (sometimes called a synapse) between two neurons is associated with a digital weight that is multiplied by the output of one neuron to produce the input received by the other neuron. Figure 1A-Figure 1C In the diagram, connections between neurons are shown by arrows. For example, connection A1->B1 connects neurons A1 and B1. Similarly, connection B3->C1 connects neurons B3 and C1.

[0026] As described above, each connection between neurons can have a corresponding weight. For example, connection A1->B1 can have a first weight, connection A1->B2 can have a second weight, connection A1->B3 can have a third weight, and so on. The corresponding weights are multiplied by the output of one neuron to produce an input received by another neuron. For example, the first weight of connection A1->B1 is multiplied by the output of neuron A1 to produce the input received by neuron B1. Similarly, the second weight of connection A1->B2 is multiplied by the output of neuron A1 to produce the input received by neuron B2.

[0027] As mentioned above, in Figure 1A-Figure 1C In the neuron-based conceptualization of the neural network model 100, the connection weights (sometimes more simply referred to herein as weights) are modified / tuned in response to training. In other words, by dynamically modifying its connection weights during training, the neural network model 100 can "learn" to produce more accurate predictions / classifications.

[0028] Although Figure 1A-Figure 1C The neuron-based conceptualization of the neural network model is helpful for understanding the neural network model, but another representation / conceptualization of the neural network model involves weight matrices and matrix multiplication. In the matrix-based representation, the weights of the neural network model 100 correspond to the elements of the weight matrix. The number of weight matrices and / or matrix multiplication operations of the neural network model 100 can correspond to the number of layers of the neural network model. For example, the neural network model 100 may include: (a) a first weight matrix that is multiplied with an input vector to produce a first output vector (the first weight matrix may be similar to Figure 1A-Figure 1C (b) a second weight matrix, which is multiplied by the first output vector to produce a second output vector (the second weight matrix can be similar to Figure 1A-Figure 1C ) a connection weight between the first hidden layer and the second hidden layer of the neural network model 100 in the neural network model 100); and (c) a third weight matrix, which is multiplied by the second output vector to generate a third output vector (the third weight matrix can be similar to Figure 1A-Figure 1C The third output vector may represent (or otherwise be used to make) the final prediction / classification of the neural network model 100. Figure 1A-Figure 1C Based on the conceptualization of neurons, the number of neurons of the input layer (i.e., two) and the number of neurons of the first hidden layer (i.e., three) of the neural network model 100 can correspond to the dimensions of the first weight matrix representing / corresponding to their connection weights (i.e., the number of columns and rows, respectively). Similarly, the number of neurons of the first hidden layer (i.e., three) and the number of neurons of the second hidden layer of the neural network model 100 can correspond to the dimensions of the second weight matrix representing / corresponding to their connection weights. Similarly, the number of neurons of the second hidden layer (i.e., three) and the number of neurons of the output layer (i.e., two) of the neural network model 100 can correspond to the dimensions of the third weight matrix representing / corresponding to their connection weights. As described above, the weights / elements of the weight matrix of the neural network model 100 are modified / tuned in response to training.

[0029] Because the neural network model 100 includes more than one "hidden layer", it may be referred to as a deep neural network (DNN) model. However, it should be understood that the neural network model 100 is a significantly simplified version of a typical DNN model. As described above, many DNN models are huge, often including billions of weights, and rely on thousands / millions of matrix multiplications to produce predictions / classifications. However (as described above), this scale comes at a severe monetary and resource cost.

[0030] For example, the processing and memory hardware required to implement large DNN models may be considerable. Therefore, DNN models are typically implemented on many physical computing units (e.g., general processing units (GPUs)). The implementation of many physical computing units may significantly increase monetary and resource costs and processing time. For example, in the case where the weights of a DNN model are stored across many physical computing units (e.g., GPUs and / or hardware accelerators), implementing the DNN model typically requires a model serving system that requests weights from multiple physical computing units as needed during runtime (e.g., for matrix multiplication). This request / movement of weights across many physical computing units will increase the huge consumption of processing resources and processing time required to implement the DNN model.

[0031] In general, the monetary and resource costs associated with implementing a neural network model (e.g., a DNN model) can be reduced by using lower precision (i.e., a smaller number of memory bits) to represent the weights of a neural network model. This is in part because, with lower precision, a greater number of weights can be stored in a given (physical or logical) memory hardware segment. In other words, the amount of memory hardware required to store the weights of a neural network model can be reduced with lower precision. The reduction in the amount of physical memory hardware required to store the weights of a neural network model can significantly reduce the monetary and resource costs required to implement the neural network model. For example, the weights of a neural network model can be stored across fewer physical computing units, thereby reducing: (1) memory hardware material costs; and (2) data delays associated with requesting and moving weights to different physical computing units as needed during runtime (e.g., for matrix multiplication). Related to the above, matrix multiplication can typically be performed faster with lower precision weights, which in some cases helps to reduce the processing time of a neural network model.

[0032] However, representing the weights of a neural network model with lower precision generally degrades the model and training quality. For example, with lower precision weights, the neural network model may never achieve weight convergence (as used herein, weight convergence can refer to the convergence of the individual weight values ​​of the neural network model, which generally indicates that training is complete). In contrast, a neural network model with relatively higher precision weights is more likely to converge, and also tends to converge faster (e.g., using fewer training iterations). However (as described above), higher precision weight representation / storage requires more and more memory hardware, which comes with all the monetary and resource costs mentioned above.

[0033] In summary, there is a pressing need for innovative techniques that intelligently balance the competing interests of memory hardware resource conservation (typically achieved by representing / storing neural network model weights at relatively low precision) and improved model / training quality (typically achieved by representing / storing neural network model weights at relatively high precision).

[0034] In this context (as described above), examples of the disclosed technology provide computerized systems and methods for dynamically adjusting the amount of precision (i.e., the number of bits) used to represent and store individual weights of a neural network model (e.g., the weights of neural network model 100, or more precisely, the weights of a larger version of neural network model 100) in response to training. The examples may use various heuristics to intelligently determine these personalized, dynamic levels of precision. For example, the heuristics may include one or more of the following: (1) a measure that quantifies the magnitude of the effect of the corresponding weight on the output of the neural network model during a recent set of training iterations (weights with relatively high effects may be represented using higher precision); and (2) a measure that quantifies the magnitude of fluctuations in the value of the corresponding weight during a recent set of training iterations (weights with relatively small value fluctuations may be represented using higher precision (i.e., finer precision)).

[0035] The examples improve upon potential alternative techniques that simply: (a) statically modify weight precision (e.g., prior to training); and / or (b) uniformly adjust the precision of all weights of a neural network model. In other words, by dynamically and individually adjusting the precision of neural network model weights, the examples can better balance the aforementioned competing interests of memory hardware resource conservation (typically achieved by representing / storing neural network model weights at relatively lower precision) and improved model / training quality (typically achieved by representing / storing neural network model weights at relatively higher precision). Relatedly, the examples improve the functionality of computer memory systems / techniques for implementing neural network models by providing methods for more efficiently storing neural network model weights in computer memory hardware.

[0036] This improvement exploits an intelligent insight that higher precision for certain weights of a neural network model (e.g., those weights that had a relatively large effect on the neural network model output during a recent set of training iterations and / or those weights that had relatively small value fluctuations during a recent set of training iterations) can improve / affect weight convergence more than higher precision for other weights of the neural network model (e.g., those weights that had a relatively small effect on the neural network model output during a recent set of training iterations and / or those weights that had relatively high value fluctuations during a recent set of training iterations). By exploiting this insight to define a heuristic method for dynamically determining the amount of precision used to represent / store individual weights of a neural network model, examples can better balance the aforementioned competing interests of memory hardware resource conservation (typically achieved by representing / storing neural network model weights with relatively lower precision) and improved model / training quality (typically achieved by representing / storing neural network model weights with relatively higher precision).

[0037] For example (again refer to Figure 1A ), the system 102 of the disclosed technology may apply a first set of training iterations to the neural network model 100. During this first set of training iterations, each weight of the neural network model 100 may be represented / stored using a first (e.g., low) number of memory bits.

[0038] In response to the first set of training iterations, the system 102 may calculate a heuristic for each weight of the neural network model 100. For example, the system 102 may calculate: (1) the magnitude of the effect of each weight on the output of the neural network model 100 during the first set of training iterations (as described above, weights with relatively higher effect magnitudes are better candidates for relatively higher precision representation / storage); and (2) the magnitude of the value fluctuation of each weight during the first set of training iterations (as described above, weights with relatively smaller value fluctuations are better candidates for relatively higher precision representation / storage). The system 102 may then use these calculations to calculate a heuristic (e.g., a number) for first classifying each weight of the neural network model 100 into a first memory block (e.g., a first memory logical or physical segment that uses a first number of memory bits to represent / store individual weights), a second memory block (e.g., a second memory logical or physical segment that uses a second number of memory bits to represent / store individual weights), or a third memory block (e.g., a third memory logical or physical segment that uses a third number of memory bits to represent / store individual weights). Here, the system 102 may first classify weights having heuristic values ​​in a first value range into the first memory block. Relatedly, the system 102 can first classify the weights with heuristic method values ​​in the second value range to the second memory block. Similarly, the system 102 can first classify the weights with heuristic method values ​​in the third value range to the third memory block. Here, the first value range to the third value range can be continuous with each other. Relatedly, the first range can have a minimum value for the heuristic method, the second range can have a second minimum value for the heuristic method, and the third range can have a highest value for the heuristic method. Here, it should be understood that the above-mentioned heuristic method calculation, memory block arrangement and classification method are only illustrative examples. Generally, various heuristic methods and heuristic method calculations, two or more memory blocks, and various classification methods can be used.

[0039] As used herein, a set of training iterations may refer to a specific (in some cases, predetermined) number of forward and / or backward passes through the training data of a neural network model during training. During each training iteration, the neural network model may process a batch of data, calculate loss and / or error based on predicted and actual results, and update its weights using an optimization algorithm. The number of training iterations in a set may be determined based on various factors, such as the size of the training data set, the batch size, etc.

[0040] As described above, prior to the second set of training iterations, system 102 may store the first classified weights of neural network model 100 according to the corresponding memory blocks of its neural network model 100. For example, individual weights that are first classified into the first memory block may be stored using a first (e.g., low) number of memory bits, respectively. Individual weights that are first classified into the second memory block may be stored using a second (e.g., medium) number of memory bits, respectively. Individual weights that are first classified into the third memory block may be stored using a third (e.g., high) number of memory bits, respectively. Figure 1A-Figure 1C In a particular example, a first number of memory bits may correspond to a relatively lower precision, a second number of memory bits may correspond to a relatively medium precision, and a third number of memory bits may correspond to a relatively higher precision.

[0041] Figure 1A-Figure 1C The thickness of the arrows showing the connections of the neural network model 100 provides a visual representation of the level of precision for the respective weights associated with the respective connections. The thinnest arrows correspond to a first (i.e., low) number of memory bits. Arrows of medium thickness correspond to a second (i.e., medium) number of memory bits. The thickest arrows correspond to a third (i.e., high) number of memory bits.

[0042] like Figure 1A As shown, during a first set of training iterations, each weight is represented / stored using a first (e.g., low) number of memory bits. Figure 1B As shown, during the second set of training iterations, system 102 represents / stores certain bits with higher precision. That is, system 102 utilizes a second (i.e., medium) number of memory bits to store individual weights associated with connections A1->B2, A2->B1, C1->D1, and C3->D2. Relatedly, system 102 utilizes a third (i.e., high) number of memory bits to store individual weights associated with connections A1->B1, B1->C1, and C1->D2. As described above, these personalized precision classifications performed dynamically during training can better balance the above-mentioned competing interests of memory hardware resource conservation and improved model / training quality compared to potential alternative techniques of merely (a) statically (e.g., before training) modifying weight precision and / or (b) uniformly adjusting the precision of all weights of a neural network model.

[0043] Since the examples of the disclosed technology are designed for understanding, the values ​​of the heuristics for the individual weights of the neural network model 100 may evolve during training. To illustrate this dynamic nature, the system 102 may iteratively / repeatedly calculate the heuristics for each weight of the neural network model 100 during training (e.g., after each set of training iterations). Therefore, in response to the second set of training iterations, the system 102 may again calculate the heuristics for each weight of the neural network model 100. Based on these heuristic calculations, the system 102 may secondarily classify the individual weights of the neural network model 100 into the first memory block to the third memory block in the same / similar manner as described above. Then, before the third set of training iterations, the system 102 may store the secondarily classified weights of the neural network model 100 according to the corresponding memory blocks of the neural network model 100. For example, the individual weights secondarily classified into the first memory block may be stored using a first (e.g., low) number of memory bits, respectively. The individual weights secondarily classified into the second memory block may be stored using a second (e.g., medium) number of memory bits, respectively. The individual weights that are secondarily sorted into the third memory block may each be stored using a third (eg, high) number of memory bits. Figure 1C As shown, this may involve adjusting the storage precision of individual weights associated with connections A1->B1, A2->B1, A2->B2, A2->B3, B1->C1, B3->C1, C1->D1, C1->D2, and C2->D1 (e.g., when Figure 1B Go to Figure 1C , see the change in thickness of the arrows for these connections).

[0044] Figure 2 An example computing component 200 for storing weights of a neural network model using different amounts of precision according to various examples of the disclosed technology is shown. In some examples, the computing component 200 can be implemented on the system 102 of FIG. 1 .

[0045] Computing component 200 may be, for example, a server computer, a controller, a general purpose processing unit (GPU), or any other similar computing component capable of processing and storing data. Figure 2 In the example implementation of , computing component 200 includes hardware processor 212 , machine-readable storage medium 214 , memory hardware segment 216 , and memory hardware segment 226 .

[0046] Hardware processor 212 may include one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for retrieving and executing instructions stored in machine-readable storage medium 214. Hardware processor 212 may fetch, decode, and execute instructions, such as instructions for storing weights of a neural network model (e.g., neural network model 100) using different amounts of precision. As an alternative or in addition to retrieving and executing instructions, hardware processor 212 may include one or more electronic circuits including electronic components for performing the functions of one or more instructions, such as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or other electronic circuits.

[0047] A machine-readable storage medium, such as machine-readable storage medium 214, may be any electronic, magnetic, optical, or other storage device that contains or stores executable instructions. Thus, machine-readable storage medium 214 may be, for example, a random access memory (RAM), a non-volatile RAM (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disk, etc. In some examples, machine-readable storage medium 214 may be a non-transitory storage medium, where the term "non-transitory" does not include a transient propagation indicator. Machine-readable storage medium 214 may be encoded with executable instructions, for example, instructions for storing weights of a neural network model (e.g., neural network model 100) using different amounts of precision.

[0048] Memory hardware segments 216 and 226 may each be a separate (physical or logical) memory hardware segment for storing weights of a neural network model. Memory hardware segments 216 and 226 may include various types of memory hardware, such as random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), storage devices, optical disks, caches, and / or other dynamic storage devices. In some examples (e.g., where compute component 200 is a GPU), memory hardware segment 216 may be implemented on a first streaming multiprocessor (SM) and memory hardware segment 226 may be implemented on a second SM (such SMs being accessed via Figure 2 ). In these examples, one or more of the hardware processors 212 may be implemented on the first SM, and one or more of the hardware processors 212 may be implemented on the second SM. Relatedly, portions of the machine-readable storage medium 214 may be implemented on the first SM and the second SM, respectively.

[0049] Each of the memory hardware segments 216 and 226 includes (physical or logical) sub-segments, as shown by the boxes within the memory hardware segments 216 and 226. The sub-segments of the memory hardware segment 216 may include a first number of memory bits, and the sub-segments of the memory hardware segment 226 may include a second number of memory bits. As visually shown by the size of the boxes in the memory hardware segments 216 and 226, respectively, the first number of memory bits is less than the second number of memory bits.

[0050] As described above, hardware processor 212 may execute instructions stored in machine-readable storage medium 214 to: (1) in response to completion of a first set of training iterations of the neural network model, first sort the weights of the neural network model into a first memory block and a second memory block according to a heuristic method; (2) in memory hardware segment 216, first store the individual weights first sorted into the first memory block using a first number of memory bits; (3) in memory hardware segment 226, first store the individual weights first sorted into the second memory block using a second number of memory bits; (4) in response to completion of a second set of training iterations of the neural network model having the first stored weights, second sort the first stored weights of the neural network model into the first memory block and the second memory block according to a heuristic method; (5) in memory hardware segment 216, second store the individual weights second sorted into the first memory block using the first number of memory bits; and (6) in memory hardware segment 226, second store the individual weights second sorted into the second memory block using the second number of memory bits. As described above, the heuristic method can include one or more of the following: (1) a measure that quantifies the magnitude of the influence of the corresponding weight on the output of the neural network model during the most recent set of training iterations (weights with relatively higher influence can be represented / stored using a second (i.e., higher) number of memory bits); and (2) a measure that quantifies the magnitude of the value fluctuations of the corresponding weight during the most recent set of training iterations (weights with relatively smaller value fluctuations can be represented using a second (i.e., higher / finer) number of memory bits).

[0051] like Figure 2 As shown, storing individual weights classified into the first memory block may include storing the individual weights classified into the first memory block in individual (physical or logical) sub-segments of the memory hardware segment 216, wherein the corresponding sub-segments of the memory hardware segment 216 include a first number of memory bits. Relatedly, storing individual weights classified into the second memory block may include storing the individual weights classified into the second memory block in individual (physical or logical) sub-segments of the memory hardware segment 226, wherein the corresponding sub-segments of the memory hardware segment 226 include a second number of memory bits. As described above, in some examples, the memory hardware segment 216 can be implemented on the first SM of the computing component 200 (see, e.g., Figure 2 ), and the memory hardware segment 226 can be implemented on the second SM of the computing component 200 (see, for example, Figure 2 right frame of the dashed line in the figure).

[0052] By grouping weights of similar precision in the same (physical or logical) memory unit (i.e., memory hardware segments 216 and 226, respectively), the computing component 200 can reduce data latency associated with moving weights between different memory units at runtime. Relatedly, the computing component 200 can improve ease / efficiency of programming and processing by having a given (physical or logical) computing unit perform matrix operations involving weights stored in a common precision.

[0053] For example, one or more of the hardware processors 212 and portions of the machine-readable storage medium 214 may be implemented on a first SM of the computing component 200 having a memory hardware segment 216 (see, e.g., Figure 2 Relatedly, one or more of the hardware processors 212 and portions of the machine-readable storage medium 214 may be implemented on a second SM of the computing component 200 having a memory hardware segment 226 (see, e.g., Figure 2 ). One or more of the hardware processors 212 implemented on the first SM may use the first number of memory bits to perform intra-memory block matrix operations involving weights stored in the memory hardware segment 216 (i.e., matrix operations involving weights stored in the same memory block / segment of the memory hardware). Relatedly, one or more of the hardware processors 212 implemented on the second SM may use the second number of memory bits to perform intra-memory block matrix operations involving weights stored in the memory hardware segment 226. Then, in response to the completion of the intra-memory block matrix operations, one or more processors (e.g., one or more central processing units) in the hardware processors 212 may perform inter-memory block matrix operations (i.e., matrix operations between weights classified into the first memory block / stored in the memory hardware segment 216 and weights classified into the second memory block / stored in the memory hardware segment 226).

[0054] Such an organized sequence of performing matrix operations can be more efficient (i.e., resulting in less weight movement across memory hardware units / segments) than an alternative approach of performing matrix operations involving weights stored across different memory hardware segments and / or different physical compute units in a less organized manner. As described above, improving matrix operation efficiency can improve processing time and, in some cases, reduce power consumption and monetary cost.

[0055] Figure 3An example computing system 300 for storing weights of a neural network model using different amounts of precision according to various examples of the disclosed technology is shown. In some examples, computing system 300 can be implemented on computing system 102 of FIG. 1 .

[0056] As shown, computing system 300 includes three separate computing components: computing component 310; computing component 320; and computing component 320. In some implementations, computing components 310-330 may include separate physical computing units / components. Each of computing components 310-330 may be, for example, a server computer, a controller, a general purpose processing unit (GPU), or any other similar computing component capable of processing and storing data. Although not shown, computing system 300 may include a bus or other communication mechanism for transmitting information / data across different computing components of computing system 300.

[0057] As shown, computing component 310 includes hardware processor 312 and machine-readable storage medium 314. These components may be combined with Figure 2 The hardware processor 212 and machine-readable storage medium 214 are the same / similar as described. As shown, the computing component 320 includes a memory hardware segment 326. As shown, the memory hardware segment 326 can be divided into sub-segments including a first number of memory bits. Although not shown, the computing component 320 can include one or more hardware processors and machine-readable storage media, which can be used to perform matrix operations involving weights of a neural network model stored in the memory hardware segment 326.

[0058] As shown, the computing component 330 includes a memory hardware segment 336. As shown, the memory hardware segment 336 can be divided into sub-segments including a second number of memory bits, the second number of memory bits being greater than the first number of memory bits. Although not shown, the computing component 330 can include one or more hardware processors and machine-readable storage media, which can be used to perform matrix operations involving weights of a neural network model stored in the memory hardware segment 336.

[0059] As described above, hardware processor 312 may execute instructions stored in machine-readable storage medium 314 to: (1) in response to completion of a first set of training iterations of the neural network model, first sort the weights of the neural network model into a first memory block and a second memory block according to a heuristic method; (2) in memory hardware segment 326, first store (or cause to be first stored) the individual weights first sorted into the first memory block using a first number of memory bits; (3) in memory hardware segment 336, first store (or cause to be first stored) the individual weights first sorted into the second memory block using a second number of memory bits; (4) in response to completion of a second set of training iterations for the neural network model having the first stored weights, secondarily classifying the first stored weights of the neural network model into the first memory block and the second memory block according to a heuristic method; (5) in the memory hardware segment 326, secondarily storing (or causing secondarily storing) the individual weights secondarily classified into the first memory block using a first number of memory bits; and (6) in the memory hardware segment 336, secondarily storing (or causing secondarily storing) the individual weights secondarily classified into the second memory block using a second number of memory bits. As described above, the heuristic method may include one or more of the following: (1) a measure that quantifies the magnitude of the influence of the corresponding weight on the output of the neural network model during the most recent set of training iterations (weights with relatively high influence may be represented / stored using a second (i.e., higher) number of memory bits); and (2) a measure that quantifies the magnitude of the value fluctuation of the corresponding weight during the most recent set of training iterations (weights with relatively small value fluctuations may be represented using a second (i.e., higher / more refined) number of memory bits).

[0060] like Figure 3 As shown, storing the individual weights classified into the first memory block may include storing the individual weights classified into the first memory block in individual (physical or logical) sub-segments of the memory hardware segment 326, wherein the corresponding sub-segments of the memory hardware segment 326 include the first number of memory bits. Relatedly, storing the individual weights classified into the second memory block may include storing the individual weights classified into the second memory block in individual (physical or logical) sub-segments of the memory hardware segment 336, wherein the corresponding sub-segments of the memory hardware segment 336 include the second number of memory bits.

[0061] By storing weights with common precision in the same (e.g., physical) computing component, computing system 300 can reduce data latency associated with requesting / moving weights across multiple (e.g., physical) computing components at runtime. Relatedly, computing system 300 can improve ease / efficiency of programming and processing by enabling a given (e.g., physical) computing component to perform matrix operations involving weights stored with common precision. For example, computing component 320 can use a first number of memory bits to perform intra-memory block matrix operations involving weights stored in memory hardware segment 326 (i.e., matrix operations involving weights stored in the same memory block / segment of memory hardware). Relatedly, computing component 330 can use a second number of memory bits to perform intra-memory block matrix operations involving weights stored in memory hardware segment 336. Then, in response to completion of the intra-memory block matrix operations, computing system 300 (e.g., computing component 310) can perform inter-memory block matrix operations (i.e., matrix operations between weights classified into a first memory block / stored in memory hardware segment 326 and weights classified into a second memory block / stored in memory hardware segment 336). This organized sequence of performing matrix operations can be more efficient (i.e., resulting in less weight movement across computing components) than an alternative method of performing matrix operations involving weights stored across different computing components in a less organized manner. As described above, improving matrix operation efficiency can improve processing time and, in some cases, reduce power consumption and monetary cost.

[0062] Figure 4 1 shows an example computing component 410 according to various examples of the disclosed technology, which can be used to dynamically adjust the amount of precision used to represent the individual weights of the neural network in response to training. In some examples, the computing component 410 can be in the computing system 102 of FIG. Figure 2 The computing component 200, Figure 3 The computing component 310 and / or Figure 3 is implemented on a computer system 300.

[0063] Reference now Figure 4 , computing component 410 can be, for example, a server computer, a controller, or any other similar computing component capable of processing data. Figure 4 In an example implementation of , computing component 410 includes a hardware processor 412 and a machine-readable storage medium 414 .

[0064] Hardware processor 412 may be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for retrieving and executing instructions stored in machine-readable storage medium 414. Hardware processor 412 may fetch, decode, and execute instructions, such as instructions 416-426. As an alternative or in addition to retrieving and executing instructions, hardware processor 412 may include one or more electronic circuits including electronic components for performing the functions of one or more instructions, such as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or other electronic circuits.

[0065] A machine-readable storage medium, such as machine-readable storage medium 414, may be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. Thus, machine-readable storage medium 414 may be, for example, a random access memory (RAM), a non-volatile RAM (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disk, etc. In some examples, machine-readable storage medium 414 may be a non-transitory storage medium, where the term “non-transitory” does not include a transient propagation indicator. As described in detail below, machine-readable storage medium 414 may be encoded with executable instructions, such as instructions 416-426. Furthermore, although Figure 4 The instructions are shown in order, but the order shown is not the only order in which the instructions may be executed. Any instructions may be executed in any order, at any time, repeatedly, and / or by any suitable device or devices.

[0066] Hardware processor 412 may execute instructions 416 to first sort the weights of the neural network model into the first memory block and the second memory block according to a heuristic method. In various implementations, the first sorting may be responsive to completion of a first set of training iterations of the neural network model. As described above, the heuristic method may include at least one of: (1) a measure that quantifies the magnitude of the effect of the corresponding weight on the output of the neural network model during the most recent set of training iterations; and (2) a measure that quantifies the magnitude of fluctuations in the value of the corresponding weight during the most recent set of training iterations.

[0067] The hardware processor 412 may execute instructions 418 to use a first number of memory bits to first store the individual weights that are first sorted into the first memory block of weights. Relatedly, the hardware processor 412 may execute instructions 420 to use a second number of memory bits to first store the individual weights that are first sorted into the second memory block of weights, wherein the first number of memory bits is less than the second number of memory bits.

[0068] As described above, using a first number of memory bits to first store individual weights that are first classified into a first memory block may include first storing the individual weights that are first classified into the first memory block in individual sub-segments of a first memory hardware segment, wherein the corresponding sub-segments of the first memory hardware segment include the first number of memory bits. Relatedly, using a second number of memory bits to first store individual weights that are first classified into a second memory block includes first storing the individual weights that are first classified into the second memory block in individual sub-segments of a second memory hardware segment, wherein the corresponding sub-segments of the second memory hardware segment include the second number of memory bits. In some implementations, the first memory hardware segment may be implemented on a first streaming multiprocessor (SM) of a general purpose processing unit (GPU), and the second memory hardware segment may be implemented on a second SM of the GPU. In other implementations, the first memory hardware segment may be implemented on a first GPU, and the second memory hardware segment may be implemented on a second GPU.

[0069] The hardware processor 412 may execute instructions 422 to secondarily sort the first stored weights of the neural network model into the first memory block and the second memory block according to the heuristic method. In various implementations, the second sorting may be in response to completion of a second set of training iterations of the neural network model having the first stored weights.

[0070] The hardware processor 412 may execute instructions 424 to use a first number of memory bits to secondarily store the individual weights of the weights secondarily sorted into the first memory block. Relatedly, the hardware processor 412 may execute instructions 426 to use a second number of memory bits to secondarily store the individual weights of the weights secondarily sorted into the second memory block.

[0071] In some implementations, the hardware processor 412 may execute additional instructions to, during a second set of training iterations: (a) perform intra-memory block matrix operations between weights first classified into the first memory block; (b) perform intra-memory block matrix operations between weights first classified into the second memory block; and (c) in response to completion of the intra-memory block matrix operations, perform inter-memory block matrix operations between weights first classified into the first memory block and weights first classified into the second memory block.

[0072] In some implementations, prior to a first set of training iterations for the neural network model, the hardware processor 412 may execute additional instructions to: (a) initially sort all weights of the neural network model into a first memory block; and (b) use a first number of memory bits to initially store individual weights of the weights initially sorted into the first memory block, so that the neural network model has initially stored weights during the first set of training iterations.

[0073] In various implementations, first classifying may further include first classifying the weights of the neural network model into a third memory block according to a heuristic method. In these implementations, first storing may further include first storing individual weights among the weights first classified into the third memory block using a third number of memory bits, wherein the second number of memory bits is less than the third number of memory bits. Relatedly, second classifying may further include second classifying the weights of the neural network model into the third memory block according to a heuristic method. Here, second storing may further include second storing individual weights among the weights second classified into the third memory block using a third number of memory bits.

[0074] Figure 5 An example flowchart 500 is shown that can be used to dynamically adjust the amount of precision used to represent individual weights of a neural network in response to training, according to one or more examples.

[0075] As shown, flowchart 500 includes operations 516-526. Figure 4 For the sake of brevity, the instructions 416-426 will not be described again. Operations 516-526 may be performed by the computing system 102 of FIG. 1, Figure 2 The computing component 200, Figure 3 The computing component 310, Figure 3 The computing system 300 and Figure 5 Any one of the computing components 410 is executed.

[0076] Each of the processes, methods, and algorithms described in the preceding sections may be embodied in a code component executed by one or more computer systems or computer processors including computer hardware, and fully or partially automated by it. One or more computer systems or computer processors may also operate to support the execution of related operations in a "cloud computing" environment or as "software as a service" (SaaS). These processes and algorithms may be implemented in part or in whole in a dedicated circuit system. The various features and processes described above may be used independently of each other, or may be combined in various ways. Different combinations and sub-combinations are intended to fall within the scope of the present disclosure, and in some implementations, certain methods or process blocks may be omitted. The methods and processes described herein are also not limited to any particular sequence, and the blocks or states associated therewith may be executed in other appropriate sequences, or may be executed in parallel, or in some other manner. Blocks or states may be added to or removed from the disclosed example embodiments. The execution of certain operations or processes may be distributed between computer systems or computer processors, not only residing in a single machine, but also deployed on multiple machines.

[0077] As used herein, circuits can be implemented using any form of hardware, software or a combination thereof. For example, one or more processors, controllers, ASIC, PLA, PAL, CPLD, FPGA, logic components, software routines or other mechanisms can be implemented to form circuits. In implementation, the various circuits described herein can be implemented as discrete circuits, or the described functions and features can be partially or completely shared between one or more circuits. Even if various features or functional elements can be individually described or claimed as separate circuits, these features and functions can be shared between one or more common circuits, and this description should not require or imply that separate circuits are needed to implement such features or functions. In the case where the circuit is implemented in whole or in part using software, such software can be implemented as operating together with a computing or processing system that can perform the functions described therein.

[0078] As used herein, the term "or" may be interpreted as inclusive or exclusive. In addition, singular resource, operation, or structure descriptions should not exclude the plural form. Unless otherwise specifically stated, or otherwise understood in the context of use, conditional language such as "can", "could", "might", or "may" is generally intended to convey that certain embodiments include and other embodiments do not include certain features, elements, and / or steps.

[0079] Unless expressly stated otherwise, the terms and phrases used in this document and variations thereof should be interpreted as open ended, not restrictive. Adjectives such as "conventional," "traditional," "normal," "standard," "known," and terms of similar meaning should not be interpreted as limiting the items described to items available in a given time period or as of a given time, but should be understood to include conventional, traditional, normal, or standard technology that may be available or known at any time now or in the future. In certain cases, the appearance of broadening words and phrases such as "one or more," "at least," "but not limited to," or other similar phrases should not be understood as an intent or need to use a narrower case in the absence of such broadening phrases.

Claims

1. A method comprising: According to a heuristic method, weights of the neural network model are first classified into a first memory block and a second memory block; using a first number of memory bits to first store individual weights of the weights first sorted into the first memory block; using a second number of memory bits to first store individual weights that were first sorted into the weights of the second memory block, wherein the first number of memory bits is smaller than the second number of memory bits; sorting the first stored weights of the neural network model secondarily into the first memory block and the second memory block according to the heuristic method; using said first number of memory bits to secondarily store individual weights secondarily sorted into said weights of said first memory block; as well as The second number of memory bits is used to secondarily store individual weights secondarily sorted into the weights of the second memory block.

2. The method according to claim 1, wherein: The first classifying is responsive to completion of a first set of training iterations for the neural network model; and The second classification is responsive to completion of a second set of training iterations for the neural network model.

3. The method according to claim 1, wherein the heuristic method comprises at least one of the following: a measure that quantifies the magnitude of the effect of the corresponding weight on the output of the neural network model during a most recent set of training iterations; and A measure that quantifies how much the value of the corresponding weight fluctuated during the most recent set of training iterations.

4. The method of claim 2, further comprising, during the second set of training iterations: performing intra-memory block matrix operations between the weights first sorted into the first memory block; performing intra-memory block matrix operations between the weights first sorted into the second memory block; and In response to completion of the intra-memory block matrix operation, an inter-memory block matrix operation is performed between the weights first sorted into the first memory block and the weights first sorted into the second memory block.

5. The method according to claim 1, wherein: Using the first number of memory bits to first store the individual weights first classified into the first memory block comprises: first storing the individual weights first classified into the first memory block in individual sub-segments of a first memory hardware segment, wherein the corresponding sub-segments of the first memory hardware segment include the first number of memory bits; and Using the second number of memory bits to first store the individual weights that are first classified into the second memory block includes: first storing the individual weights that are first classified into the second memory block in individual sub-segments of a second memory hardware segment, wherein the corresponding sub-segments of the second memory hardware segment include the second number of memory bits.

6. The method according to claim 5, wherein: The first memory hardware segment is implemented on a first streaming multiprocessor SM of a processing unit; and The second memory hardware segment is implemented on a second SM of the processing unit.

7. The method according to claim 2, further comprising: Prior to the first set of training iterations for the neural network model, initially sorting all weights of the neural network model into the first memory block; as well as Individual weights that are initially classified into the weights of the first memory block are initially stored using the first number of memory bits so that the neural network model has the weights initially stored during the first set of training iterations.

8. The method according to claim 1, wherein: The first classification further comprises: first classifying the weights of the neural network model into a third memory block according to the heuristic method; The first storing further comprises: using a third number of memory bits to first store the individual weights of the weights first classified into the third memory block, wherein the second number of memory bits is smaller than the third number of memory bits; The secondarily classifying further comprises: secondarily classifying the weights of the neural network model into the third memory block according to the heuristic method; and The secondarily storing further comprises: secondarily storing the individual weights secondarily sorted into the weights of the third memory block using the third number of memory bits.

9. A system comprising: a first memory hardware segment; a second memory hardware segment; as well as one or more processors operative to execute machine-readable instructions to: According to a heuristic method, weights of the neural network model are first classified into a first memory block and a second memory block; In said first memory hardware segment, using a first number of memory bits to first store individual weights first sorted into said weights of said first memory block; in the second memory hardware segment, using a second number of memory bits to first store individual weights that were first sorted into the weights of the second memory block, wherein the first number of memory bits is smaller than the second number of memory bits; sorting the first stored weights of the neural network model secondarily into the first memory block and the second memory block according to the heuristic method; in said first memory hardware segment, using said first number of memory bits to secondarily store individual weights secondarily sorted into said weights of said first memory block; as well as In the second memory hardware segment, the second number of memory bits is used to secondarily store individual weights secondarily sorted into the weights of the second memory block.

10. The system of claim 9, wherein: The first classifying is responsive to completion of a first set of training iterations for the neural network model; and The second classification is responsive to completion of a second set of training iterations for the neural network model.

11. The system of claim 9, wherein the heuristic method comprises at least one of the following: a measure that quantifies the magnitude of the effect of the corresponding weight on the output of the neural network model during a most recent set of training iterations; and A measure that quantifies how much the value of the corresponding weight fluctuated during the most recent set of training iterations.

12. The system of claim 10, wherein the one or more processors are further operative to execute machine-readable instructions to, during the second set of training iterations: performing intra-memory block matrix operations between the weights first sorted into the first memory block; performing intra-memory block matrix operations between the weights first sorted into the second memory block; and In response to completion of the intra-memory block matrix operation, an inter-memory block matrix operation is performed between the weights first sorted into the first memory block and the weights first sorted into the second memory block.

13. The system of claim 9, wherein: In the first memory hardware segment, first storing the individual weights first classified into the weights of the first memory block comprises: first storing the individual weights first classified into the weights of the first memory block in individual sub-segments of the first memory hardware segment, wherein the respective sub-segments of the first memory hardware segment include the first number of memory bits; and In the second memory hardware segment, first storing the individual weights that are first classified into the weights of the second memory block includes: in an individual sub-segment of the second memory hardware segment, first storing the individual weights that are first classified into the weights of the second memory block, wherein the corresponding sub-segment of the second memory hardware segment includes the second number of memory bits.

14. The system of claim 9, further comprising a processing unit, wherein: The first memory hardware segment is implemented on a first streaming multiprocessor SM of the processing unit; and The second memory hardware segment is implemented on a second SM of the processing unit.

15. The system according to claim 9, further comprising a first processing unit and a second processing unit, wherein: The first memory hardware segment is implemented on the first processing unit; and The second memory hardware segment is implemented on the second processing unit.

16. The system of claim 10, wherein the one or more processors are further operative to execute machine-readable instructions to: Prior to the first set of training iterations of the neural network model, initially sorting all weights of the neural network model into the first memory block; and In the first memory hardware segment, the individual weights initially classified into the weights of the first memory block are initially stored using the first number of memory bits so that the neural network model has the weights initially stored during the first set of training iterations.

17. A non-transitory computer readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: In response to completion of a first set of training iterations of the neural network model, weights of the neural network model are first sorted into a first memory block and a second memory block according to a heuristic method; using a first number of memory bits to first store individual weights of the weights first sorted into the first memory block; using a second number of memory bits to first store individual weights that were first sorted into the weights of the second memory block, wherein the first number of memory bits is smaller than the second number of memory bits; responsive to completion of a second set of training iterations of the neural network model having the first stored weights, secondarily sorting the first stored weights of the neural network model into the first memory block and the second memory block according to the heuristic method; using said first number of memory bits to secondarily store individual weights secondarily sorted into said weights of said first memory block; as well as The second number of memory bits is used to secondarily store individual weights secondarily sorted into the weights of the second memory block.

18. The non-transitory computer readable medium storing instructions of claim 17, wherein the heuristic method includes at least one of the following: a measure that quantifies the magnitude of the effect of the corresponding weight on the output of the neural network model during a most recent set of training iterations; and A measure that quantifies how much the value of the corresponding weight fluctuated during the most recent set of training iterations.

19. The non-transitory computer readable medium storing instructions of claim 17, further comprising instructions for, during the second set of training iterations: performing intra-memory block matrix operations between the weights first sorted into the first memory block; performing intra-memory block matrix operations between the weights first sorted into the second memory block; and In response to completion of the intra-memory block matrix operation, an inter-memory block matrix operation is performed between the weights first sorted into the first memory block and the weights first sorted into the second memory block.

20. The non-transitory computer readable medium storing instructions of claim 15, wherein: First storing the individual weights that are first sorted into the weights of the first memory block comprises: first storing the individual weights that are first sorted into the weights of the first memory block in individual physical sub-segments of a first memory hardware physical segment, wherein the corresponding physical sub-segment of the first memory hardware physical segment comprises the first number of memory bits; and First storing the individual weights that are first classified into the weights of the second memory block includes: first storing the individual weights that are first classified into the weights of the second memory block in individual physical sub-segments of a second memory hardware physical segment, wherein the corresponding physical sub-segment of the second memory hardware physical segment includes the second number of memory bits.