STORAGE METHOD FOR MIXED PRECISION WEIGHTS IN NEURAL NETWORKS
Dynamic shift registers adapt to non-constant weight bit sizes, enhancing neural network weight storage and access efficiency, thereby accelerating inference and reducing energy consumption.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-07
- Publication Date
- 2026-04-09
AI Technical Summary
Existing neural network weight storage and access architectures are inefficient, leading to suboptimal inference speed and high energy consumption, particularly in battery-powered devices.
A method involving dynamic shift registers that adjust to non-constant weight bit sizes, allowing efficient storage and reconstruction of weights in a target memory with improved access speed and reduced energy consumption.
Accelerates neural network inference speed and reduces energy consumption by optimizing weight data access in internal memory, enabling faster and more efficient processing.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present invention relates to a method for improving the data storage and access architecture.
[0002] US2018082181 discloses a reordering, weight compression, and processing procedure for a neural network. In the method disclosed in US'181, a neural network is trained to generate feature maps and associated weights. Reordering is performed to generate a functionally equivalent network. This reordering can be performed to improve weight compression, load balancing, and / or execution. In one implementation, zero-value weights are grouped so that they can be skipped during execution.
[0003] The present invention aims to solve at least some of the problems and disadvantages mentioned above.
[0004] The object of the present invention is to provide a method for processing a plurality of weights with a weight bit size of a layer of an artificial neural network in a target memory with a memory width or a method for reconstructing weights of a neural network from a target memory with improved characteristics.
[0005] This problem is solved by a method according to claim 1 and a method according to claim 14.
[0006] A specific preferred embodiment relates to an invention according to claim 11. In this embodiment, at least one of a plurality of shift registers used to store and read weight data is a dynamic shift register configured to change the number of bits it stores according to the bit size of the weight currently being processed by the output processing pipeline. This enables the processing of weights with a non-constant bit size.
[0007] The following description of the figures of specific embodiments of the invention is merely exemplary and is not intended to limit the present teachings, their application, or uses. In all drawings, corresponding reference numerals indicate identical or corresponding parts and features. They show: Fig. 1. A weight storage for 8-bit weights in a 32-bit wide memory; Fig. 2 an example of 6-bit weights, where each weight is padded to eight bits; Fig. 3 a memory in which multiple weights are stored in a single memory word; and Fig. 4 how the weights are read from memory and reconstructed in a processing unit.
[0008] The present invention relates to a method for processing a plurality of weights with a weight bit size of a layer of an artificial neural network in a target memory with a memory width. The original location of the weight data can be another memory, probably an external memory, while the target memory is an internal memory equipped with at least one processing pipeline. The method enables the transfer of data, in particular weights of a neural network model, from a low-speed memory to a high-speed memory in such a way that the data are stored in the high-speed target memory in an arrangement that makes simultaneous access to a plurality of weights both faster and more efficient. This increase in the access speed to weight data has the advantage that the inference orThe inference speed of the neural network is greatly accelerated, while the energy consumption of any device in which the method according to the present invention is applied is reduced.
[0009] Unless otherwise defined, all terms used in disclosing the invention, including technical and scientific terms, have the meanings that would normally be understood by a person skilled in the art in the field to which this invention belongs. Further definitions are provided to facilitate understanding of the teaching of the present invention.
[0010] As used herein, the following terms have the following meanings: “Einer / eine / eines” and “der / die / das”, as used herein, refer to both singular and plural cases unless the context clearly indicates otherwise. For example, “ein Fach” refers to one or more subjects.
[0011] “Approximately,” as used herein with reference to a measurable value such as a parameter, a quantity, a duration, and the like, shall include variations of ±20% or less, preferably ±10% or less, more preferably ±5% or less, more preferably ±1% or less, and yet more preferably ±0.1% or less of the specified value, insofar as such variations are suitable to be carried out in the disclosed invention. It should be noted, however, that the value to which the modifier “approximately” refers is itself also specifically disclosed.
[0012] “Exhibiting”, “exhibiting”, “indicates” and “consists of”, as used herein, are synonymous with “comprising”, “comprising”, “encompassing”, or “containing”, “containing”, “includes”, and are inclusive or open terms that specify the presence of the following component and do not exclude or prevent the presence of additional, unspecified components, features, elements, elements, or steps known in the field or disclosed therein.
[0013] Furthermore, the terms first, second, third, and the like are used in the description and in the claims to distinguish between similar elements and not necessarily to describe a sequential or chronological order, unless specified. It is pointed out that the terms used in this way are interchangeable under appropriate circumstances and that the embodiments of the invention described herein may be operated in orders other than those described or illustrated herein.
[0014] The naming of numerical ranges by endpoints includes all numbers and fractions that are grouped within this range, as well as the named endpoints.
[0015] While the terms “one or more” or “at least one”, such as one or more or at least one element(s) of a group of elements, are clear in themselves, the term, with further explanation, includes, among other things, a reference to one of the elements or to two or more of the elements, such as any ≥3, ≥4, ≥5, ≥6 or ≥7 etc. of the elements, and up to all elements.
[0016] Unless otherwise defined, all terms used in disclosing the invention, including technical and scientific terms, have the meanings that would normally be understood by a person skilled in the art in the field to which the invention belongs. Further guidance is provided by way of definition for the terms used in the description to aid in understanding the teaching of the present invention. The terms or definitions used herein are provided only to assist in understanding the invention.
[0017] The reference in this description to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearance of the phrase "in an embodiment" at various points in this description does not necessarily always refer to the same embodiment, but may. Furthermore, the particular features, structures, or characteristics in one or more embodiments may be combined in any suitable manner, as is apparent to a person skilled in the art from this disclosure.While some embodiments described herein contain some, but not different, features that are also present in other embodiments, combinations of features from different embodiments are intended to fall within the scope of the invention and to form different embodiments, as understood by those skilled in the art. For example, any of the claimed embodiments may be used in any combination in the following claims.
[0018] In a first aspect, the invention provides a method for processing a plurality of weights with a weight bit size of a layer of an artificial neural network in a target memory with a memory width, wherein the method comprises the step of processing the weights and the step of storing the processed weights in the target memory. The weights are obtained by first training the artificial neural network. Preferably, the weights are obtained by sufficiently training the neural network such that an inference performed using the neural network yields a high degree of reliability, wherein the reliability is preferably above 70%, more preferably above 75%, 80%, 85%, 90%, 95%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, and most preferably above 99.7%. The target memory is preferably an internal memory.The source storage can be any type of storage, such as internal storage, but preferably external storage. This allows a larger external storage device to provide the weights, which are then stored in a smaller internal storage device. This enables higher bandwidth access to the weights in the internal storage, and therefore faster access than with direct access to external storage.
[0019] Artificial neural networks (ANNs, also shortened to neural networks (NNs)) are a branch of machine learning models built using principles of neural organization. An ANN is based on a collection of connected units or nodes called artificial neurons, which loosely model the neurons in a biological brain. Each connection, like the synapses in a biological brain, can transmit a signal to other neurons. An artificial neuron receives signals, processes them, and can then signal connected neurons. The "signal" at a connection is a real number, and the output of each neuron is calculated as a nonlinear function of the sum of its inputs. The connections are called edges. Neurons and edges typically have a weight that adjusts as learning progresses. The weight increases or decreases the strength of the signal at a connection.Neurons can have a threshold, so that a signal is only sent if the total signal exceeds this threshold.
[0020] To facilitate the implementation of a neural network in an integrated circuit, preferably an application-specific integrated circuit, the weight of the neural network can be quantized. This advantageously reduces the memory requirements of a device used for performing inference, which further advantageously reduces the device's power consumption. This is particularly advantageous when the inference is to be performed by portable devices, where, in these typically battery-powered devices, energy savings translate into longer operating time before recharging.
[0021] In this context, inference, or deep learning inference, is the development phase in which the skills acquired during the training of a neural network are put into practice. The trained neural networks make predictions or inferences about new or novel data that the model has never seen before.
[0022] In this context, quantization is the process of reducing the precision of weights, biases, and activations so that they consume less memory.
[0023] The processing step of the weights comprises the following step: dividing each weight into at least two separate weight discs with a fixed bit size for each weight, wherein the fixed bit sizes are preferably the same for each weight disc, each having one or more bits and together defining the weight.
[0024] The step of storing the processed weights involves sequentially storing each nth separate weight disc in the target memory. Here, n ranges from 1 to the number of weight discs per weight, so that each word in the target memory contains only the mth weight discs, where m is a number between 1 and the number of weight discs per weight. This allows for more complete utilization of the available memory and bus when reading the weights.
[0025] In this context, the term "bus," shortened from the Latin "omnibus" and historically also referred to as a data highway or data bus, is understood as a communication system that transmits data between components within a computer or between computers. This term encompasses all associated hardware components, such as wire, optical fiber, etc., and software, including communication protocols.
[0026] In this context, the term processing pipeline or data pipeline is to be understood as a set of processes and tools that move data from one system to another, often involving stages of collection, processing, storage and analysis.
[0027] In one embodiment, the weight discs are stored such that the target memory has a plurality of consecutive words with a width equal to the memory width, wherein the nth word only contains nth weight discs of the weights, where n ranges from 1 to the number of weight discs per weight.Preferably, the nth weight discs are stored in the nth word until the word is filled with nth weight discs, after which successive nth weight discs are stored in the n+K-th word, where K is the number of weight discs per weight. This process is repeated by storing successive nth weight discs in the n+K·(p+1)-th word after filling the n+K·p-th word, where p is a natural number ranging from 1 to a value for which n+K·(p+1) lies between [(N·B) / W]-1 and [(NB) / W], where N is the total number of weights, B is the fixed bit size, and W is the memory width. This advantageously allows full utilization of the available internal memory and processing pipelines, which are exposed via the bus connecting the internal memory to the processing unit that uses the weights stored in the internal memory.
[0028] In one embodiment, the weight disk bit size is determined based on a number of available input processing pipelines, where the input processing pipelines are configured for the step of processing and / or storing the weights. This ensures that a maximum number of weights can be written simultaneously. In this way, the speed at which weights can be accessed and written from the source memory to the destination memory is advantageously increased.
[0029] In one embodiment, the weight disc bit size is determined based on the memory width and / or the weight bit size. In this embodiment, the term "memory" refers to the target memory. Preferably, the weight processing and storage steps are performed by a number of input processing pipelines, the number of which is set to equal the width of the target memory. This advantageously maximizes the number of weights that can be accessed simultaneously from the target memory, further increasing the speed at which a processing unit can access the weights.
[0030] In one embodiment, each output pipeline includes a shift register. Preferably, each shift register is set to a bit value equal to the bit size of the weight assigned to its corresponding output processing pipeline. In this way, multiple weights can be assigned to the same processing pipeline, since the shift register first collects all the bits of each weight before transferring the bits to a processing unit and starting the collection of bits for a subsequent weight.
[0031] In one embodiment, at least one of the shift registers is a dynamic shift register configured to change the number of bits it stores according to the bit size of the weight currently being processed by the output processing pipeline. Preferably, each weight is stored in the destination memory along with its corresponding shift register value. This allows weights with non-constant bit sizes to still be processed.
[0032] In one embodiment, each input processing pipeline stores a bit disk of its assigned weight in the destination memory during each clock cycle. Preferably, each time the number of clock cycles executed by the processing pipeline reaches the set value of the shift register assigned to its output pipeline, a new weight is assigned to an input processing pipeline.
[0033] In one embodiment, the shift register is sized according to the largest weight bit size supported by the processing pipeline, with each weight being padded with a bit size smaller than the largest weight bit size before being stored in the destination memory. This ensures that all weight bit sizes can be accommodated by the shift register, thus preventing any processing pipeline breakage errors.
[0034] A second aspect of the invention provides a method for reconstructing weights of a neural network from a target memory, wherein the weights have a known weight bit size, the target memory has a known memory width and a known number of words, and wherein each of the weights has been stored in at least two separate weight disks with a known disk bit size for each weight disk, the weights preferably being stored using the method according to claim 1, wherein the method for reconstructing weights comprises the following steps: - in each nth word of the target memory, subdividing the nth word into separate, consecutive word disks with a bit size equal to the nth weight disk bit size, repeating the step for n, which runs from 1 to the number of words in steps of 1; - Reconstructing each m-th weight by concatenating the m-th word disk of each word in the target memory to which the m-th weight is assigned, repeating the step for m, which runs from 1 to the number of weights in steps of 1.
[0035] In one embodiment, the reconstruction of each m-th weight is achieved by concatenating the m-th word disk of K consecutive words, where K is the number of weight disks per weight, and where the last of the K consecutive words is the K·p-th word, where p is a natural number such that K·p is at most the total number of words in the target memory. Preferably, a plurality of shift registers are used to perform the concatenation. More preferably, the shift registers are dynamic shift registers. This advantageously allows the reconstruction of weights with a fixed bit size, but also of weights with variable bit sizes.
[0036] In one embodiment, each shift register has a serial output. This advantageously reduces wiring complexity and thus the risk of errors on printed circuit boards. Furthermore, serial output shift registers typically have a smaller footprint, allowing their use in smaller devices.
[0037] In one embodiment, each shift register has parallel outputs, with the number of outputs being the same as the size of the shift register. This advantageously enables high-speed data transmission.
[0038] In one embodiment, each shift register is a universal shift register. In this context, a universal shift register is understood to be a shift register that can perform input-output operations in both serial and parallel modes. This allows the use of the capabilities of both parallel-output and serial-output shift registers, as well as enabling communication with devices and / or elements that require either serial or parallel input.
[0039] However, it is obvious that the invention is not limited to this application. The method according to the invention can be applied in all types of devices that require high-speed memory access.
[0040] The invention is further described by the following non-limiting examples, which further illustrate the invention and are not intended to limit its scope, nor should they be interpreted as such.
[0041] The present invention will now be described in more detail with reference to examples which are not limiting.
[0042] With the aim of better illustrating the features of the invention, the following, by way of example and in no way limiting other potential applications, presents a description of a number of preferred applications of the method for processing a plurality of weights based on the invention, wherein: Fig. Figure 1 shows a weight storage for 8-bit weights (4) in a 32-bit wide memory (1). In neural networks, weights are typically used in blocks. Fig. Figure 1 shows an example of how the weights (4) can be stored in the memory (1). In this example, weights are stored efficiently: there are no gaps in memory utilization, and a minimal set of word reads is required to retrieve the full set of 64 weights. The memory (1) is divided into eight-bit words ordered by means of a plurality of word addresses (3), each address comprising thirty-two bits (2) grouped into four words of eight bits each. Typically, a processing unit (6, not shown) reads the weights from the memory. The weights, retrieved in parallel from each of the four data words, as if seeded into a data word, are fed into separate processing pipelines (8). This is further illustrated in Figure 1. Fig. 4 illustrates. Fig. Figure 2 shows an example of 6-bit weights (4), where each weight is padded to eight bits. These eight bits correspond to the size of each data word in which each weight (4) is stored. If the size of the encoded weights (4) is not a divisor of the width of the memory (1), efficient storage becomes less trivial. This is due to the discrepancy between the word size and the bit size of each weight (4). Fig. Figure 3 shows a memory (1) containing multiple weights (4) in a single memory word. Each word comprises one of a set of thirty-two weights, with each group of thirty-two weights corresponding to a block (9, 10). For each block (9, 10), the figure shows the first word containing the first bit of each of the thirty-two weights (4), and each subsequent word of the same block (9 or 10) containing a subsequent bit of the same thirty-two weights (4). Thus, the required number of words corresponds to the size of the weight (4) in bits. The figure also shows a second block (10) containing another set of thirty-two weights (4), stored in eight words of thirty-two bits each. Fig.Figure 4 shows how the weights (4) are read from memory (1) and reconstructed in a processing unit (6). The width of memory (1) is set equal to the number of processing pipelines (8), which in this example is thirty-two. Each processing pipeline is preceded by a shift register (not shown) connected to a single bit of the data coming from memory (1). As data comes from memory (1), it is shifted into the respective shift registers. When all the bits of a set of weights (4) are read in the correct order, each shift register contains all the bits belonging to a single weight value (4). The shift register itself is sized according to the largest weight size (4) supported by the processing pipeline.
[0043] Therefore, before shifting begins, the shift register must be initialized to have a known value for the bits that will not be updated by the shift. Adjusting to a different weight size simply involves loading a different counter value into the state machine, which generates the memory read accesses and decides when all bits are shifted and can be passed to the processing pipeline (8).
[0044] The present invention is in no way limited to the given examples or to the embodiments illustrated in the figures. On the contrary, methods according to the present invention can be implemented in many different ways without deviating from the scope of protection of the invention.
[0045] It is assumed that the present invention is not limited to any previously described form of implementation and that some modifications can be added to the illustrated example without re-evaluating the appended claims. For example, the present invention has been described with reference to weights of artificial neural networks, but it is clear that the invention can be applied to other forms of data that, for example, require efficient storage and retrieval. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 2018082181
[0002]
Citation Information
Patent Citations
Neural Network Reordering, Weight Compression, and Processing
US20180082181A1