Dynamic buffers and how to allocate them

JP2025510483A5Pending Publication Date: 2026-04-08AIMOTIVE KFT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-27
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

Existing memory devices face challenges in optimizing power consumption, as reducing memory size decreases power consumption but limits capacity, while increasing memory size increases power consumption.

Method used

The method involves dynamic buffer width allocation, where memory is divided into multiple slices, and the buffer width is determined based on the requirements of arithmetic operations, such as word length, number of results, and statistical parameters, allowing for selective activation and deactivation of memory slices to minimize power usage.

Benefits of technology

This approach dynamically adjusts the buffer size to match the current arithmetic operations, reducing power consumption by only using the necessary memory slices while ensuring sufficient capacity for a wide range of operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method for dynamic buffer width allocation, comprising: providing a memory with a plurality of memory slices; determining a buffer width for a set of arithmetic operations; allocating some of the plurality of memory slices to form a buffer having at least the determined buffer width; and performing the set of arithmetic operations using the buffer. Additionally, a dynamic memory and an apparatus having a plurality of memory slices are described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to memory management, dynamic buffers, dynamic buffer allocation, and dynamic buffer width allocation. In particular, the present disclosure relates to a method and respective memory for dynamic buffer width allocation. [Background technology]

[0002] Memory or memory devices are well known in the art as electronic components suitable for storing data. Memory devices are used to store information for immediate use in electronic devices. Memory can store computer program operations or instructions, or data values ​​used by such operations or instructions. Memory devices are typically implemented as semiconductor memories, with data stored in memory cells constructed from transistors and other components on an integrated circuit. However, other physical memory architecture memory devices can include volatile and / or non-volatile memory. Examples of volatile memory include dynamic random access memory (DRAM), which can be used for primary storage, and static random access memory (SRAM), which can be used as a CPU cache, for example. Various types of memory can be used in personal computers, workstations, routers, peripherals, hard disks, routers, or dedicated hardware such as graphics hardware or AI accelerators.

[0003] Memory devices are typically organized into memory cells. The memory cells or other memory units may be grouped into words of a fixed word length, for example, 1, 2, 4, 8, 16, 32, 64, 128 bits, etc. Thus, each memory cell or memory unit may have a width that may be sufficient to accommodate a word of the fixed word length, which may include control data, such as error correction codes. Each stored word may be accessed by a binary address of several bits, allowing multiple words to be stored in the memory device. The memory cells or memory units may also be arranged as one or more arrays, which may be organized, for example, into rows and columns. A certain number of memory rows or columns may be used to store corresponding bits of data. Other semiconductor memory devices exist that may use other groupings of memory cells and corresponding addressing schemes.

[0004] Memory devices typically have a fixed storage capacity that determines the number and size of data that can be stored in the memory device. This is sometimes called the memory width. Electronic devices are usually equipped with a memory that has a sufficient capacity or width. This capacity is initially determined to accommodate the expected number of words of a desired length. After implementation, the memory capacity of the electronic device is fixed.

[0005] A buffer or data buffer typically refers to an area of ​​a memory device used for temporary storage of data. For example, a buffer may be used to store the results of an operation and provide input for such an operation. If a buffer uses the entire memory or memory device, the term buffer may be used to refer to the memory or memory device. After setup, a buffer requires the memory device in which it is located to be in full operation to store and / or retrieve data. A buffer may have a width to indicate the number of bits that can be stored in the buffer.

[0006] Memory devices consume power during operation, which may involve the power cost to refresh individual memory cells or memory units, or the power cost to access the memory, which may be proportional to the size of the memory and the number of reads or writes to the memory device. Summary of the Invention

[0007] For many electronic devices, power consumption is a major factor in product design. Reducing the size of memory can reduce power consumption. However, this can lead to insufficient memory capacity that can limit the operating capabilities of the electronic device. On the other hand, increasing the size of memory can increase power consumption.

[0008] Therefore, there is a need in the art to further optimize power consumption in memory devices.

[0009] This problem is solved by a method, at least one computer readable medium, an apparatus and a memory device as defined in the independent claims. Preferred embodiments are defined in the corresponding dependent claims.

[0010] The present disclosure may provide an approach for dynamic memory or buffer allocation. More specifically, the present disclosure may provide a technique for dynamic memory or buffer width allocation.

[0011] A first aspect of the present disclosure provides a method for buffer allocation. More specifically, the method may be for dynamic buffer width allocation. The method includes providing a memory having a plurality of memory slices, determining a buffer width for a set of arithmetic operations, allocating some of the memory slices, forming a buffer in which the memory slices have at least the determined buffer width, and performing the set of arithmetic operations using the buffer.

[0012] The memory may be configured as several memory cells or units, each capable of storing one bit, several bits, or one or several words of a particular length. One or more of the memory cells or units may be denoted as a memory slice. Thus, the memory may be organized or arranged as a plurality of memory slices. Each slice may have a slice width indicating the number of bits that may be stored in the respective slice. For example, a memory slice may have a width of 1 and store one bit. This may represent a bitwise slicing of the memory. In another example, a memory slice may have a width of 2, 4, 8, 16, 32, 64 bits, etc. Additionally, memory slices with widths that are not powers of 2 may be used. The slice width may be the same for all slices in the plurality of slices. The slice width may be different for at least one of the plurality of slices. Thus, the sum of the widths of the allocated memory slices corresponds to the buffer width.

[0013] The arithmetic operation is used to determine the required buffer width. The type of arithmetic operation, the processed data, and the result can be used to estimate the required buffer width. This is done dynamically and adapts to each new set of arithmetic operations to be performed using the buffer.

[0014] As a result, the size of the buffer is adapted to the current arithmetic operation, such that only memory slices corresponding to the determined width of the buffer are used during reading and / or writing, which significantly reduces power costs. Meanwhile, since the buffer size is extended to all memory slices, the buffer can be dynamically adapted to a wide range of arithmetic operations, including those that require large buffer sizes.

[0015] In a preferred embodiment, the memory allows selective activation and deactivation of at least some of the memory slices. The memory may include hardware components capable of controlling and setting the state of at least some of the memory cells or units in which it is located. Each memory slice may be associated with at least one memory cell or unit. Activating or deactivating a memory slice thus involves activating or deactivating the respective associated at least one memory cell or unit. The power consumption of the memory cells or units may differ depending on their state. In the active state, the memory cells or units may consume more power than in the dormant or inactive state. Initially, at least some or all of the memory cells or units may be activated and switched to an operational state. Alternatively, the memory may be initialized by switching all or at least some of the memory cells or units to a dormant or inactive state. Preferably, when activated, individual memory cells or units not used for the memory slices allocated for the buffer may be selectively deactivated and switched to a dormant or inactive state. Optionally, inactive memory cells or units may be selectively activated if the associated memory slice is allocated to the buffer. Selective activation (switching each cell or unit to an active state) and deactivation (switching each cell or unit to a dormant or inactive state) can be performed automatically when allocating or deallocating each memory slice. Selective activation and deactivation of memory slices better adapts memory power consumption to the required buffer width.

[0016] According to a particularly preferred embodiment, the method further comprises deactivating an unallocated memory slice of the plurality of memory slices. The plurality of memory slices may be partitioned into memory slices allocated for the buffer and further memory slices not allocated for the buffer. The former memory slices may be referred to as allocated memory slices. The latter memory slices may be referred to as unallocated memory slices. Thus, the plurality of memory slices may consist of allocated and unallocated memory slices. However, it should be understood that the plurality of memory slices corresponds to the allocated memory slices if all memory slices are allocated. Each of the unallocated memory slices may be associated with at least one memory cell or unit. The associated at least one memory cell or unit of the unallocated memory slice may be deactivated and switched to a dormant or inactive state to reduce the power consumption of the memory in silicon.

[0017] According to another embodiment, the buffer width is determined based on the word length of the values ​​to be processed by the set of arithmetic operations. The arithmetic operations may be analyzed or may include metadata or additional information that may define the word length. A word is a fixed-size piece of data that is treated as a unit by arithmetic operations, such as a processor instruction set or processor hardware instructions. Word length refers to the number of bits in a word, such as 8, 16, 32, 64 bits, etc. However, it should be understood that word length is not limited to powers of 2. Arithmetic operations for special purpose processor designs, such as digital signal processors, may operate with word lengths ranging from, for example, 4 to 80 bits. The buffer width may be determined as a function of the word length as a parameter.

[0018] In yet another embodiment, the buffer width is determined based on the number of results of a set of arithmetic operations. The arithmetic operations may be analyzed or may include metadata or additional information that may indicate the number of results expected to be produced by the arithmetic operations. The number of results may correspond to the number of iterations. The buffer width may be determined as a function with the number of results as a parameter. Preferably, the function may also include a word length as a further parameter. For example, analysis of the arithmetic operations and / or metadata and / or additional information, if available, may indicate that the arithmetic operations are of word length N i We define a binary accumulator that accumulates integers having M, and show that the binary accumulator accumulates up to M such integers. As an example, the required buffer width may be N i +ceil(log2M)+t, where t represents a safety margin.

[0019] According to another embodiment, the buffer width is determined based on statistical parameters related to the input data to be processed by the set of arithmetic operations. The statistical parameters may indicate the input data type and / or distribution characteristics of the input data. The statistical parameters may include any further values ​​describing statistics related to the input data. The buffer width may be determined as a function of the statistical parameters as one or more parameters. Preferably, the function may include a word length as a further parameter. Preferably, the function may include a number of results as a further parameter. Thus, the function may include parameters related to the statistical parameters, the word length, the number of results in any combination. For example, assuming a data type of random value two's complement signed input data, the word length N i For a binary accumulator that accumulates up to M (an integer) of i A value "close" to the target value should be understood in the context of this disclosure as a value that is close to the range of the target value. This can be indicated using a threshold as the distance between the proximate value and the target value being equal to or less than the threshold. Thus, given a safety margin (or threshold) t, the buffer width is N iHowever, for unsigned input data, the M cumulative random integers can be estimated as N i Use the +ceil(log2(M)) bit.

[0020] In a preferred embodiment, the method further includes collecting statistical parameters in real time during the processing of the set of arithmetic operations. The statistical parameters can be based on the statistical parameters available for analysis or an initial set of statistical parameters. The values ​​of the currently processed input data can be monitored and used to modify or update the previous statistical parameters. For example, the monitored values ​​can update a histogram or any other suitable data structure that can be used to estimate the distribution of the input data. During operation, a randomly distributed input data set can turn out to be a data set having a normal distribution. The updated statistical parameters can lead to a different estimation of the buffer width. For the previous example of M cumulative random integers, the total buffer width N i The use of +ceil(log2(M)) can only occur when the number of accumulations are both near maximum and the input data is highly correlated. The buffer width can be adjusted based on the updated statistical parameters.

[0021] In a further embodiment, the method further comprises receiving statistical parameters associated with the input data. Prior to receiving the input data, the statistical parameters may be submitted as metadata or one or more parameters. Preferably, the statistical parameters are pre-calculated for the input data. This may be performed in a dedicated process during pre-processing and / or analysis of the input data. For example, a host computing device may receive the input data, pre-process and / or analyze the input data, and determine statistics associated with the data set. The statistical parameters may be provided to respective processing hardware and / or memory to set up and allocate buffers for processing of the input data by arithmetic operations.

[0022] According to a preferred embodiment, the method further includes monitoring at least one value stored in the buffer, determining that the value exceeds a threshold, and dynamically allocating at least one further memory slice of the plurality of memory slices for the buffer. If the value exceeds the threshold, the at least one further memory slice may be dynamically allocated and assigned to the buffer. The further memory slice may be arranged in the buffer to store the most significant bit or most significant bit (MSB) of the future value. However, it should be understood that the buffer may be dynamically rearranged in a different manner, for example, according to the structure of the new set of allocated memory slices (including the old allocated slice and the further allocated slice) in the hardware. The threshold may represent a maximum or upper limit of values ​​that may be safely stored in the buffer without risk of buffer overflow as the arithmetic operations continue processing. For example, the threshold may be 50%, 60%, 70%, 80%, 90%, or 95% of the maximum value that may be stored in the buffer. It should be understood that any other percentage of the maximum value, such as 67% or 85%, may be used as the threshold. The buffer may also store multiple values, for example, an array of values. In this case, the threshold is adapted to the number of values ​​stored in the array and the values ​​stored in the array in the buffer are monitored and compared to the threshold. For example, if the array stores n integers, the threshold is adapted to t / n, where t denotes the threshold used for the entire buffer. Similar to allocating further memory slices for the buffer, a second threshold may be defined that may be used as a lower limit if the current value exceeds the threshold. If the current value stored in the buffer falls below the second threshold, at least one allocated memory slice may be deallocated from the buffer. The deallocated memory slice may represent a slice that stores the most significant bits or the least significant bits of the values ​​in the buffer. However, similar to the allocation of memory slices, different memory slices may be deallocated based on, for example, hardware requirements.Deallocated memory slices (and underlying memory cells or units) can be selectively deactivated and switched to a dormant or inactive state, allowing for highly dynamic allocation and deallocation of memory slices and further optimization of power consumption.

[0023] In one embodiment, the method may include determining a write side word width and / or monitoring a read side word width. The memory may provide a write port for writing to the memory and / or the respective memory slices of the buffer. The memory may further provide a read port for reading from the memory and / or from the respective memory slices of the buffer. The size of the write data of the write side may be continuously measured in real time to determine the write side word width. Similarly, the size of the read data of the read side may be continuously monitored to keep track of the read side word width. This may be done by dedicated logic circuit units of the write side and the read side that may control the respective write and read ports. The logic circuit units may further communicate with an auxiliary memory to store and / or retrieve configuration parameters and / or store and / or retrieve measured or determined values. Both the read side word width and the write side word width may be used to determine whether to allocate a required memory slice and / or whether to deallocate an unused memory slice of the buffer.

[0024] According to another embodiment, the set of arithmetic operations includes a number of multiply-accumulate operations, the results of which are accumulated in a buffer. The multiply-accumulate operations include operations that allow multiplication of two or more factors, where the products of several input data are accumulated to obtain a final result. These operations are well known in the art in scientific computing, graphics processing, and processing related to, for example, training or inference of neural networks. When these operations are performed on dedicated hardware, such as GPUs or accelerator hardware for neural networks, it is particularly important to provide an appropriately sized accumulator buffer with optimized consumption of power in silicon.

[0025] In a preferred embodiment, the multiply-and-accumulate operations include multiplying the input data by a number of weights, and the buffer width is determined based on the values ​​of the weights and the maximum expected value of the input data. The multiply-and-accumulate operations may be configured to multiply a large amount of input data by a fixed set of weights. For example, the input data may represent different portions of the input channels of a layer of a neural network. Thus, the required buffer width may be estimated based on the values ​​of the fixed set of weights and statistics related to the input data, such as the maximum expected value.

[0026] In a preferred embodiment, the method further includes allocating a further number of memory slices of the plurality of memory slices for a further buffer. The method may further include determining a further buffer width and performing a further set of arithmetic operations using the further buffer. Thus, the memory may be used to provide a plurality of buffers that may be dynamically allocated. Each of the plurality of buffers may be allocated a respective plurality of memory slices of the memory according to the determined buffer width. During the operation of the arithmetic operations, respective values ​​in the buffers may be monitored to determine whether further capacity is required (thus further memory slices are allocated for the respective buffer) or whether capacity of the buffer may be freed (thus freeing at least some of the memory slices of the respective buffer). Thus, in addition to an optimized utilization of the memory with reduced power consumption, the memory may also be flexibly utilized for larger demands. This provides a highly flexible memory architecture with reduced power consumption for various computational tasks.

[0027] A second aspect of the present disclosure defines a computer-readable medium having instructions stored therein, the instructions, in response to execution by a computing device, causing the computing device to perform a method according to any one of the preceding embodiments. In particular, the method may comprise providing a plurality of memory slices in a memory, determining a buffer width for a set of arithmetic operations, allocating some of the memory slices, and performing the set of arithmetic operations using the buffer, where the some of the memory slices form a buffer having at least the determined buffer width.

[0028] However, it should be understood that a computer-readable medium according to an embodiment of the second aspect of the present disclosure may store instructions that cause a computing device to perform any one of the method steps of the embodiment of the first aspect of the present disclosure in any combination.

[0029] A third aspect of the present disclosure defines an apparatus comprising a memory having a plurality of memory slices and at least one processing logic circuit, the processing logic circuit configured to execute a method according to any one of the preceding embodiments of the first aspect to configure the memory as a buffer and to perform a set of arithmetic operations using the buffer. In particular, the apparatus comprises a memory having a plurality of memory slices and at least one processing logic circuit, the processing logic circuit configured to determine a buffer width for the set of arithmetic operations, allocate some of the plurality of memory slices, and perform the set of arithmetic operations using the buffer, where the some of the memory slices form a buffer having at least the determined buffer width.

[0030] The processing logic may be any kind of hardware and / or software-based logic, which may be a programmable element capable of performing the method steps described above. The processing logic may be any kind of logic unit, such as a FSM, a combinatorial logic circuit, a look-up table, an encoder / decoder, etc. The processing logic may also be a processor, such as a general-purpose processor or a dedicated processor for a specific task.

[0031] It should be understood that an embodiment of the apparatus according to the third aspect may include configurations of the processor and / or memory according to features of the embodiment of the first aspect of the present disclosure, in any combination. Preferably, the memory may be configured according to an embodiment of the first aspect, and the processor may perform the configuration of the memory and perform the respective arithmetic operations on the allocated buffers in the memory. Furthermore, the memory may include logic circuitry that may be used to configure the memory to provide the respective buffers to the processor without further configuration steps.

[0032] For example, as described with respect to the embodiments of the first or second aspect, the processor may receive input data, process it using arithmetic operations, and further analyze the input data against statistical parameters of the input data. This information may be forwarded to logic circuits in the memory to potentially adjust the buffer width, or may be further processed by the processor to dynamically adjust the buffer width, communicate new buffer widths to the memory for dynamic allocation of buffers, or this information may be forwarded to logic circuits and further processed by the processor. Thus, processing performed by the processor in one embodiment may be at least partially performed by logic circuits in the memory in another embodiment, and vice versa.

[0033] Yet another aspect of the present disclosure discloses a dynamic memory. The dynamic memory comprises a plurality of memory slices, each memory slice having a slice width. The dynamic circuit further comprises a logic circuit. The logic circuit is configured to determine a buffer width for a set of arithmetic operations and allocate some of the memory slices of the plurality of memory slices. The memory slices are configured to form a buffer having at least the determined buffer width, provide the buffer, and perform the set of arithmetic operations using the buffer. Preferably, the logic circuit may be configured to perform the method according to any one of the embodiments of the first aspect. Furthermore, the dynamic memory may be a memory in an apparatus according to an embodiment of the third aspect. In this case, the logic circuit of the dynamic memory may communicate with the processor and provide the dynamic buffer for the apparatus.

[0034] In a preferred embodiment, the logic circuitry may be a processing unit that may be configured to perform arithmetic operations. In this embodiment, the memory may be configured to receive a set of arithmetic operations from a host computing device.

[0035] In yet another embodiment, the logic circuitry may be configured to determine a buffer width and allocate a memory slice for the buffer. A set of arithmetic operations is performed on the host computing device using the buffer.

[0036] Preferred embodiments of the dynamic memory may include the structural and / or functional features of the apparatus according to the first and third aspects of the present disclosure in any combination.

[0037] Certain features, aspects, and advantages of the present disclosure will become better understood with regard to the following description and accompanying drawings. [Brief description of the drawings]

[0038] [Figure 1] 1 shows a flowchart of a method according to an embodiment of the present disclosure. [Diagram 2] 1 illustrates a schematic layout of a memory according to an embodiment of the present disclosure. [Diagram 3] 1 illustrates a schematic layout of a memory according to an embodiment of the present disclosure. [Figure 4] 1 illustrates implementation details of at least some portions of a memory according to one embodiment of the present disclosure. [Diagram 5] 1 illustrates implementation details of at least some portions of a memory according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0039] In the following description, reference is made to the drawings which illustrate various embodiments. Various embodiments are also described below by reference to several examples. It is to be understood that the embodiments may include changes in design and structure without departing from the scope of the claimed subject matter.

[0040] 1 shows a flowchart of a method according to an embodiment of the present disclosure. Method 100 may begin at item 102. Method 100 may be a method for memory management and / or for dynamic buffer allocation and / or for dynamic buffer width allocation. In particular, method 100 may relate to a method for dynamic buffer width allocation.

[0041] In item 104, a memory having a plurality of memory slices is provided. The memory may be configured in hardware as several memory cells or units each capable of storing one bit, several bits, or one or several words of a certain length. For example, the memory may be a random access memory. A memory slice may represent several memory cells. At least some of the memory slices may be selectively activated or deactivated. This may cause activation or deactivation of the memory cells in which the memory slice is located.

[0042] The method 100 may proceed with determining a buffer width for the set of arithmetic operations in item 106. The buffer width may be determined based on one or more parameters. The one or more parameters may include word lengths of values ​​of input data processed by the set of arithmetic operations, some results of the set of arithmetic operations, and / or statistical parameters related to the input data, in any combination. At least some of the parameters may be provided as metadata. For example, the method 100 may include receiving, e.g., on the host system, metadata including pre-calculated statistical parameters for the input data. Additionally or alternatively, the parameters may be updated in real time based on currently processed values ​​of the input data and / or the computational flow of the arithmetic operations.

[0043] The arithmetic operation may include a multiply-accumulate operation that multiplies a plurality of weights by a plurality of input data. The results of the multiple multiply-accumulate (MAC) operations may be accumulated in a buffer. Thus, the buffer width may be determined based on a parameter related to the weights and a parameter related to the input data. For example, if the rate is fixed, the buffer may be determined based on the weights of the input data and the value of a statistical parameter, such as a maximum value of the input data.

[0044] In one example related to the convolutional layer of a neural network used for computer vision, one data point or "pixel" of the output feature map of the convolutional layer may be the sum of thousands or tens of thousands of multiplication products. This may involve (Kernel[x]*Kernel[y]*InputChannel[i]) of MAC cycles. Preferably, a partial sum (accumulation) is created every 10-100 multiplications. The processing logic may be dedicated hardware, general purpose hardware, or software, or a combination thereof, and may include parallel multipliers that generate parallel feature map data points (X and Y directions). The parallel multipliers may receive equal weighted but different portions of the input channels at a time. Each parallel multiplier may have a respective buffer for its individual partial sums.

[0045] The buffer width may be determined to be capable of storing and fetching partial sums. As an example, for 9-bit signed integer values, the multiplication may result in a 17-bit integer product, providing the numerically largest weight and input value. The result of the accumulation (e.g., before scaling down to INT8) may be 32 bits wide (without overflow, having a summed 100k + max product).

[0046] The memory may be partitioned into multiple memory slices with different sizes. For example, one memory slice representing 16 bits and multiple memory slices of 2 bits. To determine the buffer width, the absolute values ​​of the weights can be accumulated and the maximum input value can be numerically estimated. This can produce an upper bound width estimate for the current partial sum width used to set up the buffer. In practice, the buffer width is likely to be a few bits lower.

[0047] In item 108 of the method 100, some memory slices of the plurality of memory slices may be allocated (or activated). The some memory slices constitute a buffer having at least the determined buffer width. In the above example, the buffer may be configured to include a 16-bit slice and some 2-bit slices necessary to accommodate partial width values ​​of the estimated size. If the estimated width is 24 bits, 16-bit+2-bit+2-bit+2-bit memory slices of the buffer are configured as follows: The magnitude (useful bit count) of the weight accumulation result described above can be applied as the activation of the memory slices.

[0048] Thus, embodiments of the present disclosure may represent a fully embedded memory. The memory may represent a dynamic buffer or an array of dynamic buffers that functionally behave as a normal buffer or memory. After an initial setup, the power consumption of the embedded memory is reduced. The memory may determine slice allocation based on input data or may use a logic unit that may be integrated with the memory buffer or the host system. The logic unit may be any type of logic unit or programmable element, such as, by way of example, a MAC cycle counter. Thus, in one or more embodiments, the memory buffer does not require pre-placement or pre-configuration for future operation or host / application intervention.

[0049] In item 110, an arithmetic operation, such as one or more of the MAC operations described above, is performed or executed using a buffer. The buffer containing the first few slices can be used for several partial results that compute with different inputs but share the same weights, as described above. During writing to the buffer, only the allocated (or activated) memory slices can be accessed.

[0050] The memory may further include an auxiliary buffer, such as a 4-bit or 8-bit wide auxiliary buffer, or an auxiliary buffer of any other suitable or required width. The auxiliary buffer may store a segment pattern for each individual write. During a read, the auxiliary buffer may be read and based on the output of the auxiliary buffer, only the relevant segment used for the write may be accessed.

[0051] The memory may comprise a write side and a read side, as shown in FIG. 2, for example, where the write side includes a write port and the read side includes a read port. Data is written to the memory slices on the write side. Similarly, data is read from the memory slices on the read side. In one or more embodiments, the buffer width on the write side of the memory may be determined using available metadata, which may be provided in any suitable manner, such as, for example, a side channel on the bus, or a bus including the number of bits used for each write or parallel write, logic that checks the MSBs until a useful bit is found, or a combination thereof. Other approaches may be used. A theoretical upper limit may be approximated or estimated for an array of values ​​written in the same cycle. This may include a safety / pessimistic technique that ensures that no value can utilize more bits than is approximated. This upper limit may then be used for the allocation of memory slices of the memory. Thus, a pessimistic data width is used, so a compromise can be made between the power savings achieved and the complexity of the implementation and the power cost of the control logic.

[0052] In addition to width considerations on the write side, embodiments may also configure the read side of the memory to access and read only the necessary, i.e., allocated, memory slices.

[0053] In one or more embodiments, neither the write nor read access patterns may be random, and some correlation may exist between values ​​written to different addresses or memory slices of the buffer. For example, it may be assumed that if address X is read, the read continues at X+1, X+2, ..., X+n, where n may be any suitable integer, such as 15, or an example of a different upper limit of any value, for a set of cycles. Address X may correspond to a particular memory slice. A memory slice may include one or more addresses. This may be further based on the assumption that the maximum data width difference between each address offset may be up to 1, or a different suitable value. Based on this assumption, for at least some or all read cycles, the necessary information may be available to access only the memory slice that actually stores the written data. Using this technique, the data width written to a particular address may be approximated.

[0054] In one or more embodiments, by using an auxiliary buffer, complex control and respective additional power consumption can be avoided. The auxiliary buffer can have the same depth as the dynamic buffer or the array of memory slices representing the parallel dynamic buffer. The auxiliary buffer can include an active slice pattern for each address that can be updated with each new write. During a read, the read address can be first used to access the contents of the auxiliary buffer that includes the slice pattern. In the slice pattern, an actual read can be issued to the dynamic buffer, but only on the required memory slice. This realizes additional memory access power cost savings in the read cycle. This approach can be implemented in various ways, and the disclosure is not limited to a particular implementation. One possible implementation can include delaying read access to the dynamic buffer by only one clock cycle. Other implementations can depend on the integration environment that can be used to implement the read address look-ahead mechanism.

[0055] Although the auxiliary buffer introduces additional memory hardware that needs to be accessed in each active cycle to save power during memory access, this may still improve overall power consumption. For example, in one embodiment having four memory slices with an exemplary configuration of 20+4+4+4, as an example, the auxiliary buffer may be coded with 2 bits. The 20-bit lowest memory slice may always be active (or allocated). Assuming that the more significant memory slices may only be active (or allocated) if all the lowest memory slices are active (or allocated), the three more significant 4-bit memory slices may be coded as four values. For example, if the average width of the data written to the buffer is 24 bits, i.e., a 20-bit memory slice and a subsequent 4-bit memory slice, by adding a 2-bit auxiliary buffer, only 26 bits of memory width are accessed instead of 32 bits of all four memory slices. Thus, using an auxiliary buffer may represent a trade-off between additional memory hardware and power consumption due to added control logic. For implementations where an array of dynamic buffers are fed simultaneously with similar data patterns, the overhead generated by the auxiliary buffer becomes even more negligible since one auxiliary buffer can serve a large number of dynamic buffers, such as tens or hundreds of dynamic buffers. Independent of determining which memory slices need to be active (or allocated) for a particular write to a particular address, the auxiliary buffer can achieve a comparable savings on read accesses.

[0056] In one or more embodiments, the segment pattern may be a bit vector that includes one bit for each memory slice. Preferably, these bits are routed from the read port of the auxiliary buffer to the memory cells or units of the particular memory slice, and may participate in an AND logic relationship with the read enable input to the buffer, for example. Thus, only those memory slices will read their data that have a set bit in the segment pattern. The segment pattern may be coded using a unary code or a thermometer code. Unary coding or thermometer coding is an entropy coding that represents a natural number n with a sequence of n or (n-1) first symbols, such as 1 or 0, followed by a different symbol, such as 0 or 1, respectively. For example, a unary code (or representation) of a natural number n may be n ones followed by a zero. The code may be padded to the required width. For example, the number 3 may be coded as "111" or "1110" or "0001", etc. A padded version of the unary code of the number 3 may be "0000111", etc. Other representations exist, including other unary codes or thermometer codes, one-hot codes, or binary codes. For example, the one-hot code for the number 3 can be "000100". Those skilled in the art will appreciate that other encoding techniques and representations and codes can be used. Using thermometer encoding, the segment pattern can effectively be a thermometer encoded memory slice enable signal.

[0057] In one or other embodiments, on the write side, the segment pattern can arrive at the buffer from a write logic measurement unit or any other suitable logic on the write side. From the write port, the segment pattern can be written to an auxiliary buffer. The bits can also be routed to memory cells or units of specific memory slices and involved in an AND logic relationship with the write enable input to the buffer. Thus, only those memory slices are written with input data that has a set bit in the segment pattern.

[0058] It should be understood that the segment pattern may be implemented as a bit vector or in any other suitable manner. The segment pattern may also be referred to as a (memory) slice activation pattern. In one or other embodiments of the present disclosure, other encodings than temperature (gauge) encoding may be used to reduce the auxiliary buffer width. This may require additional combinatorial logic circuitry to generate the activation control bits per memory slice. In a preferred embodiment, the auxiliary buffer may be as wide as a controllable number of segments in the full buffer word length.

[0059] During instruction execution of the arithmetic operation in item 110, the value or some or all values ​​stored in the buffer may be monitored, preferably in real time or based on a number of clock cycles (after or before a certain amount of clock cycles) or a number of reads / writes (after or before a certain amount of reads, writes, or reads / writes) to determine whether the size of the buffer is large enough to store future results of the arithmetic operation. Additionally or alternatively, the monitored values ​​may be used to determine whether the size of the buffer is too large and may be reduced. This may be done by determining whether the monitored values ​​exceed or fall below respective thresholds that may indicate a potentially small buffer size. Thus, after instruction execution of the arithmetic operation in item 110 or in parallel with instruction execution of the arithmetic operation, method 100 may reiterate item 106 to proceed to determining an expected buffer width in item 106. If more memory is needed, the method may proceed to item 108 to dynamically allocate (or activate) at least one further memory slice of the multiple memory slices for the buffer.

[0060] Depending on the determined buffer width, in one or more embodiments, the method 100 may proceed to item 112. One or more of the allocated memory slices of the plurality of memory slices may also be deallocated (or deactivated). In the above example, at least some of the upper 2 or 4 memory slices of the buffer may be deallocated or deactivated. The method 100 may proceed to item 110.

[0061] It should be understood that the allocation (activation) and / or deallocation (deactivation) of memory slices in items 108 and 112 may be performed sequentially or in parallel to the instruction execution of the operation in item 110. Also, the allocation and deallocation may be combined in a single step. Furthermore, it should be understood that the deallocation in item 112 may be optional, as it may require additional logic circuitry and power consumption, as indicated by the dashed line. Thus, in embodiments targeting a fixed or larger memory size, the deallocation may be omitted. Furthermore, after performing steps 106 and 108, the method 100 may not loop back to item 16 during execution, i.e., during the execution of the operation in step 110. Thus, the method 100 may preconfigure the dynamic buffers in items 106 and 108 and perform the operation in 110 before proceeding to item 114, where the method 100 may end.

[0062] 2 shows a schematic layout of a memory according to an embodiment of the present disclosure. The embodiment shown in FIG. 2 shows a simple dual-port memory. However, it should be understood that the principles can be applied to any other memory layout, such as true dual-port, single-port memories, etc.

[0063] The memory 200 may define a buffer 202 including multiple memory slices 204a, 204b, ..., 204n. The widths of the memory slices 204 may vary. For example, the memory slice 204a may have a first width, such as 16 bits. The memory slices 204b-204n may have a second width, such as 2 bits or 4 bits, that may be different from the first width. One or more of the memory slices 204b-204n may also have an additional width, which may be different from the first width, the second width, and / or any other width of the memory slices. The memory 200 may include additional memory slices (not shown), each having a different width or the same width, that are not currently allocated for the buffer 202. The memory 200 may further include write logic 206 on a write side of the memory 200 and read logic 208 on a read side of the memory 200. The write logic 206 may monitor and control a write port, denoted as WR ENi, that enables writing of data to the buffer 202. The read logic 208 may monitor and control a read port, denoted as RD ENi, that enables the reading of data from the buffer 202 .

[0064] The write logic 206 and read logic 208 may also be used to update or retrieve patterns in the auxiliary buffer 210 (AUX RAM) using data lines shown as WR DATA AUX* and RD DATA AUX* in Figure 2. Similar to the embodiment of Figure 1, the auxiliary buffer 210 may store segment patterns for writing and / or reading that indicate the assigned memory slices 204a,...,204n. The patterns or segment patterns may also be referred to as (memory) slice activation patterns in one or more embodiments of the present disclosure.

[0065] The input to the logic circuitry 206 can convey word length information to the buffer 202 during the actual write cycle. This can be used to generate the segment pattern. In one or more embodiments, the input can be converted to a thermometer encoded slice enable bit in the write side logic circuitry 206. The slice enable bit can then participate in an AND relationship with the current WR EN input value in the logic circuitry 206 to individually drive the WR EN inputs of the memory slices.

[0066] Each WR_ENi input may be functionally a single bit. A write may enable both memory slices 204a,...,204n and auxiliary buffer 210 in parallel. The generated slice enable bit may also be routed from logic circuitry 206 to a write port of auxiliary buffer 210 and stored in auxiliary buffer 210 upon WR_EN assertion.

[0067] The auxiliary buffer 210 may receive the same write address WR ADDR that is also received by the memory slices 204a, ..., 204n. Thus, the auxiliary buffer 210 may store word length information or slice enable bits for at least one or every address. Thus, the auxiliary buffer 210 may store word length information or slice enable bits for a smaller number of addresses depending on the implementation in one or more embodiments of the present disclosure.

[0068] On the read side of memory 200, the read address RD ADDR and read enable RD EN are connected directly to the read port of the auxiliary buffer 210. The segment enable bits can be read in a single clock cycle. The read side logic 208 can combine the read bits with an AND relationship with the read enables and can drive the read enables for each memory slice 204a,...,204n individually.

[0069] The read address and enable bits may be delayed to memory slices 204a,...,204n by a single cycle. This is indicated in Figure 2 by the boxes labeled "-1" on each connection. Thus, each signal is coherent with the slice read enable generated by logic circuitry 208 due to the one cycle read of auxiliary buffer 210.

[0070] The read port of the auxiliary buffer 210 may also be connected to the read output of the buffer 202 with a respective delay that depends on the read latency of the memory slices 24a, 204n. An external logic circuit (not shown), such as a word length reconstructing logic circuit, may be used to generate full length data words from the read memory slices. The external logic circuit may, for example, replicate the significant most significant bit up to the most significant bit of the full word length to reconstruct the full word length. The significant most significant bit of each read may be the most significant bit of the most significant index slice.

[0071] Additionally or alternatively, the word length control may be internal to the read side logic 208 of the buffer if the logic 208 is configured to receive read data. In yet another implementation, slice enable information may be used by subsequent stages to process the read data, thereby allowing information outside the buffer to be used.

[0072] As described above, the auxiliary buffer 210 can track the data width of the read side. Additionally, the word width of the write side can be measured by the logic circuitry 206 in real time, which can be updated in the auxiliary buffer 210. The measured and monitored values ​​can be used to determine or adjust the width of the buffer 202. In a preferred embodiment, the logic circuitry 206 can receive metadata or parameters related to the write data to determine or adjust the width of the buffer 202. The parameters can include a data width parameter that can be provided every clock cycle or updated between several accesses. The number of accesses can be related to similar width accesses that can be performed in bulk.

[0073] Each memory slice 204a,...,204n can receive separate read and write controls (enable / read / write / address, etc.). The read and write controls can be driven by real-time metadata such as statistics. The statistics can be collected on the input values, or pre-calculated and sent with each data, or pre-calculated and sent in bulk (one such statistic for many data). This setup allows to always keep track (or pre-calculate) the word length used, which allocates only the amount of memory slice required for storage. This saves considerable power, since the remaining memory slices of the memory 200 can remain unavailable (disabled) or waiting (idle). The auxiliary buffer 210 can be utilized to realize similar savings on the read side.

[0074] Thus, logic circuitry 206 and / or logic circuitry 208 may be used to monitor values ​​stored in the buffer and monitor and / or update metadata that may be used to dynamically adjust the size of buffer 202, e.g., as described with respect to the embodiment of FIG. 1.

[0075] 2, and as noted above, the boxes labeled "-1" may represent one cycle delays (register stages) placed on the actual signals. They cause the auxiliary buffer 210 to first read and then provide its read word length value to be used as a subsequent read of the allocated (or activated) memory slice. Thus, the schematic diagram shown in FIG. 2 assumes a one cycle read latency of the auxiliary buffer 210.

[0076] 1 and 2 describe several methods for determining the required buffer width on the write side of the buffer 202, for example based on readily available metadata (such as a side channel on the bus or a bus containing the number of bits used for each write or parallel write), calculating the required buffer width for each buffer (logic circuitry checking the most significant bit until a useful bit is found), and / or a combination of the two. For example, a theoretical upper limit can be approximated / estimated for an array of values ​​written in the same cycle, and approximated using a safe / pessimistic technique that ensures that no value can utilize more bits than the one that is approximated. The upper limit can then be applied to the array of memory slices, making a compromise between the achieved power savings based on the pessimistic data width and the implementation complexity and power cost of the control logic. Other techniques for determining the required buffer width may be used. These approaches can be advantageously applied to the write side, for example using logic circuitry 206.

[0077] The logic circuitry 206 may receive an additional 3-bit wide bus with each bit representing an additional slice that needs to be active (or assigned) above slice 0, such as memory slice 204a. The logic circuitry 206 may process input and control write enable or simple enable signals for each individual memory slice 204b, ..., 204n to activate (or assign) one or more of the additional memory slices 204b, ..., 204n. Additionally, the logic circuitry 206 may encode each 3-bit value into 2 bits for the auxiliary buffer 210.

[0078] Alternatively or additionally, depending on the situation or use case, logic circuitry 206 can receive the data to be written to buffer 202 and determine the required slice combination on the fly. As previously mentioned, logic circuitry 206 can control individual “enable” signals for memory slices 204 a,...,204 n and encode values ​​for auxiliary buffer 210.

[0079] In yet another preferred implementation, logic circuitry 206 may represent an on-the-fly approximation logic circuit that extrapolates the maximum required bit width for an array of memory slices operating in parallel. As previously described, logic circuitry 206 may control individual “enable” signals for memory slices 204 a,...,204 n and encode values ​​for auxiliary buffer 210.

[0080] The read side logic 208 can perform operations corresponding to those performed by the write side logic 206. The logic 208 can decode the contents of the auxiliary buffer 210, which may be decimal (or otherwise) coded content, and recover the slice access pattern for a particular read address. The logic 208 can control the timing and drive each of the enable signals for the memory slices 204a,...,204n on the read side of the buffer 202. The logic 208 can also perform word length reconfiguration.

[0081] The embodiments of the present disclosure may not require such a specific addressing scheme. In one or more implementations, the dynamic memory 200 may be used just like a normal memory with any addressing. However, it should be understood that the embodiments may define the addressing scheme. For example, a typical addressing pattern may be used, and a typical data width distribution and data width pattern may be analyzed as further parameters when configuring and implementing the dynamic buffer 202 of the memory 200 for a particular application or use case. This may be done in a pre-configuration phase to determine which control method and which slice configuration may yield the best results.

[0082] It should be understood that either logic circuitry 206 or logic circuitry 208 may be optional and / or provided as one logic circuit unit that can monitor values, parameters, and determine appropriate word widths on the write side, read side, and / or both sides. Buffer 202, memory 200, including preferably one or more of logic circuits such as logic circuitry 206 and / or logic circuitry 208, and auxiliary buffer 210 may be configured to provide functionality as described above with respect to FIG.

[0083] FIG. 3 illustrates a schematic layout of a memory according to an embodiment of the present disclosure. Similar to memory 200 of FIG. 2, memory 300 may include a buffer 302 including multiple memory slices 304a, 304b, ..., 304n, as shown in FIG. 3. The width of memory slices 304 may vary as described above with respect to FIG. 2. Memory 300 may include additional memory slices (not shown) not currently allocated for buffer 302. Memory 300 may further include write logic 306. As indicated by the dashed line, memory 300 may optionally include read logic 306. Write logic 306 may monitor and control a write port that allows data to be written to buffer 302. Optional read logic 308 may monitor and control a read port that allows data to be read from buffer 302.

[0084] The word width of the write side may be measured by logic circuit 306 in real time. The measured and monitored values ​​may be used to determine or adjust the width of buffer 302. If logic circuit 308 is not provided in memory 300, the full width of buffer 302 may always be read. In a preferred embodiment, memory 300 includes logic circuit 308. Logic circuit 306 may receive parameters related to the write data to determine or adjust the width of the buffer. Thus, the reading of data is controlled by logic circuit 308.

[0085] Logic circuitry 306 and / or 308 may be configured similarly to the embodiment of Figure 2. Logic circuitry 306 and / or logic circuitry 308 may be used, for example, to monitor values ​​stored in the buffer and to monitor and update metadata that may be used to dynamically adjust the size of buffer 302, as described with respect to the embodiment of Figure 1.

[0086] The memory 200, 300 according to the embodiment of Figs. 2 and 3 may represent a memory device or arrangement that may be provided as dedicated hardware or as part of dedicated hardware that may be embedded in, coupled to, and / or connected to a host device (not shown). The host device may be a general-purpose computing device or dedicated hardware such as an accelerator. The host device may provide metadata and further information to set up the buffers 202, 302 in the memory 200, 300 and to write data to the memory 200, 300. The data may represent the result of an arithmetic operation as described above with respect to Fig. 1. The arithmetic operation may be performed on the host device and / or in the dedicated hardware and / or on the logic circuit unit (not shown) of the memory 200, 300.

[0087] By tracking the actual size of the data to be stored, embodiments of the present disclosure limit the power costs on silicon associated with storing and retrieving data to a minimum, based on the fact that buffers typically need to accommodate a theoretical maximum size value, but in practice this maximum, or sizes close to this maximum, are rarely reached in many applications.

[0088] Embodiments of the present disclosure can be based on the statistical distribution of sizes of numeric data (such as different partial sums), especially compared to the maximum sum. This can be used to allocate memory slices or storage elements of memory and dynamically control them based on current data size to save significant power associated with buffer accesses. According to embodiments of the present disclosure, measurements reveal power savings of up to 30-40% for memory.

[0089] For example, in CNN acceleration techniques, significant power cost savings can be achieved. The embodiments of the present disclosure can be used in deep neural network training and / or inference. When performing a convolution, a single output value can be the result of tens of thousands of multiply-add operations. This can result in a large number of accumulations, such as tens of thousands of accumulations. However, the (tens of thousands of) input data that contribute to this singular output are also used to calculate other outputs. Therefore, it is recommended to load a chunk of input data and perform all calculations on all contributing output data. Partial results can be stored in a dynamic buffer with a width adjusted to the expected size of the result. This significantly reduces the associated power cost of the buffer.

[0090] However, it should be understood that embodiments of the present disclosure are not limited in this respect and may be advantageously used in, for example, graphic processing or scientific computing devices, units, and / or accelerators. Other areas may include digital signal processing, video and audio encoding and decoding, blockchain technology, etc.

[0091] Throughout this disclosure, the terms "word width" and "word length" may be used interchangeably to refer to the number of bits required to encode a word or data word. To store a word or data word having a particular word width or word length, the respective buffer or memory must provide at least a buffer width or memory width that can accommodate the number of bits of the word or data word.

[0092] Figures 4 and 5 show implementation details of at least some portions or components of a memory according to respective embodiments of the present disclosure. Figures 4 and 5 may represent silicon implementations of some of the embodiments described with respect to Figures 1, 2, and 3. Thus, one or more components shown in Figures 4 and 5 may correspond to implementations of one or more components shown in Figures 2 and 3, or may implement a process corresponding to one or more steps of the method described with respect to Figure 1.

[0093] Figure 4 shows implementation details of a dynamic memory 400. At least some components of the dynamic memory 400 may correspond to components of the memory 200, as shown in Figure 2. In Figure 4, small vertical rectangles on the connections represent one-cycle delay register stages.

[0094] The dynamic memory 400 may include a circuit 402 that may implement an auxiliary buffer, such as the auxiliary buffer 210 (AUX RAM) of Figure 2. The circuit 402 may include a read port RDin for loading data from a memory unit 404 and a write port WRin for storing data to the memory unit 404.

[0095] The dynamic memory 400 may further include a circuit 406 that may include one or more memory units 408 that may correspond to the memory slices 204a, ..., 204n of the buffer 202 of FIG. 2. The circuit 406 may include a write port WRin for storing data in the one or more memory units 408. The circuit 406 may further include read ports RDin and RDout for loading data from the one or more memory units 408. The read port RDin may receive a data pattern or bit vector, such as a slice enable pattern, stored in the memory unit 404 to indicate active portions of the one or more memory units 408. The contents of the active portions of the one or more memory units 408 are provided via the read port RDout.

[0096] The dynamic memory 400 may further include a circuit 410 that performs word length reconfiguration (WL reconfiguration) on data read from the active one or more memory units 408 according to the slice enable pattern provided by the memory units 404 to regenerate full length data words from the read data. A completely self-contained implementation of the dynamic memory 400 presents itself as a memory with the maximum data width at all of its ports.

[0097] 5 shows implementation details of a circuit 500 applicable in one or more embodiments of the present disclosure. The circuit 500 can perform magnitude approximation with coefficient absolute value accumulation. As in FIG. 4, the small vertical rectangles on the connections in FIG. 5 represent one-cycle delay register stages.

[0098] The magnitude approximation may be performed in parallel with the MAC operation (not shown). Thus, the result of the MAC operation, and the result of the circuit 500 that provides the approximated magnitude for the MAC operation, may be fixed to a cycle-by-cycle timing relationship on the result side of the circuit 500. Thus, the operations implemented by the circuit 500 may represent the MAC coefficient absolute sum used for cycle-by-cycle buffer slice allocation in dynamic memories, such as memories 200 and 300, as described above with respect to FIGS. 2 and 3.

[0099] It should be understood that the implementation details provided in Figures 4 and 5 represent preferred examples. Other implementations using different components, circuits, connections, and links can be used, and the present disclosure is not limited by a particular implementation in silicon.

[0100] The techniques described herein may be implemented in various computing systems, examples of which are described in more detail below. Such systems generally involve the use of appropriately configured computing devices that implement several modules, each providing one or more operations required to complete the execution of such techniques. For example, each module may be implemented to provide the functionality of a method step according to an embodiment of the present disclosure. Each module may be implemented in its own way, and need not all be implemented in the same way. As used herein, a module performs an operational role, but is a structural component of an instantiated system, which may be part or the whole of a software element (e.g., a function of a process, a discrete process, or any other suitable embodiment). A module may include computer-executable instructions and may be encoded on a computer storage medium. Modules may be executed in parallel or serially, as appropriate, and may pass information to each other using shared memory on the computer on which they are executed, using a message passing protocol, or in any other suitable manner. Although exemplary modules are described below to perform one or more tasks, it should be appreciated that the described modules and division of tasks are merely illustrative of the types of modules that may implement the exemplary techniques described herein, and the invention is not limited to being implemented with any particular number, division, or type of modules. In some implementations, all functionality may be implemented within a single module. Furthermore, while the modules are described below as all executing on a single computing device for clarity, it should be understood that in some implementations the modules may be implemented on separate computing devices adapted to communicate with each other. For example, one computing device may be adapted to execute an identification module for identifying available networks. A connection module on the other computing device may retrieve information on available networks from the computing device before establishing a connection.

[0101] Although several embodiments have been described in detail, it should be understood that aspects of the present disclosure can take many forms. In particular, the claimed subject matter may be practiced or implemented differently from the specific examples described, and the described features and characteristics may be practiced or implemented in any combination. The embodiments shown herein are illustrative, rather than limiting, of the invention as defined by the claims.

Claims

1. A method for dynamic buffer width allocation, The steps include providing a memory having multiple memory slices, The steps include determining the buffer width for the set of arithmetic operations, A step of allocating some memory slices from a plurality of memory slices, wherein the some memory slices form a buffer having at least a determined buffer width. A method that includes the steps of performing a set of arithmetic operations using a buffer.

2. The method according to claim 1, wherein the memory enables selective activation and deactivation of at least some of the plurality of memory slices.

3. The method according to claim 1 or 2, further comprising the step of deactivating an unallocated memory slice among the plurality of memory slices.

4. The method according to claim 1 or 2, wherein the buffer width is determined based on the word length of the values ​​processed by the set of arithmetic operations.

5. The method according to claim 1 or 2, wherein the buffer width is determined based on the result of the set of arithmetic operations.

6. The method according to claim 1 or 2, wherein the buffer width is determined based on statistical parameters related to the input data processed by the set of arithmetic operations.

7. The method according to claim 6, further comprising the step of collecting the statistical parameters in real time during the processing of the set of arithmetic operations.

8. The method according to claim 6, further comprising the step of receiving the statistical parameters related to the input data.

9. The method according to claim 8, wherein the statistical parameters are pre-calculated for the input data.

10. The method according to claim 1 or 2, further comprising the steps of: monitoring at least one value to be stored in the buffer; determining that the value has exceeded a threshold; and dynamically allocating at least one further memory slice from the plurality of memory slices for the buffer.

11. The method according to claim 1 or 2, wherein the set of arithmetic operations includes a plurality of sum-of-products operations, and the results of the plurality of sum-of-products operations are accumulated in the buffer.

12. The method according to claim 11, wherein the plurality of sum-of-products operations include multiplying input data by a plurality of weights, and the buffer width is determined based on the values ​​of the weights and the assumed maximum value of the input data.

13. The method according to claim 1 or 2, further comprising the step of allocating some further memory slices from the aforementioned memory slices to a further buffer.

14. It is dynamic memory, Several memory slices, each having a slice width, A dynamic memory comprising a logic circuit configured to perform a method for dynamic buffer width allocation, The aforementioned method, The steps include determining the buffer width for the set of arithmetic operations, A step of allocating some memory slices from a plurality of memory slices, wherein the some memory slices form a buffer having at least a determined buffer width. Dynamic memory, including the step of performing a set of arithmetic operations using a buffer.

15. The dynamic memory according to claim 14, wherein the logic circuit enables selective activation and deactivation of at least some of the plurality of memory slices.

16. The method for dynamically allocating buffer width, further comprising the step of deactivating an unallocated memory slice among the plurality of memory slices, according to claim 14 or 15.

17. The dynamic memory according to claim 14 or 15, wherein the buffer width is determined based on the word length of the values ​​processed by the set of arithmetic operations.

18. The dynamic memory according to claim 14 or 15, wherein the buffer width is determined based on the result of the set of arithmetic operations.

19. The dynamic memory according to claim 14 or 15, wherein the buffer width is determined based on statistical parameters related to the input data processed by the set of arithmetic operations.

20. The dynamic memory according to claim 19, wherein the method for allocating the dynamic buffer width further comprises the step of collecting the statistical parameters in real time during the processing of the set of arithmetic operations.

21. The method for allocating the dynamic buffer width further comprises the step of receiving the statistical parameters related to the input data, according to claim 19.

22. The dynamic memory according to claim 8, wherein the statistical parameters are pre-calculated with respect to the input data.

23. The method for dynamically allocating a buffer width according to claim 14 or 15, further comprising the steps of: monitoring at least one value to be stored in the buffer; determining that the value has exceeded a threshold; and dynamically allocating at least one further memory slice from the plurality of memory slices for the buffer.

24. The dynamic memory according to claim 14 or 15, wherein the set of arithmetic operations includes a plurality of multiply-accumulate operations, and the results of the plurality of multiply-accumulate operations are accumulated in the buffer.

25. The dynamic memory according to claim 24, wherein the plurality of sum-of-accumulate operations include multiplying input data by a plurality of weights, and the buffer width is determined based on the values ​​of the weights and the assumed maximum value of the input data.

26. The method for dynamically allocating buffer width according to claim 14 or 15, further comprising the step of allocating some further memory slices from the plurality of memory slices to a further buffer.

27. A device comprising the dynamic memory described in claim 14.