Method for Compressing Operands, Method and System for Decompressing a Compressed Data Sequence

By compressing and decompressing neural network operations through reordering and removing redundant exponents, the method addresses the challenge of high energy consumption and bandwidth in neural networks, achieving efficient computation with maintained accuracy.

CN116957086BActive Publication Date: 2025-07-15MEDIATEK INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210553758.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-20
Filing Date
2022-05-19
Publication Date
2025-07-15
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

There are problems with high power consumption and high memory bandwidth requirements in neural network computing, especially when it is difficult to balance low power consumption and low memory bandwidth requirements while maintaining computing accuracy.

Method used

By reordering and compressing the operands of neural networks, redundant exponents are removed, compressed data sequences are generated, and operands are restored during decompression, lossless compression and decompression are achieved.

Benefits of technology

Reduces memory footprint and bandwidth requirements while maintaining computational accuracy without retraining the neural network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116957086B_ABST
    Figure CN116957086B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a method for compressing operand, a method for decompressing a compressed data sequence, and a system. A method for compressing an operand for neural network computing provided by the present invention includes: receiving a plurality of operands, where each operand has a floating-point representation including a sign bit, an exponent, and a fraction; reordering the plurality of operands into a first sequence composed of sign bits, a second sequence composed of exponents, and a third sequence composed of fractions; and compressing the first sequence, the second sequence, and the third sequence to at least remove duplicate exponents, thereby losslessly generating a compressed data sequence. Implementing the embodiments of the present invention can losslessly generate a compressed data sequence and can losslessly recover a plurality of operands.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to neural network processing, and more particularly, to methods for compressing operands, methods for decompressing compressed data sequences, and systems therefor. Background Art

[0002] A deep neural network is a neural network having an input layer, an output layer, and one or more hidden layers located between the input layer and the output layer. Each layer performs operations on one or more tensors. A tensor is a mathematical object that can be zero-dimensional (also known as a scaler), one-dimensional (also known as a vector), two-dimensional (also known as a matrix), or multi-dimensional. Some layers apply weights to tensors, such as in a convolution operation. Typically, the tensors produced by a neural network layer are stored in a memory and fetched by the next layer from the memory as inputs. Storing and fetching tensors, as well as storing and fetching any applicable weights, may consume a large amount of data bandwidth of the memory.

[0003] Neural network computations require intensive computing and bandwidth requirements. Modern calculators typically use floating-point numbers with a large bit width (e.g., 16 bits or 32 bits) in numerical computations to achieve high precision. However, high precision comes at the cost of high power consumption and high memory bandwidth. Balancing low power consumption and low memory bandwidth requirements while maintaining an acceptable accuracy of neural network computations is a challenge.

[0004] For example, the amount of bandwidth measurement and the number of multiply-and-add (MAC) operations have increased steadily at a rapid rate in the past decade. The types of neural network applications have evolved from image classification, object detection, image segmentation, depth / pose and motion estimation to image quality enhancement such as super-resolution. The latest neural network applications may require up to 10 trillion MAC operations and 1 gigabit per second of bandwidth. Summary of the Invention

[0005] Embodiments of the present invention provide methods for compressing operands, methods for decompressing compressed data sequences, and systems therefor.

[0006] A method for compressing operands for neural network computing provided by the present invention includes: receiving a plurality of operands, where each operand has a floating-point representation including a sign bit, an exponent, and a fraction; reordering the plurality of operands into a first sequence composed of sign bits, a second sequence composed of exponents, and a third sequence composed of fractions; and compressing the first sequence, the second sequence, and the third sequence to at least remove duplicate exponents, thereby losslessly generating a compressed data sequence.

[0007] A method for decompressing a compressed data sequence provided by the present invention includes: decompressing the compressed data sequence into a first sequence composed of N sign bits, a second sequence composed of N exponents, and a third sequence composed of N fractions, where N is a positive integer, and where the compressed data sequence represents N operands and does not contain duplicate exponents; reordering the first sequence composed of the N sign bits, the second sequence composed of the N exponents, and the third sequence composed of the N fractions into a restored sequence composed of N floating-point numbers representing the N operands; and sending the restored sequence composed of the N floating-point numbers to an accelerator for neural network computing.

[0008] An operand processing system provided by the present invention includes: an accelerator circuit; and a compressor circuit coupled to the accelerator circuit, the compressor circuit being configured to: receive a plurality of operands, where each operand has a floating-point representation including a sign bit, an exponent, and a fraction; reorder the plurality of operands into a first sequence composed of sign bits, a second sequence composed of exponents, and a third sequence composed of fractions; and compress the first sequence, the second sequence, and the third sequence to at least remove duplicate exponents, thereby losslessly generating a compressed data sequence. In an alternative embodiment, the compressor circuit is further configured to: decompress the compressed data sequence to losslessly restore the plurality of operands in floating-point format.

[0009] As described above, embodiments of the present invention at least remove duplicate exponents when compressing operands, and since the removed content is duplicate, a compressed data sequence is generated losslessly. When decompressing the data sequence, the opposite steps of compression are performed, so that a plurality of operands can be restored losslessly. Description of the Drawings

[0010] Figure 1 is a block diagram showing a system 100 for performing neural network operations according to one embodiment.

[0011] Figure 2A and Figure 2B shows the interaction between an accelerator 150 and a memory 120 according to two alternative embodiments.

[0012] Figure 3 is a diagram illustrating the conversion operations performed by a compressor according to one embodiment.

[0013] Figure 4 FIG. is a diagram showing data movement in a CNN operation according to an embodiment.

[0014] Figure 5 FIG. is a flowchart showing a method 500 for compressing floating-point numbers according to an embodiment.

[0015] Figure 6 FIG. is a flowchart showing a method 600 for decompressing data into floating-point numbers according to an embodiment. DETAILED DESCRIPTION

[0016] In the specification and claims, certain terms are used to refer to specific components. Those skilled in the art should understand that hardware manufacturers may use different terms to refer to the same component. The specification and claims do not use the difference in names as a way to distinguish components, but rather use the difference in functions of components as the criterion for distinction. The terms "comprising" and "including" mentioned throughout the specification and claims are open-ended terms and should be interpreted as "including but not limited to". "Substantially" means within an acceptable error range. Those skilled in the art can solve the technical problem within a certain error range and basically achieve the technical effect. In addition, the term "coupled" herein includes any direct and indirect electrical connection means. Therefore, if it is described in the text that a first device is coupled to a second device, it means that the first device can be directly electrically connected to the second device, or indirectly electrically connected to the second device through other devices or connection means. The following describes the preferred embodiments contemplated by the present invention. These descriptions are for explaining the general principles of the present invention and should not be used to limit the present invention. The protection scope of the present invention should be determined based on reference to the claims of the present invention.

[0017] The following description is of the optimal embodiments contemplated by the present invention. These descriptions are for explaining the general principles of the present invention and should not be used to limit the present invention. The protection scope of the present invention should be determined based on reference to the claims of the present invention.

[0018] Embodiments of the present invention provide a mechanism for compressing floating-point numbers used for neural network calculations. The compression exploits redundancy in floating-point operands (or operands for short) of neural network operations to reduce memory bandwidth. For example, multiple values of input activations (e.g., input feature maps) of a neural network layer may be distributed over a relatively small range of values that may be represented by one or a few exponents. Compression removes redundant bits representing the same exponent in an efficient and adaptable manner, so that compression and corresponding decompression may be performed dynamically during the inference phase of neural network calculations without requiring the neural network model to be retrained.

[0019] In one embodiment, compression and corresponding decompression are applied to operands (i.e., floating point operands) of a convolutional neural network (CNN). Typically, in a convolutional neural network, a tensor can be cut into multiple batches of operation arrays, each of which can include N operands (where N is a non-zero integer). For example, when N is 4, the 4 operands can be used for subsequent Figure 3 A0, A1, A2 and A3 shown in . Operands may include inputs (also called input activations), outputs (also called output activations), and weights of one or more CNN layers. Taking the input of a CNN neural network as an example, an input of each layer may be a tensor, so that each operand may include a segment of a CNN input. The compression method described in an embodiment of the present invention completes the compression of an input by performing compression on multiple batches of operation arrays respectively. Similarly, one or more weights or a CNN output may also be a tensor, and each operand includes the one or more weights or a segment of the output. It can be understood that compression and corresponding decompression are applicable to neural networks not limited to CNN.

[0020] In one embodiment, floating point compression is performed by a compressor that compresses floating point numbers (e.g., the aforementioned operands or floating point operands) into a compressed data sequence for storage or transmission. The compressor also decompresses the compressed data sequence into floating point numbers for the deep learning accelerator to perform neural network calculations. The compressor can be a hardware circuit, software (e.g., a processor that executes software instructions), or a combination of both.

[0021] The floating-point compression disclosed herein can be specifically used for workloads in neural network computing. The workload includes computations in floating-point operations. Floating-point operations are widely used in scientific computing and applications that require precision. As used herein, the term "floating-point representation" refers to a number representation having a fraction (also referred to as "mantissa" or "coefficient") and an exponent. The floating-point representation may also include a sign bit. Examples of floating-point representations include, but are not limited to, the IEEE 754 standard formats, such as 16-bit floating-point numbers, 32-bit floating-point numbers, 64-bit floating-point numbers, or other floating-point formats supported by some processors. The compression disclosed herein can be applied to floating-point numbers in various floating-point formats (e.g., IEEE 754 standard formats, their variants, and other floating-point formats supported by some processors). The compression is applied to both zero and non-zero values.

[0022] The compression disclosed herein can reduce memory footprint and bandwidth and can be easily integrated into deep learning accelerators. The compression is lossless; thus, there is no loss of precision. There is no need to retrain a neural network that has already been trained. The compressed data format adapts to the input floating-point numbers; more specifically, the compression applies to floating-point numbers with exponents of any bit width. There is no hard requirement for the bit width of the exponent of each floating-point number. The input floating-point numbers also do not have to share a predetermined fixed number of exponents (e.g., a single exponent).

[0023] Figure 1 is a block diagram showing a system 100 for performing neural network operations according to one embodiment. The system 100 includes processing hardware 110, which further includes one or more processors 170, such as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Digital Processing Unit (DSP), a Field-Programmable Gate Array (FPGA), and other general-purpose and / or special-purpose processors. The system 100 also includes a deep learning accelerator 150 (hereinafter referred to as "accelerator 150") for performing neural network operations; for example, tensor operations. Examples of neural network operations include, but are not limited to: convolution, deconvolution, fully connected operations, normalization, activation, pooling, resizing, element-wise arithmetic, concatenation, etc. In one embodiment, the processing hardware 110 can be part of a System-On-a-Chip (SOC).

[0024] In one embodiment, system 100 includes compressor 130 coupled to accelerator 150. Compressor 130 is used to convert floating-point numbers into a compressed data sequence that uses fewer bits than the floating-point numbers before conversion.

[0025] Compressor 130 can compress the floating-point numbers generated by accelerator 150 (e.g., the aforementioned operands or floating-point operands) for memory storage or transmission. In addition, compressor 130 can also decompress the received or acquired data from the compressed format into floating-point numbers for accelerator 150 to perform neural network operations. The following description focuses on compression for reducing memory bandwidth. However, it can be understood that compression can save transmission bandwidth when the compressed data is transmitted from system 100 to another system or device.

[0026] In Figure 1 the example of, compressor 130 is shown as a hardware circuit external to processor 170 and accelerator 150. In an alternative embodiment, software module 172 executed by processor 170 can perform the functions of compressor 130. And in yet another embodiment, circuit 152 within accelerator 150 or software module 154 executed by accelerator 150 can perform the functions of compressor 130. The following description explains the operation of the compressor using compressor 130 as a non-limiting example.

[0027] Processing hardware 110 is coupled to a memory 120, which may include on-chip memory and off-chip memory devices such as Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash memory, and other non-transitory machine-readable storage media; e.g., volatile or non-volatile storage devices. As used herein, the term "on-chip" refers to on the same System-on-Chip (SOC) where the processing hardware 110 is located, and the term "off-chip" as used herein refers to outside the SOC. For simplicity of illustration, the memory 120 is shown as a single block; however, it should be understood that the memory 120 may represent an architecture of multiple memory components such as caches, local memory of the accelerator 150, system memory, solid-state or magnetic storage devices, etc. The processing hardware 110 executes instructions stored in the memory 120 to perform operating system functions and run user applications. A Deep Neural Network (DNN) model 125 may be represented by a computational graph including multiple layers, the multiple layers including an input layer, an output layer, and one or more hidden layers therebetween. An example of the DNN model 125 is a multi-layer CNN. The DNN model 125 may have been trained to have weights associated with one or more of the layers.

[0028] In one embodiment, the system 100 may receive input data 123 and store it in the memory 120. The input data 123 may be an image, video, voice, etc. The memory 120 may store instructions, and when the instructions are executed by the accelerator 150, the accelerator 150 performs neural network computations on the input data 123 according to the DNN model 125. The memory 120 may also store compressed data 127, e.g., the output of the neural network computations.

[0029] Figure 2A and Figure 2B illustrates the interaction between the accelerator 150 and the memory 120 according to two alternative embodiments. Figure 2A illustrates that the accelerator 150 may use the compressed data stored in the memory 120 to perform neural network operations. For example, the accelerator 150 may obtain the compressed data from the memory 120, such as input activations and weights for the first CNN layer. The compressor 130 decompresses the obtained data into floating-point numbers for the accelerator 150. The accelerator 150 may store the output of the first CNN layer in the memory 120. Before storing in the memory, the compressor 130 compresses the output of the accelerator from floating-point numbers into a compressed format. In Figure 2BIn this case, the Direct Memory Access (DMA) engine 210 can store compressed data to and retrieve compressed data from the memory 120 for the accelerator 150.

[0030] The floating-point compression performed by the compressor 130 sequentially includes an unpacking stage, a shuffling stage, and a data compression stage. The unpacking stage and the shuffling stage can be collectively referred to as a re-ordering stage. Decompression sequentially includes a data decompression stage, an un-shuffling stage, and a packing stage. The un-shuffling stage and the packing stage can also be collectively referred to as a re-ordering stage. As described below with reference to Figure 3 the compressor 130 transforms data in each stage.

[0031] Figure 3 is a diagram illustrating the transformation operations performed by the compressor according to an embodiment. Also referring to Figure 1 , the compressor 130 can compress the floating-point operands of the CNN layer through a series of unpacking, shuffling, and compression operations. Each floating-point operand is represented by a floating-point representation of three fields (S, E, F), where S is the sign bit, E is the exponent, and F is the fraction. The bit widths of each of E and F can be determined by the neural network application to be executed on the accelerator 150. The compressor 130 does not impose requirements on the bit widths of each of the E and F fields. The compressor 130 can operate on E and F of any bit width.

[0032] In the unpacking stage, the compressor 130 extracts the floating-point fields to combine each of the S, E, and F fields, thereby generating a first sequence composed of sign bits, a second sequence composed of exponents, and a third sequence composed of fractions. The shuffling operation identifies floating-point operands with the same exponent and re-orders these operands so that at least their same exponents are adjacent to each other. Then the compressor 130 compresses the unpacked and re-ordered operands to at least remove duplicate exponents from the sequences. The duplicate exponents are redundant exponents; thus, they can be removed without losing information. In addition, the data compression operation can also remove any duplicate sign bits and / or duplicate fractions from the sequences. The compressor 130 can use any known lossless data compression algorithm (e.g., run-length encoding, LZ compression, Huffman coding, etc.) to reduce the number of bits in the representation of the compressed data sequence.

[0033] In Figure 3In the example, the sequence of floating-point operands includes A0, A1, A2, and A3. All four operands A0, A1, A2, and A3 have the same sign bit (S0) (i.e., S0 = S1 = S2 = S3). A0 and A3 have the same exponent (E0) (i.e., E0 = E3), and A1 and A2 have the same exponent (E1) (i.e., E1 = E2). The compressor 130 unpacks the four operands by forming a first sequence of four consecutive sign bits, a second sequence of four consecutive exponents, and a third sequence of four consecutive fractions. These three sequences can be arranged in any order. The compressor 130 shuffles the operands by moving E0 next to E3 and E1 next to E2. The shuffling operation also moves the corresponding sign bits and fractions; i.e., S0 is next to S3, S1 is next to S2, F0 is next to F3, and F1 is next to F2. Then the compressor 130 removes at least the duplicate exponents (e.g., E2 and E3). The compressor 130 can also remove duplicate sign bits (e.g., S1, S2, and S3). A data compression algorithm can be applied to these three sequences to generate a compressed data sequence. The exponent sequence can be compressed separately from the fraction sequence. Alternatively, the three sequences can be compressed together into one sequence. The compressor 130 generates metadata that records the parameters of the shuffling; e.g., the position or the shift distance of the re-ordered elements (e.g., sign bits, exponents, or fractions) to track the position of each element (e.g., sign bits, exponents, or fractions) in the shuffled sequence. The metadata can also include the parameters of the data compression algorithm; e.g., the compression ratio. The compressed data sequence and the metadata can be stored in the memory 120. In the case where the compressed data sequence exceeds the on-chip memory capacity of the accelerator 150, the compressed data sequence can be stored in off-chip memory (e.g., DRAM).

[0034] The compressor 130 can also perform decompression to convert the compressed data sequence into the original form of floating-point numbers (e.g., A0, A1, A2, and A3 in the example). According to the metadata, the compressor 130 decompresses the compressed data sequence by reversing the above compression operations. In one embodiment, the compressor 130 can decompress the compressed data sequence through a series of data decompression, unshuffling, and packing operations. In the data decompression stage, the compressor 130 decompresses the compressed data sequence based on the data compression algorithm previously used in data compression. The decompression restores the duplicate exponents removed previously and any other duplicate or redundant data elements (e.g., duplicate sign bits or fractions). In the unshuffling stage, the compressor 130 restores the original order of the floating-point numbers represented by the three sequences of sign bits, exponents, and fractions according to the metadata. In the packing stage, the compressor 130 combines the three fields of each floating-point number to output the original sequence of floating-point numbers (e.g., A0, A1, A2, and A3 in the example).

[0035] A neural network layer may have a large number of floating - point operands; for example, several gigabytes of data. Thus, the compressor 130 can process N operands at a time; for example, N can be 256 or 512. Figure 3 An example uses N = 4 for illustrative purposes.

[0036] In an alternative embodiment, the compressor 130 can identify those operands that share the same exponent and remove duplicate exponents without performing a shuffling operation. In yet another embodiment, the compressor 130 can identify those operands that share the same exponent and remove duplicate exponents without the need for unpacking and shuffling operations.

[0037] Figure 4 is a diagram illustrating the data movement in a CNN operation according to an embodiment. For illustrative purposes, the operations of two CNN layers are shown.

[0038] The accelerator 150 begins the operation of CONV1 by fetching the initial input data 425 and the operation parameters 413 of the first CNN layer (i.e., CONV1) from the memory 120 (step S1.1). The accelerator 150 also fetches or requests to fetch the compressed CONV1 weights 411 from the memory 120 (step S1.2). In this embodiment, the initial input data 425 (e.g., an image, video, voice, etc.) is not compressed, while the CONV1 weights 411 (and CONV2 weights 412) have been compressed. In an alternative embodiment, either the initial input data 425 or any of the weights 411, 412 can be compressed or uncompressed. The compressor 130 decompresses the compressed CONV1 weights 411 and passes the CONV1 weights to the accelerator 150 (step S1.3). The decompression can be performed in the three transformation stages described previously Figure 3 (i.e., data decompression, de - shuffling, and packing). The accelerator 150 performs the first CNN layer operation and generates the CONV1 output (step S1.4). The CONV1 output is the input activation for the second CNN layer. The compressor 130 compresses the CONV1 output and generates metadata associated with the compressed CONV1 output 424 (step S1.5). The compression can be performed in the three transformation stages described previously Figure 3 (i.e., unpacking, shuffling, and data compression). Then the compressed CONV1 output 424 and the metadata are stored in the buffer 420 of the memory 120. The compressed CONV1 output 424 is the compressed CONV2 input.

[0039] The accelerator 150 starts the operation of CONV2 by obtaining from or triggering the obtaining of the compressed second CNN layer (i.e., CONV2) input, metadata, the compressed CONV2 weights 412, and the operation parameters 413 for CONV2 from the memory 120 (step S2.1). The compressor 130 decompresses the compressed CONV2 input according to the metadata, decompresses the CONV2 weights 412, and passes the decompressed output to the accelerator 150 (step S2.2). The accelerator 150 performs the second CNN layer operation and generates the CONV2 output (step S2.3). The compressor 130 compresses the CONV2 output and generates metadata associated with the compressed CONV2 output 426 (step S2.4). Then the compressed CONV2 output 426 and the metadata are stored in the buffer 420.

[0040] Figure 4 The example of shows that the data stream moving in and out of the memory 120 mainly includes compressed data. More specifically, the accelerator output activation from one neural network layer is also the compressed accelerator input activation for the next neural network layer. Therefore, in a deep neural network containing multiple layers, since the activation data between layers is transmitted online in the memory in a compressed form, the memory bandwidth can be greatly reduced.

[0041] Figure 5 is a flowchart illustrating a method 500 for compressing floating-point numbers according to an embodiment. The method 500 may be performed by Figure 1 , 2A and / or an embodiment of 2B.

[0042] The method 500 starts at step 510, at which time the system (e.g., Figure 1 system 100) receives a plurality of operands, each of which has a floating-point representation including a sign bit, an exponent, and a fraction. At step 520, the compressor in the system reorders the operands into a first sequence composed of sign bits, a second sequence composed of exponents, and a third sequence composed of fractions. At step 530, the compressor compresses the first sequence, the second sequence, and the third sequence to at least remove duplicate exponents, thereby generating a compressed data sequence losslessly.

[0043] In one embodiment, elements sharing the same exponent, such as multiple sign bits, multiple exponents, and multiple fractions, are reordered to adjacent spatial positions in each of a first sequence, a second sequence, and a third sequence, respectively. In one embodiment, the reordering and compression are performed on N operands in multiple batches, where N is a non-negative integer. For example, sequences A0, A1, A2, and A3 can be considered as one batch of operands, and other operand sequences, such as A4, A5, A6, and A7, can be considered as another batch of operands, and so on. The reordering and compression rules for each other batch of operands are similar to the corresponding operations performed on sequences A0, A1, A2, and A3. After multi-batch processing, the compression of a tensor is generally completed. The compressor generates metadata indicating the parameters used in the reordering and compression. The first sequence can be compressed to remove duplicate sign bits.

[0044] In one embodiment, the operands can include the weights of a CNN layer. The operands can include the activation outputs from an accelerator that executes a CNN layer. The compressor can store the compressed data sequence in a memory and can retrieve the compressed data sequence for decompression for the accelerator to execute subsequent layers of the CNN. The compression of floating-point numbers is applicable to exponents of any bit width.

[0045] Figure 6 is a flowchart illustrating a method 600 for decompressing data into floating-point numbers according to one embodiment. Method 600 can be performed by Figure 1 、 2A and / or embodiments of 2B.

[0046] Method 600 begins at step 610, where a compressor in a system (e.g., Figure 1 system 100) decompresses a compressed data sequence into a first sequence of N sign bits, a second sequence of N exponents, and a third sequence of N fractions (N is a positive integer). The compressed data sequence represents N operands and does not contain duplicate exponents. At step 620, the compressor reorders the first sequence of N sign bits, the second sequence of N exponents, and the third sequence of N fractions into a restored sequence of N floating-point numbers representing N operands. At step 630, the restored sequence of N floating-point numbers is sent to an accelerator (e.g., Figure 1 accelerator 150 in

[0047] for neural network computations. In one embodiment, after decompression at step 610, elements (such as sign bits, exponents, or fractions) sharing the same exponent in each of the first sequence, the second sequence, and the third sequence are located at adjacent spatial positions. Figure 1 、 2A and 2B's exemplary embodiments have been describedFigure 5 and 6 the operations of the flowchart of. However, it should be understood that Figure 5 and 6 the operations of the flowchart of can be performed by Figure 1 , 2A and the embodiments of 2B, and Figure 1 , 2A and the embodiments of the present invention other than those of 2B, and Figure 1 , 2A and the embodiments of 2B can perform operations different from those discussed with reference to the flowchart. Although Figure 5 and 6 the flowchart of shows a specific order of operations performed by certain embodiments of the present invention, it should be understood that this order is exemplary (e.g., alternative embodiments may perform operations in a different order, combine certain operations, overlap certain operations, etc.).

[0048] Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the scope of the present invention. Any person with ordinary knowledge in the technical field to which the present invention pertains, without departing from the spirit and scope of the present invention, may make some modifications and refinements. Therefore, the protection scope of the present invention shall be subject to that defined by the scope of the patent application.

Claims

1. An operand processing system, characterized in that, Comprising: An accelerator circuit; And A compressor circuit coupled to the accelerator circuit, the compressor circuit being configured to: Receive a plurality of operands, each operand having a floating-point representation including a sign bit, an exponent, and a fraction; Reorder the plurality of operands into a first sequence composed of sign bits, a second sequence composed of exponents, and a third sequence composed of fractions, wherein a plurality of sign bits, a plurality of exponents, and a plurality of fractions sharing the same exponent are respectively reordered to adjacent spatial positions in each of the first sequence, the second sequence, and the third sequence; And Compress the first sequence, the second sequence, and the third sequence to at least remove duplicate exponents, thereby generating a compressed data sequence losslessly.

2. The operand processing system according to claim 1, wherein The compressor circuit is further configured to: Perform the reordering and the compression on N operands in multiple batches, where N is a non-negative integer.

3. The operand processing system according to claim 1, wherein When compressing the first sequence, the second sequence, and the third sequence, the compressor circuit is further configured to remove duplicate sign bits.

4. The operand processing system according to claim 1, characterized in that, The compressor circuit is further configured to: Generate metadata indicating parameters used in the reordering and the compression.

5. The operand processing system according to claim 1, characterized in that, The plurality of operands include a plurality of weights of a layer of a convolutional neural network.

6. The operand processing system according to claim 1, wherein The plurality of operands include activation outputs from the accelerator circuit executing a layer of a convolutional neural network.

7. The operand processing system according to claim 6, wherein Further comprising: A memory, wherein the compressor circuit is further configured to: Store the compressed data sequence in the memory; And Retrieve the compressed data sequence for decompression for the accelerator circuit to execute a subsequent layer of the convolutional neural network.

8. The operand processing system according to claim 1, wherein The compressor circuit is further configured to: Decompress the compressed data sequence to losslessly restore the plurality of operands in floating-point format.

Citation Information

Patent Citations

  • Data compression and decompression method and equipment

    CN102457283A

  • Lossless index and lossy mantissa weight compression for training deep neural networks

    CN114341882A

  • Method and System for Efficient Floating-Point Compression

    US20200225948A1