Method and data processing system
By separating Golomb-Rice codes into variable-length and fixed-length parts and using control rules to manage the size of the first part's sub-stream within processed streams, the method addresses the inefficiencies in decoding variable-length codes, enhancing processor performance.
Patent Information
- Application Number
- JP2021017501
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-04
- Filing Date
- 2021-02-05
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2041-02-05
AI Technical Summary
Golomb-Rice codes, used for reversible compression, are variable-length codes that can be slow to decode due to their serial dependency and variable unary part size, leading to inefficiencies in processor decoding.
A method is introduced where variable-length codes are separated into a first variable-length part and a second part, and for each chunk of a processed stream containing data from the first part, a set of control rules is used to form a sub-stream with a predictable size, enabling more efficient processing.
This approach allows for more efficient decoding of variable-length codes by predicting the size of the first part's sub-stream, thereby improving processor performance and reducing decoding time.
Smart Images

Figure 0007688981000001 
Figure 0007688981000002 
Figure 0007688981000003
Abstract
Description
Technical Field
[0001]
[0001] The present invention relates to a data processing system, and more particularly to a data processing system that processes variable-length codes.
Background Art
[0002]
[0002] One way to perform reversible compression data compression known to those skilled in the art is to convert values to Golomb-Rice codes. To convert a numerical value to a Golomb-Rice code, a parameter known as the divisor is selected. To generate a Golomb-Rice code, the numerical value is divided by the divisor to generate two parts. The first part is the quotient of how many times the divisor divides the numerical value completely. The second part is the remainder, which is the remaining number if any after the divisor divides the numerical value completely.
[0003]
[0003] An example of a Golomb-Rice code is shown in FIG. 1. In the example shown in FIG. 1, values from 0 to 10 are shown as Golomb-Rice codes. The first part of the Golomb-Rice code, the quotient, is represented in a unary format. In this format, the number is represented by the number of "1"s equal to the value of the quotient, followed by a stop bit which is "0". For example, the unary part of the number 9 is "1110" followed by three "1"s because 3 divides 9 three times, followed by the stop bit "0". The second part of the Golomb-Rice code is a fixed-length binary part. Since the divisor in this example is "3", the remainder can only be 0, 1, or 2. Therefore, this can be represented by a 2-bit fixed-length binary. The last 2 bits in each of the Golomb-Rice codes represent the remainder in binary form. The remainder is sometimes referred to as the "mantissa" of the Golomb-Rice code because it appears after the stop bit of the unary part of the Golomb-Rice code.
[0004]
[0004] Since the size of the unary part of the Golomb-Rice code changes, the Golomb-Rice code is a type of variable-length code. Such variable-length codes can be slow to decode in a processor because each code has to be considered separately for decoding.
Summary of the Invention
[0005]
[0005] According to a first aspect, a method for use by a processing element is provided, the method comprising: obtaining a plurality of variable-length codes, each variable-length code having a first variable-length part and a second part; separating the variable-length code into the first variable-length part and the second part of the variable-length code; and for each chunk of a processed stream containing data from the first part of the variable-length code, using a set of control rules to form a sub-stream within the chunk of the processed stream such that the data from the first part has a size determined according to the control rules.
[0006]
[0006] According to a second aspect, a method for decoding a processed data stream is provided, the method comprising: obtaining a processed data stream containing data related to a plurality of variable-length codes, each variable-length code having a first variable-length part and a second part, the processed stream being formed within chunks, at least one chunk of the processed stream containing a sub-stream of data formed from the first part of the variable-length code; and using a set of flow control rules to identify and extract the sub-stream within the chunk of the processed data stream.
[0007]
[0007] According to a third aspect, a data processing system is provided that includes a processing element and a storage device. The storage device stores a code portion that, when executed by the processing element, causes the data processing system to obtain a plurality of variable-length codes. Each variable-length code has a variable-length first portion and a second portion. The variable-length code is separated into the first portion of the variable-length code and the second portion of the variable-length code. For each chunk of the processed stream that contains data from the first portion of the variable-length code, a set of control rules is used to form a sub-stream within the chunk of the processed stream such that the data from the first portion forms a sub-stream within the chunk of the processed stream having a size determined according to the control rules.
[0008]
[0008] The present technology will be further described by referring only to the examples as shown in the accompanying drawings.
Brief Description of the Drawings
[0009]
Figure 1
[0009] FIG. 1 is a table showing Golomb-Rice codes for values from 0 to 10 using a divisor of 3.
[0010]
Figure 2
[0010] FIG. 2 shows the components of a neural processing unit that writes activation data to a storage device.
[0011]
Figure 3
[0011] FIG. 3 is a flowchart showing processing a stream of Golomb-Rice codes for storage.
[0012]
Figure 4
[0012] FIG. 4 shows the structure of a data stream during separation of Golomb-Rice codes.
[0013]
Figure 5
[0013] FIG. 5 shows a plurality of chunk structures used by an encoder and a decoder.
[0014]
Figure 6
[0014] Figure 6 is a flowchart showing the steps executed by the encoder.
[0015]
Figure 7
[0015] Figure 7 is a flowchart showing the steps of determining the chunk structure when decoding the first chunk of the stream encoded by the encoder.
[0016]
Figure 8
[0016] Figure 8 is a flowchart showing the steps of determining the chunk structure when decoding subsequent chunks of the stream encoded by the encoder.
[0017]
Figure 9a
[0017] Figure 9a is a flowchart showing the steps of decoding single-item data.
[0018]
Figure 9b
[0018] Figure 9b shows the bit processing when decoding a set of single-item values.
[0019]
Figure 10a
[0019] Figure 10a shows a mobile device.
[0020]
Figure 10b
[0020] Figure 10b is a diagram showing the hardware of the mobile device.
[0021]
Figure 11
[0021] Figure 11 is a diagram showing the system architecture installed on the mobile device.
[0022]
Figure 12
[0022] Figure 12 is a diagram showing the components of the neural processing unit.
[0023]
Figure 13
[0023] Figure 13 is a table showing a flow control rule for encoding single-term data related to weight values.
[0024]
Figure 14
[0024] Figure 14 is a table showing a flow control rule for decoding single-term data related to weight values.
DETAILED DESCRIPTION OF THE INVENTION
[0025]
[0025] Before discussing the embodiments with reference to the accompanying drawings, the following description of the embodiments and related advantages is provided.
[0026]
[0026] According to one embodiment, a method for use by a processing element is provided, the method comprising: obtaining a plurality of variable-length codes, each variable-length code having a first variable-length portion and a second portion; separating the variable-length code into the first variable-length portion of the variable-length code and the second portion of the variable-length code; for each chunk of a processed stream containing data from the first portion of the variable-length code, using a set of control rules to form a sub-stream within the chunk of the processed stream such that the data from the first portion forms a sub-stream within the chunk of the processed stream having a size determined according to the control rules. In this way, the processed stream can have a sub-stream of the first portion of the variable-length code that is predictable from the control rules. This can enable more efficient processing of the first portion of the variable-length code by a processor.
[0027]
[0027] The first part of the variable-length code may be a single-term part, and the second part of the variable-length code may be a mantissa part. In some cases, the mantissa part of the variable-length code is a truncated binary part. The single-term part of each variable-length code may represent the quotient of the original value represented by the variable-length code. Further, in some cases, the mantissa part of each variable-length code is a fixed-length binary code representing the remainder of the original value represented by the variable-length code. In some implementations, the variable-length code is a Golomb-Rice code.
[0028]
[0028] Each chunk of the processed stream may include data from the first part of the variable-length code and data from the second part of the variable-length code, may include data from the first part of the variable-length code, and may not include data from the second part of the variable-length code, or may include data from the second part of the variable-length code and may not include data from the first part of the variable-length code.
[0029]
[0029] The control rules may be configured to serially determine the size of the single-term sub-stream within each chunk. In some cases, the control rules are configured to determine the size of the single-term sub-stream within each chunk based on the amount of the single-term part of the variable-length code remaining without being added to the processed stream. For some implementations, the processed stream is formed from cells of a predetermined length, and the measurement of the single-term part of the variable-length code remaining without being added to the processed stream is the number of bits of the single-term part of the variable-length code associated with the cells remaining without being added to the processed stream. By processing the data within the cells, the flow control rules can enable the size of the single-term sub-stream to be predicted by the decoder.
[0030] The control rule may be a rule for selecting among a set of predefined chunk structures. Each chunk structure may define the type of data that should be at each position within the chunk. The set of chunk structures may include at least one chunk structure for a first chunk within a cell. The at least one chunk structure for the first chunk within a cell may include a header portion that includes information regarding the data included in the cell.
[0031]
[0031] In some ways according to the first embodiment, the processed stream is formed from cells of a predefined length, and each cell of the processed stream has a header and a plurality of chunks, and the header indicates the length of the cell and the length of the unary portion of the variable-length code within the cell. This can enable the decoder to track the amount of unary data remaining in the cell during decoding.
[0032]
[0032] The first embodiment may further include obtaining a second plurality of variable-length codes and separating the second plurality of variable-length codes into a first portion of the variable-length code and a second portion of the variable-length code, and the step of forming the processed stream within the chunk uses a set of control rules such that for each chunk of the processed stream including data from at least one first portion of the first plurality of variable-length codes and the second plurality of variable-length codes, the flow control rule determines that the number of first portions of the first plurality of variable-length codes and the number of first portions of the second plurality of variable-length codes are included in the chunk.
[0033]
[0033] In some implementations, forming the processed stream includes maintaining a balance value that records the difference between the number of first portions of the first plurality of variable-length codes included in the chunk and the number of first portions of the second plurality of variable-length codes included in the chunk, and the flow control rule determines the number of first portions of the first plurality of variable-length codes and the number of first portions of the second plurality of variable lengths included in the chunk based on the balance value.
[0034]
[0034] In some applications, a plurality of variable-length codes represent weight values for use in a neural network. In other applications, the plurality of variable-length codes represent values within an activation stream that is the output of a layer of a neural network.
[0035]
[0035] According to a second embodiment, a method for decoding a processed data stream is provided. The method includes obtaining a processed data stream that includes data associated with a plurality of variable-length codes, each variable-length code having a first variable-length portion and a second portion, the processed stream being formed in chunks, at least one chunk of the processed stream including a sub-stream of data formed from the first portion of the variable-length code, and using a set of flow control rules to identify and extract the sub-stream within a chunk of the processed data stream.
[0036]
[0036] In some implementations, each first portion of the variable-length code is encoded as unary data having a stop bit. The method may further include obtaining a plurality of identified and extracted sub-streams from a plurality of chunks of processed data, converting each extracted sub-stream into an intermediate format representing a list of bit positions of stop bits within the sub-stream, combining the plurality of converted sub-streams to form an extended list of bit positions, and measuring distances between stop bits within the extended list of bit positions to recover the value of the first portion of the variable-length code within the first stream.
[0037]
[0037] In some embodiments, the flow control rule is configured to determine the size of the unary sub-stream within each chunk to be decoded based on the amount of the unary portion of the variable-length code remaining without being decoded from the processed stream. The processed stream may be formed from cells of a predetermined length, and the amount of the unary portion of the variable-length code remaining without being decoded from the processed stream may be the number of bits of the unary portion of the variable-length code associated with the cells remaining without being decoded from the processed stream.
[0038]
[0038] In other embodiments, the processed stream may additionally include a first portion of the variable-length code belonging to a second plurality of variable-length codes. The flow control rule is configured to determine, for each chunk of the processed stream to be decoded, which includes data of at least one first portion of the plurality of variable-length codes and the second plurality of variable-length codes, whether the number of the first portions of the plurality of variable-length codes and the number of the first portions of the second plurality of variable-length codes are included in the chunk to be decoded.
[0039]
[0039] Identifying and extracting the sub-stream within the chunk of the processed stream may include maintaining a balance value that records the difference between the number of the first portions of the plurality of variable-length codes extracted from the processed stream and the number of the first portions of the second plurality of variable-length codes extracted from the processed stream. The flow control rule may be used to determine the number of the first portions of the plurality of variable-length codes and the number of the first portions of the second plurality of variable-length codes to be extracted from the chunk based on the balance value.
[0040]
[0040] According to a third embodiment, a non-transitory computer-readable storage medium storing a code portion is provided, the code portion, when executed on a processing element, causes the processing element to perform a method of obtaining a plurality of variable-length codes, each variable-length code having a variable-length first portion and a second portion, obtaining, separating the variable-length code into the first portion of the variable-length code and the second portion of the variable-length code, and for each chunk of a processed stream containing data from the first portion of the variable-length code, forming a sub-stream within the chunk of the processed stream having a size determined according to a control rule using a set of control rules such that the data from the first portion forms the sub-stream within the chunk of the processed stream.
[0041]
[0041] According to a fourth embodiment, a non-transitory computer-readable storage medium storing a code portion is provided, the code portion, when executed on a processing element, causes the processing element to perform a method of decoding a processed data stream, the method including obtaining a processed data stream containing data associated with a plurality of variable-length codes, each variable-length code having a variable-length first portion and a second portion, the processed stream being formed in chunks, at least one chunk of the processed stream including a sub-stream of data formed from the first portion of the variable-length code, and identifying and extracting a sub-stream within a chunk of the processed data stream using a set of flow control rules.
[0042]
[0042] According to a fifth embodiment, a data processing system including a processing element and a storage device is provided. The storage device stores a coded portion, which, when executed by the processing element, causes the data processing system to obtain a plurality of variable-length codes. Each variable-length code has a variable-length first portion and a second portion. The variable-length code is separated into the first portion of the variable-length code and the second portion of the variable-length code. For each chunk of the processed stream containing data from the first portion of the variable-length code, a set of control rules is used to form a sub-stream within the chunk of the processed stream such that the data from the first portion forms a sub-stream within the chunk of the processed stream having a size determined according to the control rules.
[0043]
[0043] According to a sixth embodiment, a data processing system including a processing element and a storage device is provided. The storage device stores a coded portion, which, when executed by the processing element, causes the data processing system to obtain a processed data stream including data related to a plurality of variable-length codes. Each variable-length code includes a variable-length first portion and a second portion. The processed stream is formed within chunks, and at least one chunk of the processed stream includes a sub-stream of data formed from the first portion of the variable-length code. A set of flow control rules is used to identify and extract the sub-stream within the chunk of the processed data stream.
[0044]
[0044] Here, specific embodiments are described with reference to the drawings.
[0045]
[0045] FIG. 2 shows a part of component 2 of a neural processing unit (NPU), not all of it. Component 2 is a special chip that executes calculations related to artificial intelligence applications, particularly calculations related to neural networks. In other words, the NPU enables hardware acceleration of specific calculations related to neural networks. Component 2 is a component that writes activation values to an external DRAM (not shown) of the NPU.
[0046]
[0046] When performing calculations related to a neural network, the calculations may be performed for each layer of the neural network. Those calculations generate an output known as activation data, which may be large in volume and needs to be stored before further calculations can be performed using that data. Storing activation data in memory and retrieving activation data from memory can be relatively slow processes due to the constraints on data transfer from external memory to the processor. Therefore, in order to improve processor performance, it is desirable to compress the data from the activation layer using a Golomb-Rice code.
[0047]
[0047] Component 2 is configured to process activation data for storage. The activation data is received and grouped into tiles of data. The tiles of data are defined as 8×8 groups of elements, where the elements are 8-bit uncompressed activation data. Processing elements in the form of encoder 20 are configured to compress the received activation data by converting it to a Golomb-Rice code. Further steps described below are then performed to make it easier to decode the activation data.
[0048]
[0048] When decoding a variable-length code such as a Golomb-Rice code, it is difficult to parse at a high rate. This is because there is a serial dependency between Golomb-Rice codes such that the length of the preceding Golomb-Rice code needs to be known before the next Golomb-Rice code can be identified and decoded. Therefore, a typical hardware implementation for decoding a Golomb-Rice code can achieve a rate of one or two Golomb-Rice codes per clock cycle when directly parsing using a single parser.
[0049]
[0049] The technique described in the first specific embodiment takes a different approach. FIG. 3 is a flowchart showing the steps performed by an encoder 20 that receives uncompressed activation data. In step S30, in this case, a Golomb-Rice code is obtained by conversion by the encoder 20. Next, in step S31, the encoder 20 separates the Golomb-Rice code into a stream of unary values and a stream of remainder values, and stores them in the RAM 21 shown in FIG. 2.
[0050]
[0050] FIG. 4 shows three data streams. The source data stream 40 is a stream of Golomb-Rice codes. The source data stream 40 includes a series of Golomb-Rice codes indicated by values GR1 to GR5. Each Golomb-Rice code has a variable-length unary part and a fixed-length binary part of the type described in the explanation of the related art. A 3-bit fixed-length binary part is shown in FIG. 3, but the length of the binary part is not important, and other lengths may be used. The encoder 20 splits the Golomb-Rice code into two parts to generate two further streams 41 and 42 shown in FIG. 4. The first stream 41 is a unary stream, and the second stream 42 is a remainder stream, and each binary has a fixed length.
[0051]
[0051] In step S32, the stitch processor 22 shown in FIG. 2 stitches the first stream and the second stream together to form a processed stream. This is done for each cell, and each cell represents 32 tiles (2,048 elements) of uncompressed data stored in a 2112-byte slot. To allow for some overhead and round up to a total of 64 bytes, the slot is larger than the cell.
[0052]
[0052] Each cell is formed within a plurality of chunks by the stitch processor 22. The first chunk of a cell always includes a header. The unary data from stream 4 is always stitched by the stitch processor 22 into the chunks of the processed stream in 32-bit portions.
[0053]
[0053] As described herein with reference to FIGS. 5 and 6, cells are formed by the stitch processor 22 using a set of flow control rules. FIG. 5 shows different structures of chunks that can be used by the stitch processor 22 to form cells, and FIG. 6 shows the steps that would be performed by the stitch processor 22 when forming cells.
[0054]
[0054] As mentioned above, the first chunk of a cell needs to include a header, and the header provides information regarding the length of the cell and the length of the unary sub-stream contained within the cell. The length of the remainder value within the cell is not included in the header, but can be derived from the length of the cell and the length of the unary sub-stream.
[0055]
[0055] FIG. 5 shows available chunk formats designed to allow each chunk to be consumed in a single clock cycle of a decoder that decodes Golomb-Rice codes for parsing efficiency. The shown chunk structures are divided into two categories. The top two chunk structures 51 and 52 shown in FIG. 5 are the first chunk structures for a cell and may be selected for use when forming the first chunk of a cell. Both chunk format structures include a 32-bit long header portion. Below the first chunk structures 51 and 52 in FIG. 5, the next five chunk structures 53 - 57 are used after the first chunk within the cell has been emitted by the encoder 20 to include the remaining unary data and remainder data from the first and second streams corresponding to 32 tiles of uncompressed data that will be included in the cell.
[0056]
[0056] To select an appropriate chunk structure to use, the encoder 20 uses a set of flow control rules. When a chunk structure for the next chunk is identified, the chunk structure may be populated or emitted with appropriate data. The flow control rules used by the stitch processor 22 are as follows. When selecting the first chunk structure for a cell, the first chunk structure 51 shown in FIG. 5 is selected when there are more bits in the cell available to the stitch processor 22 for inclusion in the chunk than 32 bits of unary data. Otherwise, the chunk structure 52 is used to form the first chunk of the cell because the chunk 52 does not require any unary data. A situation where 32 bits of unary data are not available may occur when none of the uncompressed activation data for the cell has a unary portion, i.e., each value is less than the divisor used to generate a Golomb-Rice code. In this case, the encoder 20 does not encode the stop bit as unary data.
[0057] After the first chunk is emitted, if there are more bits of the remaining unary data for inclusion in the cell than 128 bits for the subsequent chunks in the cell emitted by the encoder 20, the chunk structure 53 is used. Since the unary data is included in the chunks within the 32-bit portion, ultimately, there are less than 128 bits of the remaining unary data that remain unencoded for the cell. If there are 96 bits of the remaining unary data that remain unencoded for the cell, then the chunk structure 54 is used, if there are 64 bits of the remaining unary data that remain unencoded for the cell, then the chunk structure 55 is used, if there are 32 bits of the remaining unary data that remain unencoded for the cell, then the chunk structure 56 is used. In the case where all of the unary data of the cell is encoded, then the chunk structure 57 is used to emit the remaining data. It should be noted that the unary data included in the above chunk structures is selected regardless of the grouping of the tiles and elements below the activation data below, so that unary data from different tiles and / or elements can be included in the same chunk.
[0058]
[0058] The above method is shown in FIG. 6. In step S60, the first chunk structure is selected from the chunk structures 51 and 52 shown in FIG. 5. This selection depends on the availability of 32 bits of unary data as described above. After selecting the chunk structure, the stitch processor 22 generates a header portion. The stitch processor 22 evaluates the length of the unary sub-stream that will be included in the cell based on 32 tiles of uncompressed data and adds information indicating that length to the header portion. The length of all the data that will be included in the cell is also evaluated and added to the header portion. If required, the data from the first stream of unary data 41 and the data from the second stream of remaining data are added to the chunks according to the selected chunk structure selected by the stitch processor 22.
[0059]
[0059] In step S61, by selecting an appropriate chunk structure from the chunk structures 53 to 57 according to the flow control rules described above, the next chunk of the processed data stream is formed. After selecting the chunk structure, the chunk is formed by filling the relevant part of the chunk structure with the data from the first stream of the single-term data 41 and the data from the second stream of the remainder data.
[0060]
[0060] In step S62, the stitch processor 22 determines whether more data is to be formed into chunks. If more data is to be formed into chunks, the method proceeds to S61 to form the next chunk. If no more data is to be processed, the method proceeds to S63 and ends.
[0061]
[0061] The method described above assumes that it is possible to form a complete 128-bit chunk from the 32 tiles within a cell and to stitch the single-term part into 32-bit parts. In practice, those conditions may not be met in cases where the first stream of the single-term part and the second stream of the remainder part are padded using the stop bit "0" for padding until they reach the desired size. As the length of the data added to the cell is stored in the header, it is possible to identify the length of the data within the cell and the location where padding starts when decoding the processed data stream.
[0062]
[0062] Next, a method of decoding the activation data stored by the decoder will be described with reference to FIGS. 7 and 8. In this case, the decoder is part of the NPU that enables the activation data to be read from the DRAM for use in further calculations. The decoder stores a replica of the chunk structure shown in FIG. 5 that was used by the encoder 20 to store the processed stream in the RAM. In step S70 of FIG. 7, the decoder receives the first chunk of activation data cells from the RAM for decoding. The decoder reads the header and identifies the length of the unary data within the cell. In step S71, the decoder identifies whether the length of the unary data specified in the header is 32 bits or more. If the unary length is 32 bits or more, then the first chunk of the cell may be formed according to the chunk structure 51, and the unary data and the remainder data may be extracted from the chunk according to the known chunk structure. If the length of the unary data identified in the header is less than 32 bits (being zero because the unary data is stitched in the 32-bit portion), the first chunk is formed according to the chunk structure 52, and accordingly, the first chunk is decoded. When decoding the data, the decoder maintains the parameter U_left, which is initially set to the value of the unary length within the cell when the cell header is examined, and is updated each time unary data is taken out of the chunk to record the amount of unary data remaining within the cell. Thus, when the chunk structure 51 is used for the first chunk, then after extracting 32 bits of unary data from the first chunk, the parameter U_left decreases by 32 bits.
[0063]
[0063] Figure 8 shows a method used by a decoder to determine the chunk structure for each subsequent chunk. In step S80, a subsequent chunk of the stored activation data is received. In step S81, the parameter U_left is examined to determine whether the amount of unary data remaining without being extracted for a cell is 128 bits or more. If the amount of unary data to be extracted is 128 bits or more, the decoder determines that chunk structure 53 is used. In step S82, the decoder extracts data from the chunk according to chunk structure 53 and updates the parameter U_left to constitute the amount of the extracted unary data.
[0064]
[0064] If less than 128 bits of unary data remain without being extracted, the method proceeds to step S83. In step S83, the parameter U_left is examined to determine whether the amount of unary data remaining without being extracted for a cell is equal to 96 bits. If the amount of unary data to be extracted is equal to 96 bits, the decoder determines that chunk structure 54 is used. In step S84, the decoder extracts data from the chunk according to chunk structure 54 and updates the parameter U_left to constitute the amount of the extracted unary data.
[0065]
[0065] If more than 96 bits of unary data remain without being extracted, the method proceeds to step S85. In step S85, the parameter U_left is examined to determine whether the amount of unary data remaining without being extracted for a cell is equal to 64 bits. If the amount of unary data to be extracted is equal to 64 bits, the decoder determines that chunk structure 55 is used. In step S86, the decoder extracts data from the chunk according to chunk structure 55 and updates the parameter U_left to constitute the amount of the extracted unary data.
[0066]
[0066] If bits less than 64 bits of the single-item data are not extracted and remain, the method proceeds to step S87. In step S87, the parameter U_left is inspected to determine whether the amount of single-item data remaining without being extracted for the cell is equal to 32 bits. If the amount of single-item data to be extracted is equal to 32 bits, the decoder determines that the chunk structure 56 is used. In step S86, the decoder extracts data from the chunk according to the chunk structure 56 and updates the parameter U_left to constitute the amount of the extracted single-item data. If the amount of single-item data to be extracted is not equal to 32 bits (equal to zero), the decoder determines that the chunk structure 57 is used. In step S89, the decoder extracts data from the chunk according to the chunk structure 57.
[0067]
[0067] Based on the processes described above in connection with FIGS. 7 and 8, the decoder can efficiently regenerate the first stream of the single-item data 41 and the second stream of the remainder data 42 from the processed data stream stored in the RAM by the encoder 20. By using the flow control rules shown in FIGS. 7 and 8, the decoder can determine the type of data found at any point in the incoming stream without the bit cost indicating the data type within the incoming compressed data stream.
[0068]
[0068] After extracting the single-item data and the remainder data from the processed data stream, the decoder needs to decode the Golomb-Rice code to regenerate the uncompressed activation data. The second stream 42 of the remainder data is an array of fixed-length binary values and is simple to decode using techniques known in the art. Therefore, this process is not further discussed here.
[0069]
[0069] Here, decoding the first stream of the single-item data 41 is described in connection with FIG. 9a. In step S90, the 8-bit block of the single-item data is converted into a binary format, and the binary format indicates the position or positions of the stop bits within the binary block. This is done by using a look-up table. In step S91, since the single-item code may span across 8-bit blocks, four 8-bit blocks are combined into a 32-bit block, and then four 32-bit blocks are combined into a 128-bit block. The 128-bit block is again a list of the positions of the stop bit positions. To extract the single-item value, in step S92, the difference between the values of each adjacent stop bit position giving the value of the single-item code is taken.
[0070] The method of FIG. 9a can be achieved by using a look-up table to analyze the 8-bit blocks of the unary data into an intermediate form. This is shown in FIG. 9b where the top row 94 identifies the bit positions within each byte shown below it. The first bit in each byte is '0' and the last bit is '7'. The second row 95 shows the bytes of the unary data. Recall that the stop bit in the unary data is '0'. In the intermediate form shown in the third row 96, each byte is extended into a list of up to 3-bit codes indicating the position of the stop bit within the byte. A Radix-4 combination of four 8-bit segments into a 32-bit segment is performed and shown in the fourth row 97 and fifth row 98 of FIG. 9b. In the fourth row 97, pairs of bits to identify are added as the most significant bit (MSB) to the code. For the first byte, the value '00' is added to the 3-bit code, for the third byte, the value '10' is added to the 3-bit code, and so on. In the fifth row, 5-bit codes are concatenated to form a list of stop bit positions within the 32-bit word. A subsequent Radix-4 combination of four 32-bit segments into a 128-bit segment uses a similar process to generate a list of 7-bit codes indicating the position of the unary stop bits. As in step S92, subtraction of adjacent values yields the length of the unary data and thus the value of the unary data.
[0071] As described above, the first particular embodiment combines the unary and remainder portions of the Golomb-Rice code within a cell. Each cell may include both a unary portion and a remainder portion. Mixing the unary and remainder portions of the activation data within this cell has the advantage of spreading the unary and remainder portions across the processed data stream retrieved from the DRAM. This makes it possible to reduce the size of the parsing buffer in the decoder for storing the unary data prior to decoding, thereby reducing the hardware requirements.
[0072]
[0072] In the first specific embodiment, the compression of activation data was discussed. In the second specific embodiment, the technique is applied to the compression of weight values. FIG. 10a shows a mobile device 10 of the second specific embodiment. Although the mobile device 10 is described herein, the techniques described are applicable to any type of computing device that extracts weight values associated with a neural network, including but not limited to, tablet computers, laptop computers, personal computers (PCs), servers, and the like. FIG. 10b shows the hardware of the mobile device 10. The mobile device 10 includes a processing element in the form of a CPU 100 and a special processor 101 in the form of a neural processing unit (NPU). The mobile device 10 further includes a storage device in the form of a random access memory (RAM) 102. Although not shown in FIG. 10b, additional non-volatile storage devices are also provided. The mobile device 10 includes a display 103 for displaying information to the user, and a communication system 104 that enables the mobile device 10 to connect to transfer and receive data through various data networks using technologies such as Wi-Fi (trademark) and LTE (trademark).
[0073]
[0073] FIG. 11 shows the system architecture installed on the mobile device 10 associated with the NPU 101. The system architecture enables a software application 110 to access the NPU 101 for hardware acceleration of computations related to neural networks. The system architecture is an Android (registered trademark) software architecture for use on, for example, a mobile phone or a tablet computer.
[0074]
[0074] For the hardware acceleration of specific processes related to neural network processing, a software application 110 that utilizes a machine learning library 111 has been developed. A runtime environment 112, known as the Android (registered trademark) Neural Network Runtime, is provided under the library to receive instructions and data from the application 110. The runtime environment 112 is an intermediate layer that is involved in the communication between the software application 110 and the NPU 101, and the scheduling of execution tasks for the most appropriate hardware. Immediately below the runtime environment 112, at least one processor driver and a related special processor, in this case the NPU 101, are provided. A plurality of processors and related drivers, such as a digital signal processor, a neural network processor, and a graphics processor (GPU), may be provided immediately below the runtime environment 112. However, to avoid redundant explanations, only the NPU 101 and the related processor driver 113 are described in relation to the second specific embodiment.
[0075]
[0075] FIG. 12 shows the partial components of the NPU 101. The NPU 101 includes a weight decoder 120 connected to a direct memory access component 121 that handles data transfer on the external interface to the RAM 102 of the mobile device 10. The decoded values from the weight decoder 120 are sent to a multiplier accumulator unit 122 for subsequent processing by the NPU 101.
[0076]
[0076] In the second specific embodiment, the processor driver 113 stores the weight values in the RAM 102 according to a chunk structure determined by a set of flow control rules. Subsequently, the direct memory access component 121 retrieves the weight values from the RAM 102, and the weight decoder 120 extracts data from the chunk structure.
[0077]
[0077] The processor driver 113 obtains a set of uncompressed (untreated) weight values for the neural network. The source of the uncompressed weight values is not considered for the purposes of the technology discussed herein. However, in one embodiment, the uncompressed weight values may be provided to the Android Neural Network Runtime by the application 110.
[0078]
[0078] The weight values are received in an uncompressed format such as binary data. The first step performed by the processor driver 113 is zero run coding. Zero run coding has advantages when weight values having the value 0 are frequent in the weight stream. For a sequence of weight values including n non-zero weight values, an array of weight values (weight_values) is formed by the processor driver 113 as the sequence of non-zero weight values. The processor driver 113 also identifies an array of zero run lengths (zruns) between the non-zero weight values. The array of zero runs has a length of n + 1. In the sequence of zruns, zruns[0] is the first zero run length and zruns[n] is the last zero run.
[0079]
[0079] For example, consider the following weight sequence: 0, 5, 6, 0, 0, 0, 7, 0. The processor driver 113 codes n = 3 because there are three non-zero values. The sequence of weight values is weight_values = {5, 6, 7}, and the sequence of zero runs is zruns = {1, 0, 3, 1}. From these two sequences, the original weight sequence may be reconstructed. In this way, the processor driver 113 separates the incoming stream of weight values into a sequence of weight_values and a sequence of zruns.
[0080]
[0080] The weight_values are converted to Golomb-Rice codes using a first divisor, and the zruns are converted to Golomb-Rice codes using a second divisor. In the same manner as in the first particular embodiment, the Golomb-Rice codes are separated by the processor driver 113 into a unary stream and a remainder stream. Thus, the processor driver 113 generates four different data streams that will be included in chunks for storage in the RAM 102, the unary part of weight_values (wunary), the remainder part of weight_values (wremain), the unary part of zruns (zunary), and the remainder part of zruns (zremain). Those different data types are added to the chunk structure using a set of flow control rules, as described below.
[0081]
[0081] The weight values received by the processor driver 113 are coded in a slice having a slice header at the beginning of each slice. The slice header contains information regarding the divisor used to generate the weight_value Golomb-Rice code and the divisor used to generate the zrun Golomb-Rice code. The slice header also contains information regarding the length of the slice. After the slice header, several chunks follow that code the different data types mentioned above. The chunks alternate between containing unary values and containing remainder values, and the remainder values (wremain and zremain) are included in a chunk after the chunk containing the corresponding unary values (wunary and zunary). Each chunk that codes unary values (wunary and zunary) codes the maximum of 12 weight symbols and 12 zero run symbols (each symbol corresponding to the first part of the Golomb-Rice code). The reason the number varies is that the unary value corresponding to a symbol is a variable-length value, and fewer unary values can be coded when the length of the symbol is long compared to the case where the length of the symbol is short. The length of the unary chunk has the maximum predefined value selected based on the characteristics of the decoder. Following the unary data, the chunks containing remainder values are variable-length chunks because they contain the remainder values corresponding to the unary symbols contained in the chunks that precede them, as mentioned above.
[0082]
[0082] For a chunk containing unary data, the addition of unary values is controlled by recording a balance that is the value obtained by subtracting the number of zunary values added to the chunks in the slice so far from the number of wunary values added to the chunks in the slice so far. If the balance is 8 or more, then only zunary values are included in the next unary chunk. If the balance is less than 0, then only wunary values are included in the next unary chunk.
[0083]
[0083] To form the first chunk containing unary data, up to 12 symbols valued as wunary are added to the chunk as long as the maximum chunk size is not exceeded. The remainder of the chunk is then filled with the zunary value corresponding to the zrun symbol. The second chunk contains the remainder values (wremain and zremain) associated with the unary part of the symbols added to the first chunk. When forming the third chunk, there are three possibilities. First, if the balance described above is between 0 and 7, then the same process for the first chunk continues with up to 12 symbols of the wunary data added to the chunk, followed by filling with the zunary value corresponding to the zrun symbol. Second, if the balance is 8 or more, then the chunk is filled with zunary values only. This allows the zunary values to catch up when more wunary values are encoded. Third, if the balance is less than 0, then the chunk is filled with wunary values only. This allows the wunary values to catch up when more zunary values are encoded. The fourth chunk contains the remainder values corresponding to the symbols included in the third chunk. The logic of the flow control rules described above is shown in FIG. 13. This process continues until all the data in the slice is encoded.
[0084]
[0084] Subsequently, the weight data is retrieved by the direct memory access component 121 for use in the weight decoder 120. The weight decoder 120 identifies the extracted values from the stream of processed data retrieved by the direct memory access component 121 as follows. The logic for decoding the chunk is shown in FIG. 14. Similar to the encoding described previously, the balance value is maintained by the weight decoder 120. The balance value in the weight decoder 120 is the value obtained by subtracting the number of zunary values extracted from the chunks in the slice so far from the number of wunary values extracted from the chunks in the slice so far.
[0085]
[0085] In the first chunk to be decoded, the weight decoder 120 starts extracting the wunary value either until all data from the chunk are extracted or until 12 symbols valued as wunary values are extracted. Subsequent values within the first chunk are extracted as zunary values.
[0086]
[0086] As discussed in relation to encoding, the chunks following the unary chunk contain the remainder values corresponding to the symbols within the preceding chunk. The number of symbols valued as wunary values within the preceding chunk is known. Thus, the same number of wremain values corresponding to the wunary values within the preceding chunk are extracted, and any subsequent values are extracted as zremain values. For subsequent unary chunks, the balance is inspected. If the balance is less than zero, the weight decoder 120 extracts all unary values as wunary values. If the balance is 8 or greater, the weight decoder extracts all values as zunary values. If the balance is between 0 and 7, the first unary data of up to 12 symbols are extracted as wunary values, and any subsequent values are extracted as zunary values.
[0087]
[0087] In this way, the weight decoder 120 extracts both the unary part and the remainder part of both the zrun data and the weight sequence data without the bit cost in the data stream for identifying the data type.
[0088]
[0088] The extracted zrun data and weight sequence data are then decoded to recover the lower weight values. The decoding of the unary part and the remainder part of the Golomb - Rice code has been discussed in relation to the first embodiment and its description is not repeated here.
[0089]
[0089] The above-described embodiments are to be understood as exemplary examples. Further embodiments are envisioned. For example, in the first embodiment, the single-item data is added to the chunks within the 32-bit portion. However, the size of the portion is not important, and the chunk structure shown in FIG. 5 is adapted according to the specific requirements of the processor that will parse the chunks.
[0090]
[0090] The second embodiment utilizes the Android neural network architecture. However, the techniques described herein may be applied to different software architectures depending on the situation. For example, different software architectures may be used in the context of a server-based implementation. [Item 1] A method for use by a processing element, comprising: obtaining a plurality of variable-length codes, each variable-length code having a variable-length first part and a second part; separating the variable-length code into the first part of the variable-length code and the second part of the variable-length code; for each chunk of a processed stream containing data from the first part of the variable-length code, forming a sub-stream within the chunk of the processed stream such that the data from the first part has a size determined according to a set of control rules; A method comprising the above steps. [Item 2] The method according to Item 1, wherein the first part of the variable-length code is a sign part, and the second part of the variable-length code is a mantissa part. [Item 3] The method according to Item 2, wherein the mantissa part of the variable-length code is a truncated binary part. [Item 4] The method according to Item 2 or 3, wherein the sign part of each variable-length code represents the quotient of the original value represented by the variable-length code. [Item 5] The method according to any one of Items 2 to 4, wherein the mantissa part of each variable-length code is a fixed-length binary code representing the remainder of the value represented by the variable-length code. [Item 6] The method according to any one of Items 1 to 5, wherein the variable-length code is a Golomb-Rice code. [Item 7] The method according to any one of Items 1 to 6, wherein each chunk of the processed stream can contain data from the first part of the variable-length code and data from the second part of the variable-length code, can contain data from the first part of the variable-length code, and not contain data from the second part of the variable-length code, or can contain data from the second part of the variable-length code and not contain data from the first part of the variable-length code. [Item 8] The method according to any one of Items 2 to 5, wherein the control rules are configured to serially determine the size of the sign sub-stream within each chunk. [Item 9] The method according to item 8, wherein the control rule is configured to determine the size of the size of the unary sub-stream within each chunk based on the amount of the unary portion of the variable-length code that remains without being added to the processed stream. [Item 10] The method according to item 9, wherein the processed stream is formed from cells of a predetermined length, and the amount of the unary portion of the variable-length code that remains without being added to the processed stream is the number of bits of the unary portion of the variable-length code associated with the cells that remain without being added to the processed stream. [Item 11] The method according to any one of items 1 to 10, wherein the processed stream is formed from cells of a predetermined length, and each cell of the processed stream has a header and a plurality of chunks, and the header indicates the length of the cell and the length of the unary portion of the variable-length code within the cell. [Item 12] Obtaining a second plurality of variable-length codes; Separating the second plurality of variable-length codes into a first portion of the variable-length code and a second portion of the variable-length code; The step of forming a processed stream within a chunk uses a set of the control rules such that for each chunk of the processed stream including data from at least one first portion of the first plurality of variable-length codes and the second plurality of variable-length codes, the flow control rule determines that the number of first portions of the first plurality of variable-length codes and the number of first portions of the second plurality of variable-length codes are included in the chunk. The method according to item 1. [Item 13] Forming a processed stream includes maintaining a balance value that records the difference between the number of first portions of the first plurality of variable-length codes included in a chunk and the number of first portions of the second plurality of variable-length codes included in the chunk, and the flow control rule is used to determine the number of first portions of the first plurality of variable-length codes and the number of first portions of the second plurality of variable-length codes included in the chunk based on the balance value. The method according to item 12. [Item 14] The method according to any one of items 1 to 13, wherein the plurality of variable-length codes represent weight values for use in a neural network. [Item 15] The method according to any one of items 1 to 14, wherein the plurality of variable-length codes represent values in an activation stream that is an output of a layer of a neural network. [Item 16] A method for decoding a processed data stream, comprising: obtaining a processed data stream including data related to a plurality of variable-length codes, each variable-length code having a variable-length first part and a second part, the processed stream being formed within a chunk, and at least one chunk of the processed stream including a sub-stream of data formed from the first part of the variable-length code; identifying and extracting the sub-stream within the chunk of the processed data stream using a set of flow control rules; A method comprising the above. [Item 17] Each first part of the variable-length codes is encoded as single-bit data having a stop bit, and the method further comprises: obtaining a plurality of identified and extracted sub-streams from a plurality of chunks of the processed data; converting each extracted sub-stream into an intermediate format representing a list of bit positions of stop bits within the sub-stream; combining the plurality of converted sub-streams to form an extended list of bit positions; measuring the distances between the stop bits within the extended list of bit positions to recover the values of the first parts of the variable-length codes in the first stream; The method according to item 16, further comprising the above. [Item 18] A data processing system comprising a processing element and a storage device, the storage device storing a code portion which, when executed by the processing element, causes the data processing system to: obtain a plurality of variable-length codes, each variable-length code having a variable-length first part and a second part; separate the variable-length codes into the first part of the variable-length code and the second part of the variable-length code; for each chunk of the processed stream including data from the first part of the variable-length code, cause the processed stream within the chunk to form a sub-stream within the chunk of the processed stream having a size determined according to a control rule, using the set of control rules; A data processing system.
Claims
1. A method for use by a processing element, comprising: obtaining a plurality of variable-length codes, each variable-length code having a variable-length mantissa part and a fixed-length exponent part; separating the variable-length code into the mantissa part of the variable-length code and the exponent part of the variable-length code; for each chunk of a data stream including data from the mantissa part of the variable-length code, forming a sub-stream within the chunk of the data stream using a selected chunk structure such that the sub-stream has a size determined based on the amount of the mantissa part of the variable-length code remaining without being added to the data stream; A method comprising the above steps.
2. The method according to claim 1, wherein the exponent part of the variable-length code is a truncated binary part.
3. The method according to claim 1 or 2, wherein the mantissa part of each variable-length code represents the quotient of the original value represented by the variable-length code.
4. The method according to any one of claims 1 to 3, wherein the exponent part of each variable-length code is a fixed-length binary code representing the remainder of the value represented by the variable-length code.
5. The method according to any one of claims 1 to 4, wherein the variable-length code is a Golomb-Rice code.
6. Each chunk of the data stream can include data from the mantissa part of the variable-length code and data from the exponent part of the variable-length code, include data from the mantissa part of the variable-length code and not include data from the exponent part of the variable-length code, or include data from the exponent part of the variable-length code and not include data from the mantissa part of the variable-length code. The method according to any one of claims 1 to 5.
7. The data stream is formed from cells of a predetermined length, and the amount of the mantissa part of the variable-length code remaining without being added to the data stream is the number of bits of the mantissa part of the variable-length code associated with the cells remaining without being added to the data stream. The method according to claim 1.
8. The data stream is formed from cells of a predetermined length, and each cell of the data stream has a header and a plurality of chunks. The method according to any one of claims 1 to 7, wherein the header indicates the length of the cell and the length of the single-term portion of the variable-length code within the cell.
9. obtaining a second plurality of variable-length codes; further comprising separating the second plurality of variable-length codes into a single-term portion of the variable-length code and a mantissa portion of the variable-length code; Forming a data stream within a chunk, for each chunk of the data stream including data from at least one single-term portion of the first plurality of variable-length codes and the second plurality of variable-length codes, including the number of single-term portions of the first plurality of variable-length codes and the number of single-term portions of the second plurality of variable-length codes in the chunk The method according to claim 1, comprising:
10. Forming a data stream includes maintaining a balance value that records a difference between the number of single-term portions of the first plurality of variable-length codes included in a chunk and the number of single-term portions of the second plurality of variable-length codes included in the chunk, The method according to claim 9, wherein the balance value is used to determine the number of single-term portions of the first plurality of variable-length codes and the number of single-term portions of the second plurality of variable-length codes included in a chunk.
11. The method according to any one of claims 1 to 10, wherein the plurality of variable-length codes represent weight values for use in a neural network.
12. The method according to any one of claims 1 to 11, wherein the plurality of variable-length codes represent values within an activation stream that is an output of a layer of a neural network.
13. A method for decoding a data stream, comprising: obtaining a data stream including data related to a plurality of variable-length codes, each variable-length code having a variable-length single-term portion and a fixed-length mantissa portion, the data stream being formed within a chunk, and at least one chunk of the data stream including a sub-stream of data formed from the single-term portion of the variable-length code; identifying and extracting the sub-stream within the chunk of the data stream using a chunk structure selected based on the amount of the single-term portion of the variable-length code remaining without being extracted from the data stream; A method comprising:
14. Each single-bit portion of the variable-length code is encoded as single-bit data having a stop bit, and the method comprises: obtaining a plurality of identified and extracted sub-streams from a plurality of chunks of the data stream; converting each extracted sub-stream into an intermediate format representing a list of bit positions of stop bits within the sub-stream; combining the plurality of converted sub-streams to form an extended list of bit positions; measuring the distances between the stop bits in the extended list of bit positions to recover the values of the single-bit portions of the variable-length codes in the data stream; The method according to claim 13, further comprising. **Claim 15** A data processing system comprising a processing element and a storage device, the storage device storing a code portion which, when executed by the processing element, causes the data processing system to: obtain a plurality of variable-length codes, each variable-length code having a variable-length single-bit portion and a fixed-length exponent portion; separate the variable-length code into the single-bit portion of the variable-length code and the exponent portion of the variable-length code; for each chunk of a data stream containing data from the single-bit portion of the variable-length code, form a sub-stream within the chunk of the data stream having a size determined based on the amount of the single-bit portion of the variable-length code remaining without the data from the single-bit portion being added to the data stream, using a selected chunk structure; A data processing system that causes the above to be executed.
Citation Information
Patent Citations
Entropy encoding and decoding scheme
JP2017118547A
Lossless compression of sparse activation maps of neural networks
US20190370667A1
Method for a hybrid golomb-elias gamma coding
WO2010000662A1