Computing core, data processing method, electronic device, and storage medium
Through the conversion module in the calculation core, the floating-point data is compressed and decomposed, and the problem of storage difficulties of non-8-bit integer multiple floating-point numbers is solved, efficient bandwidth and storage space utilization is achieved, and hardware costs are reduced.
Patent Information
- Application Number
- CN202510704189.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-29
AI Technical Summary
When storing floating point data that is not 8-bit integer multiples, the prior art has problems of reading difficulties and resource waste, and cannot effectively reduce storage overhead.
The data is compressed by the conversion module in the calculation core, decompose it into multiple sub-data and merge it into a data format that conforms to an integer multiple relationship, and stored in a register, including the operations of the receiving unit, the decomposition unit and the compression storage unit.
It realizes efficient compressed storage for non-standard data formats, saves bandwidth and storage space, reduces hardware overhead, supports a variety of non-standard data formats, and improves flexibility.
Smart Images

Figure CN120233982B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a computing core, a data processing method, an electronic device, and a non-transitory computer-readable storage medium. Background Art
[0002] Floating-point numbers (FP) are mainly used to represent decimals and usually consist of three parts, namely, a sign bit, an exponent part, and a mantissa part. The exponent part can also be referred to as the exponent value part. For example, a floating-point number V can usually be expressed in the following form:
[0003]
[0004] Among them, the sign bit s can be 1 bit, which determines whether the floating-point number V is negative or positive; M represents the mantissa part. The mantissa part can include multiple bits and is in the form of a binary decimal, which defines the precision of the floating-point number; E represents the exponent (also called the exponent value), which is used to weight the floating-point number and reflects the position of the decimal point in the floating-point number V, defining the value range of the floating-point number.
[0005] Currently, general-purpose hardware usually stores data (such as storing to registers) in units of 8 bits (1 byte). The read / write bit width or storage bit width is usually a multiple of 8, such as 4 bits, 16 bits, etc. For data formats such as FP10 and FP12 whose bit widths are not integer multiples of 8, when storing, if stored as FP10 continuously, since the bit width of the storage component is a multiple of 8, it is difficult to find the start and end positions of each data during reading, increasing the reading overhead. If each FP10 occupies 16-bit width during storage, it will cause waste of resources during storage and cannot achieve the purpose of reducing storage overhead. Summary of the Invention
[0006] At least one embodiment of the present disclosure provides a computing core, including a conversion module, configured to, in response to data compression being performed on data in a first data format obtained by processing the computing core, and when the data bit width m indicated by the data compression is not an integer multiple of 8, read data to be compressed from a register for encoding operations, and store multiple merged compressed data obtained through the encoding operations into the register, where the data bit width n indicated by the first data format is an integer multiple of 8 or a factor of 8, m and n are positive integers and m is less than n; wherein the encoding operations include: converting the data to be compressed into a second data format indicated by the data compression, decomposing each data to be compressed after format conversion to obtain multiple sub-data, and compressing and merging all the decomposed sub-data into multiple merged data in the first data format for storage into the register, where the bit width of each sub-data is an integer multiple of 8 or a factor of 8.
[0007] For example, in the computing core provided by at least one embodiment of the present disclosure, the conversion module includes a receiving unit, a decomposing unit, and a compressing and storing unit; the receiving unit is configured to send a first read address to the register and receive multiple first data stored in T1 registers indicated by the first read address as the data to be compressed; the decomposing unit is configured to convert the multiple first data from the first data format to the second data format, obtain multiple second data respectively corresponding to the multiple first data, and for each second data, decompose the second data into L sub-data, where the sum of the bit widths of the L sub-data is equal to m, and L is a positive integer greater than 1; the compressing and storing unit is configured to compress and merge the multiple sub-data obtained by decomposing the multiple second data into the first data format and store them into T2 registers, where T1 / T2 = n / m, and T1 and T2 are positive integers.
[0008] For example, in the computing core provided by at least one embodiment of the present disclosure, the compressing and storing unit executes compressing and merging the multiple sub-data obtained by decomposing the multiple second data into the first data format and storing them into T2 registers, including performing the following operations: splicing a sub-data into a merged data in the first data format and storing the merged data into a corresponding storage location in the T2 registers, where the a sub-data come from different a second data, and the a sub-data are in the same bit range in their respective second data, a is a positive integer greater than 1 and equal to n / wb, and wb represents the bit width of any one of the a sub-data.
[0009] For example, in the computing core provided by at least one embodiment of the present disclosure, the T1 registers and the T2 registers are both vector registers. Each thread can operate on multiple vector registers simultaneously. Multiple storage units in one vector register are controlled by multiple threads respectively. The a second data come from different vector registers and are stored in the storage units controlled by the same thread.
[0010] For example, in the computing core provided by at least one embodiment of the present disclosure, the multiple sub-data obtained by decomposing each second data include first sub-data. The first sub-data includes bits from the b1-th bit to the b2-th bit in the second data, where b1 and b2 are positive integers. For any one of the T1 registers, b first data are stored in any one of the storage units in the any one register, where b is a positive integer. The b first sub-data obtained based on the b second data corresponding to the b first data are stored in the first storage unit in the T2 registers. The relative position relationship of the b first sub-data in the first storage unit is the same as the relative position relationship of the b first data in the any one storage unit. Among the b first sub-data, the least significant bits of two adjacent first sub-data differ by 16 bits.
[0011] For example, in the computing core provided by at least one embodiment of the present disclosure, the decomposition unit performs converting the multiple first data from the first data format to the second data format to obtain the multiple second data respectively corresponding to the multiple first data, including performing the following operations: for each first data: taking the least significant bit of the first data as the 1st bit, and retaining the n-th bit to the (n - m + 1)-th bit of the first data as the second data corresponding to the first data; or performing a rounding operation on the (n - m + 2)-th bit to the 1st bit of the first data based on a preset rounding rule, adjusting the n-th bit to the (n - m + 1)-th bit of the first data according to the rounding operation result, and retaining the adjusted n-th bit to the (n - m + 1)-th bit of the first data as the second data corresponding to the first data.
[0012] For example, the computing core provided by at least one embodiment of the present disclosure is further configured to: after the receiving unit receives the multiple first data in the T1 registers, clear the data stored in the T1 registers.
[0013] For example, in the computing core provided by at least one embodiment of the present disclosure, T1 × t = n, T2 × t = m, where t is a positive integer and represents the product of all common divisors of n and m.
[0014] For example, in the computing core provided by at least one embodiment of the present disclosure, when L = 2, in response to m = 10, the bit widths of the two sub-data into which the second data is decomposed are 2 bits and 8 bits respectively; in response to m = 12, the bit widths of the two sub-data into which the second data is decomposed are 4 bits and 8 bits respectively.
[0015] For example, the computing core provided by at least one embodiment of the present disclosure further includes a denormalization processing unit, which is configured to: in response to the number of bits of the exponent part indicated by the first data format being different from the number of bits of the exponent part indicated by the second data format, perform a denormalization processing operation on the plurality of first data, and send the plurality of first data after the denormalization processing operation to the decomposition unit for subsequent operations.
[0016] For example, in the computing core provided by at least one embodiment of the present disclosure, the conversion module is further configured to, in response to data decompression, perform a decoding operation on the data directly read from the register to convert the directly read data into the first data format and store it in the register for the computing core to read and use; wherein, the decoding operation includes: splitting and reorganizing the directly read data according to the positional relationship during storage, and performing bit expansion to obtain the data in the first data format.
[0017] For example, in the computing core provided by at least one embodiment of the present disclosure, the conversion module includes a receiving unit and a decoding unit. The receiving unit is configured to send a second read address to the register and receive a plurality of merged data stored in T2 registers indicated by the second read address; the decoding unit is configured to split and reorganize the plurality of merged data to obtain a plurality of second data, expand the bit width of each second data to n bits to convert it into the first data format, where the data bit width of each second data is m; the receiving unit is further configured to send the data after format conversion to T1 registers, where T2 and T1 are positive integers, and T1 / T2 = n / m.
[0018] For example, in the computing core provided by at least one embodiment of the present disclosure, the decoding unit is implemented by an OR module, an AND module, and a shift module; for the plurality of merged data used to obtain the target data, the shift module is used to shift the received plurality of merged data to expand the bit width and adjust the positions of the bits according to the positional relationship; the AND module is used to extract the corresponding bits from the plurality of merged data according to the positional relationship; the OR module is used to splice the extracted corresponding bits to obtain the target data, where the data format of the target data is the first data format.
[0019] For example, in the computing core provided by at least one embodiment of the present disclosure, the OR module, the AND module, and the displacement module are reused when obtaining different target data.
[0020] At least one embodiment of the present disclosure provides a data processing method, including: in response to data compression of the received data in the first data format, and when the data bit width m of the compressed data indicated by the data compression is not an integer multiple of 8, reading the data to be compressed in the first data format from a register for encoding operations, and storing the multiple merged compressed data obtained through the encoding operations into the register, where the data bit width n indicated by the first data format is an integer multiple of 8 or a factor of 8, m and n are positive integers and m is less than n; where the encoding operations include: converting the data to be compressed into the second data format indicated by the data compression, decomposing each data to be compressed after format conversion to obtain multiple sub-data, and compressing and merging all the decomposed sub-data into multiple merged data in the first data format to be stored in the register, where the bit width of each sub-data is an integer multiple of 8 or a factor of 8.
[0021] At least one embodiment of the present disclosure provides a data processing method, including: in response to data decompression, performing a decoding operation on the data directly read from the register to convert the directly read data into the first data format and storing it in the register; where the decoding operations include: splitting and reorganizing the directly read data according to the positional relationship during storage, and performing bit expansion to obtain the data in the first data format.
[0022] At least one embodiment of the present disclosure provides an electronic device, including: a memory that stores computer-executable instructions non-transiently; a processor configured to run the computer-executable instructions, where the computer-executable instructions, when run by the processor, implement the data processing method according to at least one embodiment of the present disclosure.
[0023] At least one embodiment of the present disclosure provides a non-transient computer-readable storage medium, where the non-transient computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the data processing method according to at least one embodiment of the present disclosure.
[0024] The computing core provided by at least one embodiment of the present disclosure can perform encoding operations on non-standard data formats (for example, the bit width is not an integer multiple of 8), and compress and store them in multiple registers according to the data formats supported by the hardware. In general-purpose hardware, non-standard data formats can be used for compressed storage to save bandwidth, or they can be compressed and stored in local storage components such as hard disks to save storage space, achieving a balance between computing efficiency and storage space. Moreover, the computing core provided by at least one embodiment of the present disclosure does not require a dedicated storage component for non-standard data formats, reducing hardware overhead and R & D costs, and supporting different non-standard data formats, with stronger flexibility. Brief Description of the Drawings
[0025] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0026] Figure 1 It is a schematic structural diagram of a general-purpose graphics processing unit (GPGPU);
[0027] Figure 2 It is a schematic structural diagram of a vector register;
[0028] Figure 3 It is a schematic storage structure diagram of a vector register;
[0029] Figure 4 It is a schematic structural diagram of the computing core provided by at least one embodiment of the present disclosure;
[0030] Figure 5 It is a schematic structural diagram of the conversion module 101 provided by at least one embodiment of the present disclosure;
[0031] Figure 6A It is a schematic storage structure diagram of T1 registers provided by one embodiment of the present disclosure;
[0032] Figure 6B It is a schematic structural diagram of T2 vector registers provided by one embodiment of the present disclosure;
[0033] Figure 7A It is a schematic storage structure diagram of T1 registers provided by another embodiment of the present disclosure;
[0034] Figure 7B It is a schematic structural diagram of T2 vector registers provided by another embodiment of the present disclosure;
[0035] Figure 8 It is a schematic structural diagram of the computing core provided by at least one embodiment of the present disclosure;
[0036] Figure 9ASchematic diagram of the decoding process provided by an embodiment of the present disclosure;
[0037] Figure 9B Schematic diagram of the decoding process provided by another embodiment of the present disclosure;
[0038] Figure 10 Schematic flowchart of the data processing method provided by at least one embodiment of the present disclosure;
[0039] Figure 11 Schematic block diagram of an electronic device provided by an embodiment of the present disclosure;
[0040] Figure 12 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners
[0041] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0042] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to indicate relative positional relationships, and when the absolute positions of the objects being described change, the relative positional relationships may also change accordingly. To keep the following description of the embodiments of the present disclosure clear and concise, some details of known functions and known components are omitted in the present disclosure.
[0043] Traditional floating-point numbers include three formats, namely half-precision floating-point numbers (FP16), single-precision floating-point numbers (FP32), and double-precision floating-point numbers (FP64), and their exponent parts and mantissa parts have different numbers of bits.
[0044] AI (Artificial Intelligence) accelerators and the like have been widely used in the training of deep learning models. For the common convolution operations in deep learning models, special optimizations have been made in both software and hardware designs to accelerate calculations. For example, in the fields of artificial intelligence or deep learning, various floating-point data formats have been optimized and developed, such as FP8 (8-bit floating point), BF16 (brain floating point 16, with a bit width of 16 bits), TF32 (Tensor Float 32, with a bit width of 19 bits), etc. These data formats can significantly reduce the computing resources and power consumption required for computing, especially for matrix multiplication or convolution multiplication operations.
[0045] Table 1 below shows the data formats of several floating-point precision types.
[0046] Table 1 Data Formats
[0047]
[0048] As shown in Table 1, for the floating-point precision type FP32, the total number of bits is 32 bits, including 1 sign bit, the exponent part (i.e., the exponent) includes 8 bits, and the mantissa part includes 23 bits. For FP16, the total number of bits is 16 bits, including 1 sign bit, the exponent part includes 5 bits, and the mantissa part includes 10 bits. For BF16, the total number of bits is 16 bits, including 1 sign bit, the exponent part includes 8 bits, and the mantissa part includes 7 bits. For FP8, there are two data types, which differ in the bit width of the exponent part and the mantissa part. For FP10, the total number of bits is 10 bits, including 1 sign bit, the exponent part includes 5 bits, and the mantissa part includes 4 bits. For FP12, the total number of bits is 12 bits, including 1 sign bit, the exponent part includes 5 bits, and the mantissa part includes 6 bits.
[0049] Since FP8 has a total bit width of only 8 bits, the numerical range it can represent is limited, and the precision is insufficient in some applications. The computational and storage overheads of FP16 and BF16 are both large. Therefore, data formats such as FP10 and FP12 can be used. Their represented precision range is more than that of FP8, but the computational and storage overheads they occupy are smaller than those of FP16 or BF16, and they can achieve a balance between precision and storage overhead.
[0050] Figure 1 It is a schematic structural diagram of a general-purpose graphics processing unit (GPGPU).
[0051] As Figure 1As shown, a general-purpose graphics processing unit is actually an array of programmable multi-processors. For example, the programmable multi-processor can be a Streaming Processor Cluster (SPC), such as including Figure 1 the streaming processor clusters 1, ..., M as shown, where M is a positive integer greater than 1. In the general-purpose graphics processing unit, one streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. Data sharing between multiple streaming processor clusters is achieved through a global cache or global memory.
[0052] As Figure 1 shown, taking the streaming processor cluster 1 as an example, one streaming processor cluster includes multiple computing units, such as Figure 1 the computing units 1, 2, ..., N in it, where N is a positive integer. Each computing unit (Compute Unit, CU for short) is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A computing unit includes multiple cores (also called computing cores), and each computing core includes an Arithmetic Logic Unit (ALU), a floating-point computing unit, etc. The computing core is used to perform specific computing tasks. In addition, the computing unit also includes registers (such as Figure 1 the register file in it) and shared memory, which are used to hierarchically store the source data and destination data related to the computing task. The shared memory in one computing unit is used to share data between the cores of this computing unit.
[0053] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before being executed in the general-purpose graphics processing unit (or called parallel computing processor), and then multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in the figure). All threads in one thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundle (or simply called thread bundle, warp), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.
[0054] In each computing unit, the thread bundle scheduling / distribution module ( Figure 1The warp is scheduled and allocated (not shown in the figure) so that multiple computing cores in the computing unit can run the warp. According to the number of computing cores in the computing unit, multiple warps in a thread block can be executed simultaneously or time-shared. Multiple threads in each warp will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the intermediate-level cache or global cache or global memory for read and write operations, etc.
[0055] A vector register can be provided in the processor. Figure 2 It is a schematic structural diagram of a vector register.
[0056] As Figure 2 shown, the processor provides M memory banks (banks) corresponding one-to-one with M threads. For example, M = 32, and the M threads belong to the same warp and are used to execute the same task in parallel. For example, single instruction multiple threads (SIMT) in a GPGPU or GPU. Each memory bank includes N double words, and N is a positive integer. The bit width of each vector register is M double words (Dword, 1 Dword = 32 bits = 4 bytes). For example, the M double words included in 1 vector register come from M memory banks respectively. Thus, each thread can operate N double words in N vector registers, and the M double words in one vector register can be operated in parallel by M threads.
[0057] Figure 3 It is a schematic storage structure diagram of a vector register.
[0058] As Figure 3 shown, vector register 0 includes 32 double words, and each thread operates one double word. For example, Figure 3 T0, T1,..., T31 in it respectively represent 32 threads, and each box represents a double word that the corresponding thread can operate.
[0059] For example, taking the stored data as 16-bit data as an example, two data can be stored in each double word. For example, Figure 3 the numbers 0, 1, 2, 3, etc. in it all represent a 16-bit data.
[0060] General hardware usually operates in units of 8 bits (1 byte) when storing (such as storing to registers), and the read / write bit width or storage bit width is usually a multiple of 8, such as 4 bits, 16 bits, etc. For data formats with bit widths that are not integer multiples of 8, such as FP10 and FP12, when storing, if stored as FP10 and continuously stored, since the bit width of the storage component is a multiple of 8, it is difficult to find the start and end positions of each data during reading, increasing the reading overhead. If each FP10 occupies a 16-bit width during storage, it will cause waste of resources during storage and cannot achieve the purpose of reducing storage overhead.
[0061] At least one embodiment of the present disclosure provides a computing core, a data processing method, an electronic device, and a non-transitory computer-readable storage medium. The computing core includes a conversion module, and the conversion module is configured to, in response to data compression of data in a first data format obtained by processing the computing core, and the compressed data bit width m indicated by the data compression not being an integer multiple of 8, read the data to be compressed from the register for encoding operations, and store the multiple compressed and combined data obtained through the encoding operations into the register, where the data bit width n indicated by the first data format is an integer multiple of 8 or a factor of 8, m and n are positive integers and m is less than n; where the encoding operations include: converting the data to be compressed into a second data format indicated by the data compression, decomposing each data to be compressed after format conversion to obtain multiple sub-data, and compressing and combining all the decomposed sub-data into multiple combined data in the first data format for storage in the register, where the bit width of each sub-data is an integer multiple of 8 or a factor of 8.
[0062] The computing core can perform encoding operations on non-standard data formats (such as bit widths that are not integer multiples of 8), and compress and store them in multiple registers according to the data formats supported by the hardware. In general hardware, non-standard data formats can be used for compressed storage to save bandwidth, or can be compressed and stored in local storage components such as hard disks to save storage space, achieving a balance between computing efficiency and storage space. Moreover, the computing core provided by at least one embodiment of the present disclosure does not require a dedicated storage component for non-standard data formats, reducing hardware overhead and R & D costs, and supporting different non-standard data formats with stronger flexibility.
[0063] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.
[0064] Figure 4 It is a schematic structural diagram of the computing core provided by at least one embodiment of the present disclosure.
[0065] As Figure 4 shown, the computing core 100 includes a conversion module 101.
[0066] For example, the computing core 100 may be a vector computing core, a tensor computing core, etc., and the present disclosure does not specifically limit the structure and function of the computing core.
[0067] For example, the conversion module 101 is configured to, in response to data compression of data in the first data format processed by the computing core 100, and when the data bit width m of the compressed data indicated by the data compression is not an integer multiple of 8, read the data to be compressed from the register for encoding operation, and store the multiple merged compressed data obtained through the encoding operation into the register. For example, the data bit width n indicated by the first data format is an integer multiple of 8 or a factor of 8, m and n are positive integers and m is less than n.
[0068] For example, the encoding operation includes: converting the data to be compressed into the second data format indicated by the data compression, decomposing each data to be compressed after format conversion to obtain multiple sub-data, and compressing and merging the decomposed sub-data into multiple merged data in the first data format for storage into the register, where the bit width of each sub-data is an integer multiple of 8 or a factor of 8.
[0069] For example, the computing core 100 receives an encoding request, and the encoding request requires the conversion module to perform data compression. At the same time, the storage address of the data to be compressed in the register, that is, the first read address, and the write address of the compressed data in the register, that is, the first write address, are sent to the computing core.
[0070] For example, the data in the first data format may be processed by the computing core. For example, when the calculated data needs to be transferred to other computing cores, a streaming processor cluster, a local disk, etc., data compression can be performed through the conversion module and stored in the register. Then, the compressed data can be transmitted to other computing cores, a streaming processor cluster, etc. to save the transmission bandwidth during transmission, or the compressed data can be stored in the local disk, hard disk, etc., internal memory or external memory to save storage space. The register is usually the most basic component in the processor. The compressed data can be first placed in the register and then other operations can be performed. Of course, the compressed data can also be stored in the cache. The present disclosure does not specifically limit this.
[0071] For example, the encoding request requires compressing the FP16 format data into the FP10 format data. Of course, the present disclosure is not limited to this. The first data format may be FP4, FP8, BF16, FP32, etc., as long as the data bit width n thereof is an integer multiple of 8 or a factor of 8; the second data format indicated by the data compression may be FP10, FP12, etc., as long as the data bit width m indicated by the second data format is not an integer multiple of 8.
[0072] For example, the conversion module 101 reads the data to be compressed in the first data format from the register through the first read address. The conversion module 101 converts the data to be compressed into the second data format required by the encoding request, and decomposes each data to be compressed after format conversion to obtain a plurality of sub-data. For example, the bit width of each sub-data is an integer multiple of 8 or a factor of 8. All the sub-data obtained by decomposition are compressed and merged into a plurality of merged data in the first data format as the compressed data, and stored in the position indicated by the first write address in the register.
[0073] For example, the register 102 may be a vector register. For example, the structure of the register 102 may refer to the relevant content described above. Figure 2 For example, each thread can operate multiple vector registers simultaneously. Multiple storage units in a vector register are controlled by multiple threads respectively. For example, one storage unit may be 1 double word in the vector register. For example, Figure 2 it is the double word with label 0 in storage block 0 in it.
[0074] Figure 5 It is a schematic structural diagram of the conversion module 101 provided by at least one embodiment of the present disclosure.
[0075] As Figure 5 shown, the conversion module 101 includes a receiving unit 1011, a decomposing unit 1012, and a compression storage unit 1013.
[0076] For example, the receiving unit 1011 is configured to send the first read address to the register 102, and receive a plurality of first data stored in T1 registers indicated by the first read address as the data to be compressed.
[0077] For example, the first read address includes the address information of T1 registers. The receiving unit receives a plurality of first data read from T1 registers according to the T1 address information. The plurality of first data are in the first data format, and the data bit width of the first data format is n bits, for example, 16 bits.
[0078] The decomposing unit 1012 is configured to convert the plurality of first data from the first data format to the second data format, obtain a plurality of second data respectively corresponding to the plurality of first data, and for each second data, decompose the second data into L sub-data. For example, the sum of the bit widths of the L sub-data is equal to m, and L is a positive integer greater than 1.
[0079] For example, the decomposition unit 1012 performs the conversion of multiple first data from the first data format to the second data format, obtaining multiple second data respectively corresponding to the multiple first data, including performing the following operations: for each first data: taking the least significant bit of the first data as the 1st bit, and retaining the nth bit to the (n - m + 1)th bit of the first data as the second data corresponding to the first data; or performing a rounding operation on the (n - m + 2)th bit to the 1st bit of the first data based on a preset rounding rule, adjusting the nth bit to the (n - m + 1)th bit of the first data according to the result of the rounding operation, and retaining the adjusted nth bit to the (n - m + 1)th bit of the first data as the second data corresponding to the first data.
[0080] For example, to simplify the implementation process, the lower n - m bits of the first data can also be directly truncated, and the first m bits of the first data are taken as the second data.
[0081] For example, a rounding operation can be performed on the lower n - m bits of the first data according to a preset rounding rule, and the first m bits of the first data are adjusted according to the result of the rounding operation, and the adjusted first m bits of the first data are retained as the second data. The present disclosure does not make specific limitations on the preset rounding rule, and any feasible rounding rule can be adopted, such as rounding, etc., which will not be elaborated here.
[0082] For example, the number of L can be set as needed. For each second data, it is ensured that the sum of the bit widths of the decomposed sub - data is equal to m, and the bit width of each sub - data is an integer multiple of 8 or a factor of 8.
[0083] For example, in some embodiments, L = 2, m = m1 + m2, the second data can be divided into two sub - data, and the bit widths of the two sub - data are m1 and m2 respectively. m satisfies the following condition: m % 8 ≠ 0, m1 satisfies the following condition: (m1 % 8 == 0) or (8 % m1 == 0), and m2 satisfies the following condition: (m2 % 8 == 0) or (8 % m2 == 0), where "%" represents the remainder operation.
[0084] For example, in one embodiment, the first data format is FP16, the second data format is FP10, L = 2, and the 2 sub - data into which each second data is decomposed can have bit widths of 2 bits and 8 bits respectively. For example, taking the least significant bit of the second data as the 1st bit, the 1st bit and the 2nd bit of the second data can be taken as one sub - data, and the remaining 8 bits as another sub - data, or the 1st bit to the 8th bit of the second data can be taken as one sub - data, and the remaining 2 bits as another sub - data. The present disclosure does not make specific limitations on this.
[0085] For example, in another embodiment, the first data format is FP16, the second data format is FP12, L = 2, and each second data is decomposed into 2 sub-data with bit widths of 4 bits and 8 bits respectively. For example, taking the least significant bit of the second data as the first bit, the first bit to the fourth bit can be used as one sub-data, and the remaining 8 bits as another sub-data. Alternatively, the first bit to the eighth bit of the second data can be used as one sub-data, and the remaining 4 bits as another sub-data. The present disclosure does not make specific limitations on this.
[0086] The partitioning method can be set as needed, and the present disclosure does not make specific limitations on this.
[0087] For example, in some embodiments, when the bit width of the exponent part of the second data format is the same as that of the exponent part of the first data format, denormalization may not be considered during format conversion. However, if the bit width of the exponent part of the second data format is different from that of the exponent part of the first data format, the denormalization processing unit in the computing core needs to perform relevant processing, and then the decomposition unit performs subsequent operations such as decomposing sub-data.
[0088] The denormalization processing unit is configured to, in response to the number of bits of the exponent part indicated by the first data format being different from the number of bits of the exponent part indicated by the second data format, perform denormalization processing operations on multiple first data, and send the multiple first data after the denormalization processing operations to the decomposition unit for subsequent operations.
[0089] The denormalization processing operations may include, for example, identifying denormalized numbers, such as checking whether the exponent part is all 0 to identify denormalized numbers; calculating the actual value of the denormalized number using the formula for denormalized numbers; and performing rounding processing on the denormalized number. The rounding processing may include truncation, rounding towards 0, rounding towards the nearest number, rounding towards positive or negative infinity, and rounding towards an even number, etc.
[0090] The compression storage unit 1013 is configured to compress and merge the multiple sub-data obtained by decomposing multiple second data in the first data format and store them in T2 registers.
[0091] For example, T1 / T2 = n / m, where T1 and T2 are positive integers. For example, T1×t = n, T2×t = m, and t is a positive integer representing the product of all common divisors of n and m.
[0092] For example, in one embodiment, n = 16, m = 10, then T1 = 8, T2 = 5. For example, in one embodiment, n = 16, m = 12, then T1 = 4, T2 = 3.
[0093] For example, in a processor, a large input matrix / tensor can usually be divided into multiple small matrices, and each small matrix can be called a "TILE". For example, a matrix including 8×8 FP16-format data can be called a TILE. The size of 8×8 FP16 data is 32 Dwords and can be stored in a vector register.
[0094] For example, the matrix including 16×32 FP16 stored in 8 vector registers can be compressed into FP10-format data and stored in 5 vector registers, thereby reducing the storage space occupancy and reducing the transmission bandwidth.
[0095] For example, the matrix including 16×16 FP16 stored in 4 vector registers can be compressed into FP12-format data and stored in 3 vector registers, thereby reducing the storage space occupancy and reducing the transmission bandwidth.
[0096] For example, after the receiving unit receives multiple first data in T1 registers, the data stored in the T1 registers can also be cleared to reduce the storage space occupancy and release the T1 registers for other operations.
[0097] For example, the compression storage unit executes multiple sub-data obtained by decomposing multiple second data, compresses and merges them in the first data format, and stores them in T2 registers, including performing the following operations: splicing a sub-data into a merged data in the first data format and storing the merged data in the corresponding storage location in the T2 registers, where the a sub-data come from different a second data, and the a sub-data are in the same bit range in their respective second data, a is a positive integer greater than 1 and equal to n / wb, and wb represents the bit width of any one of the a sub-data.
[0098] For example, each second data can be decomposed into a first sub-data, a second sub-data, etc. For example, the first sub-data includes the bits from the b1-th bit to the b2-th bit in the second data, b1 and b2 are positive integers, and the difference between b1 and b2 is wb, that is, the sub-bit width of the first sub-data. For example, the first sub-data in a second data can be spliced into a merged data, the bit width of the merged data is n, and the merged data is stored in the corresponding storage location in the T2 registers. The selection of the combination and the setting of the splicing order can be determined according to actual needs, and the present disclosure does not make specific limitations thereon.
[0099] For example, in some embodiments, to reduce the read / write overhead caused by data exchange between threads, a second data are from different vector registers and stored in storage units controlled by the same thread. For example, sub-data obtained by decomposing the second data selected from the first data in multiple storage units controlled by thread 0 can be concatenated into merged data, thereby enabling data to flow within the thread, avoiding cross-thread interaction, reducing read / write overhead, and improving compression efficiency.
[0100] For example, when storing the merged data into a register, it can be stored in the storage unit controlled by the same thread, thereby reducing cross-thread data interaction and read / write overhead.
[0101] For example, in one embodiment, multiple sub-data obtained by decomposing each second data include first sub-data, and the first sub-data includes bits from the b1-th bit to the b2-th bit in the second data, where b1 and b2 are positive integers.
[0102] For any one of the T1 registers, b first data are stored in any one of the storage units in the register, where b is a positive integer. b first sub-data obtained based on the b second data corresponding to the b first data are stored in the first storage unit of the T2 registers. The relative position relationship of the b first sub-data in the first storage unit is the same as the relative position relationship of the b first data in the any one storage unit. Among the b first sub-data, the least significant bits of two adjacent first sub-data differ by 16 bits.
[0103] For the merged data satisfying the above arrangement relationship, its compression efficiency is the highest. When reading, data extraction can be completed using simple Boolean operations and shift operations, and the logic for extracting different second data is consistent, enabling the reuse of corresponding modules or logic. One thread can extract b data in one clock cycle, achieving the highest throughput and improving compression efficiency and decompression efficiency.
[0104] Next, in conjunction with the accompanying drawings, the working process of the compression and merge unit will be specifically described.
[0105] Figure 6A It is a schematic diagram of the storage structure of T1 registers provided in an embodiment of the present disclosure.
[0106] For example, assume that the encoding request requires compressing data in FP16 format and storing it in the register in FP10 format, where T1 = 8 and T2 = 5.
[0107] Figure 6AA schematic diagram of storing data in 8 registers is shown. For example, each small square with a number represents an FP16 data. The thick outer frames around two adjacent boxes (such as 0 and 1, 2 and 3, 4 and 5, etc.) indicate that two data are stored in the same storage unit, for example, stored in Figure 2 a double word of the storage block shown.
[0108] Figure 6A R0, R1, R2,..., R7 in it represent 8 vector registers, and each vector register can store 8×8 FP16 data.
[0109] For example, when compressing the FP16 data in 8 vector registers into the FP10 format, the total size of these data is 160 double words, occupying 5 vector registers.
[0110] Figure 6B This is a schematic structural diagram of T2 vector registers provided by an embodiment of the present disclosure.
[0111] As Figure 6B shown, (a) shows multiple first data stored in T1 vector registers. As described above, for example, to reduce the read and write overhead caused by data exchange between threads, the first data in the storage units controlled by the same thread can be selected for compression when compressing.
[0112] Figure 6B T0R0 in it represents a storage unit in the R0 register controlled by thread 0, and 0 and 1 represent two first data with the FP16 data format stored in this storage unit. T0R1 represents a storage unit in the R1 register controlled by thread 0, and 8 and 9 represent two first data with the FP16 data format stored in this storage unit. T0R4 represents a storage unit in the R4 register controlled by thread 0, and 16 and 17 represent two first data with the FP16 data format stored in this storage unit. The definitions of T0R5, T0R2, T0R3, T0R6, and T0R7 are the same and will not be elaborated here.
[0113] Assume that when the decomposition unit decomposes, the least significant bit of the first data is regarded as the 1st bit, and the 16th bit and the 15th bit of each first data are selected as a sub-data (sub-data 1), the 14th bit to the 7th bit are used as another sub-data (sub-data 2), and the 6th bit to the 1st bit are truncated, thus obtaining multiple sub-data.
[0114] As Figure 6BAs shown in (c), the compression storage unit concatenates the sub-data 1 obtained based on the first data 0 ("0" in (c)), the sub-data 1 obtained based on the first data 8 ("8" in (c)), the sub-data 1 obtained based on the first data 16 ("16" in (c)),..., the sub-data 1 obtained based on the first data 281 ("281" in (c)) into a combined data, and stores it in a storage unit controlled by thread 0 in the R12 register.
[0115] For example, if the first data 0 and the first data 1 are stored in the same storage unit, then the sub-data 1 of the first data 0 and the sub-data 1 of the first data 1 are stored in the same storage unit T0R12. In this storage unit, the sub-data 1 of the first data 0 is at the lower position, and the sub-data 1 of the first data 1 is at the higher position. The relative position relationship is the same as the relative position relationship between the first data 0 and the first data 1 in the storage unit T0R0. Also, the lowest bit of the sub-data 1 of the first data 0 and the lowest bit of the sub-data 1 of the first data 1 differ by 16 bits. Thus, it can be ensured that decompression can be performed with the highest throughput during reading. For specific details, refer to the description of the decoding process later.
[0116] As Figure 6B As shown in (b), the compression storage unit concatenates the sub-data 2 obtained based on the first data 0 ("0" in (b)) and the sub-data 2 obtained based on the first data 8 ("8" in (b)) into a combined data, and concatenates the sub-data 2 obtained based on the first data 1 ("1" in (b)) and the sub-data 2 obtained based on the first data 9 ("9" in (b)) into a combined data. The two combined data are jointly stored in a storage unit controlled by thread 0 in the R8 register.
[0117] For example, if the first data 0 and the first data 1 are stored in the same storage unit, then the sub-data 2 of the first data 0 and the sub-data 2 of the first data 1 are stored in the same storage unit. In this storage unit, the sub-data 2 of the first data 0 is at the lower position, and the sub-data 2 of the first data 1 is at the higher position. The relative position relationship is the same as the relative position relationship between the first data 0 and the first data 1 in the storage unit T0R0. Also, the lowest bit of the sub-data 2 of the first data 0 and the lowest bit of the sub-data 2 of the first data 1 differ by 16 bits. Thus, it can be ensured that decompression can be performed with the highest throughput during reading. For specific details, refer to the description of the decoding process later.
[0118] The compression methods and arrangement logics for other first data are the same as those for the first data 0 and the first data 1, and will not be elaborated here.
[0119] Figure 7A It is a schematic diagram of the storage structure of T1 registers provided by another embodiment of the present disclosure.
[0120] For example, assume that the encoding request requires compressing data in FP16 format and storing it in FP12 format in a register, with T1 = 4 and T2 = 3.
[0121] Figure 7A A schematic diagram of storing data in 4 registers is shown. For example, each small square with a number represents an FP16 data, and two adjacent squares (such as 0 and 1, 2 and 3, 4 and 5, etc.) are stored in the same storage unit, such as stored in Figure 2 a double word of the storage block shown.
[0122] Figure 7A R0, R1, R2, and R3 in represent 4 vector registers, and each vector register can store 8×8 FP16 data.
[0123] For example, compressing the FP16 data in 4 vector registers into FP12 format, the total size of these data is 96 double words, occupying 3 vector registers.
[0124] Figure 7B A schematic structural diagram of T2 vector registers provided by another embodiment of the present disclosure.
[0125] Assume that when the decomposition unit decomposes, the least significant bit of the first data is used as the 1st bit, and the 16th to 13th bits of each first data are selected as a sub-data (sub-data 1), the 12th to 5th bits are used as another sub-data (sub-data 2), and the 4th to 1st bits are truncated, thus obtaining multiple sub-data. For example, sub-data 1 includes Figure 7B the 8 - 10 and "S" (sign bit) marked in, and sub-data 2 includes Figure 7B the 0 - 7 marked in.
[0126] As Figure 7B shown, the compression storage unit splices the sub-data 1 obtained based on the first data 0, the sub-data 1 obtained based on the first data 8, the sub-data 1 obtained based on the first data 128,..., the sub-data 1 obtained based on the first data 137 into a combined data and stores it in a storage unit controlled by thread 0 in the R4 register.
[0127] For example, if the first data 0 and the first data 1 are stored in the same storage unit, then the sub-data 1 of the first data 0 and the sub-data 1 of the first data 1 are stored in the same storage unit T0R4. In this storage unit, the sub-data 1 of the first data 0 is located at the lower position, and the sub-data 1 of the first data 1 is located at the higher position. The relative position relationship is the same as the relative position relationship between the first data 0 and the first data 1 in the storage unit T0R0. Moreover, the lowest bit of the sub-data 1 of the first data 0 and the lowest bit of the sub-data 1 of the first data 1 differ by 16 bits. Thus, it can be ensured that decompression can be performed with the highest throughput during reading. For specific details, reference can be made to the description of the decoding process in the following text.
[0128] As Figure 7B shown, the compression storage unit concatenates the sub-data 2 obtained based on the first data 0 and the sub-data 2 obtained based on the first data 8 into a combined data, and concatenates the sub-data 2 obtained based on the first data 1 and the sub-data 2 obtained based on the first data 9 into a combined data. The two combined data are stored in a storage unit controlled by thread 0 in register R5.
[0129] As Figure 7B shown, the compression storage unit concatenates the sub-data 2 obtained based on the first data 128 and the sub-data 2 obtained based on the first data 136 into a combined data, and concatenates the sub-data 2 obtained based on the first data 129 and the sub-data 2 obtained based on the first data 137 into a combined data. The two combined data are stored in a storage unit controlled by thread 0 in register R6.
[0130] For example, if the first data 0 and the first data 1 are stored in the same storage unit, then the sub-data 2 of the first data 0 and the sub-data 2 of the first data 1 are stored in the same storage unit. In this storage unit, the sub-data 2 of the first data 0 is located at the lower position, and the sub-data 2 of the first data 1 is located at the higher position. The relative position relationship is the same as the relative position relationship between the first data 0 and the first data 1 in the storage unit T0R0. Moreover, the lowest bit of the sub-data 2 of the first data 0 and the lowest bit of the sub-data 2 of the first data 1 differ by 16 bits. Thus, it can be ensured that decompression can be performed with the highest throughput during reading. For specific details, reference can be made to the description of the decoding process in the following text.
[0131] The compression methods and arrangement logics for other first data are the same as those for the first data 0 and the first data 1, and will not be elaborated here.
[0132] Of course, as described above, the above embodiments provide an arrangement method with the highest throughput and the highest efficiency. Of course, the present disclosure is not limited thereto, and other arrangement methods can also be used for compression and combination, which will not be elaborated here.
[0133] In some embodiments, if the data received from other computing cores, streaming processor clusters, processors, etc. has been compressed and merged through the above encoding process, the conversion module can restore the data and convert it back to the first data format.
[0134] For example, the conversion module is further configured to, in response to data decompression, perform a decoding operation on the data directly read from the register to convert the directly read data into the first data format and store it in the register for the computing core to read and use; wherein, the decoding operation includes: splitting and reorganizing the directly read data according to the positional relationship during storage, and performing bit expansion to obtain the data in the first data format.
[0135] Figure 8 Schematic structural diagram of a computing core provided by at least one embodiment of the present disclosure.
[0136] As Figure 8 shown, the computing core 100 receives a decoding request, which requests the conversion module to perform data decompression. At the same time, the storage address of the data to be decoded in the register, that is, the second read address, and the write address of the decoded data in the register, that is, the second write address, are sent to the computing core.
[0137] For example, the data to be decoded may be compressed and stored by the computing core, or the data to be decoded may be obtained from other components such as other computing cores, streaming processor clusters, processors, memory, and external storage, and it has also undergone the above-mentioned encoding process in other components. After decoding the data to be decoded, it can be restored to the data in the first data format for the computing core or other components to use and read.
[0138] For example, the conversion module 101 reads the merged data from the register through the second read address. The conversion module 101 splits and reorganizes the merged data, and performs bit expansion to obtain the data in the first data format as the decoded data, and stores it at the position indicated by the second write address in the register.
[0139] For example, the conversion module includes a receiving unit and a decoding unit. For example, the receiving unit may share the receiving unit 1011 described above.
[0140] For example, the receiving unit is configured to send the second read address to the register and receive a plurality of merged data stored in T2 registers indicated by the second read address.
[0141] The decoding unit is configured to split and reorganize multiple merged data to obtain multiple second data, and expand the bit width of each second data to n bits to convert it into the first data format, where the data bit width of each second data is m; the receiving unit is further configured to send the data after format conversion to T1 registers, where T2 and T1 are positive integers, and T1 / T2 = n / m.
[0142] The T1 registers here may be the same as or different from the T1 registers that stored the data to be compressed before.
[0143] The process of splitting and reorganizing is the reverse process of the aforementioned decomposition and merging process, and can be carried out according to the positional relationship during storage.
[0144] For example, the decoding unit is implemented by an OR module, an AND module, and a shift module; for multiple merged data used to obtain the target data, the shift module is used to shift the received multiple merged data to expand the bit positions and adjust the positions of the bit positions according to the positional relationship; the AND module is used to extract the corresponding bit positions from the multiple merged data according to the positional relationship during storage; the OR module is used to splice the extracted corresponding bit positions to obtain the target data. The target data here has been subjected to format conversion, and its data format is the first data format.
[0145] In some embodiments, a special sub - data arrangement method during encoding can obtain the optimal throughput and the highest decoding efficiency. In this embodiment, the OR module, the AND module, and the shift module are reused when obtaining different target data, that is, different second data can be implemented using the same Boolean operation and shift operation logic during decoding, improving the decoding efficiency.
[0146] Next, regarding Figure 6B the following decoding process of the shown compression arrangement method will be described.
[0147] Figure 9A This is a schematic diagram of the decoding process provided by an embodiment of the present disclosure.
[0148] In Figure 9A , referring to Figure 6B the relevant content, the register R12 stores the sub - data 1 part of the first data 0, including the sign bit s and the highest bit of the data part, and the register R8 stores the sub - data 2 part of the first data 0, including the remaining 8 bit positions of the data part. In Figure 9A , the 0 in the bit position part represents the least significant bit, and S represents the sign bit, located at the most significant bit.
[0149] As Figure 9A shown, for Figure 9AFor the data 0 in it, the combined data stored in register R12 can be shifted left by 14 bits to adjust the position of sub-data 1 of data 0 to the high position, that is, the 16th and 15th bits of the target data 0. The combined data stored in register R8 is shifted left by 6 bits to expand the bit positions, and 0s are filled in the low positions. The AND module is used to extract the sign bit s and the highest bit of the data part of data 0 in R12, and the 0-7 bits of the data part of data 0 in register R8 are extracted. The OR module is used to splice the extracted bits to obtain the target data 0, and the target data 0 is 16 bits.
[0150] As Figure 9A shown, for Figure 9A the data 1 in it, the combined data stored in register R12 can be shifted left by 14 bits to adjust the position of sub-data 1 of data 1 to the high position, that is, the 16th and 15th bits of the target data 1. The combined data stored in register R8 is shifted left by 6 bits to expand the bit positions, and 0s are filled in the low positions. The AND module is used to extract the sign bit s and the highest bit of the data part of data 1 in R12, and the 0-7 bits of the data part of data 1 in register R8 are extracted. The OR module is used to splice the extracted bits to obtain the target data 1, and the target data 1 is 16 bits.
[0151] The target data 0 and the target data 1 can be stored in register R1.
[0152] For example, the above process can be expressed by the following formula:
[0153] R1 = ((R12 << 14) & 0xC000C000) | ((R8 << 6) & 0x3FC03FC0)
[0154] Here, R1 represents a storage unit for storing the target data 0 and the target data 1. Adopting the arrangement method as Figure 6B shown, each target data can be obtained through shift operations and Boolean operations during decoding.
[0155] Next, the decoding process is described for the Figure 7B shown compressed arrangement method.
[0156] Figure 9B is a schematic diagram of the decoding process provided by another embodiment of the present disclosure.
[0157] In Figure 9B , referring to the relevant content of Figure 7B , the register R4 stores the sub-data 1 part of the first data 0, including the sign bit s and the 8-10 bits of the data part, and the register R5 stores the sub-data 2 part of the first data 0, including the remaining 8 bits of the data part. In Figure 9BAmong them, 0 in the bit part represents the least significant bit, and S represents the sign bit, which is located at the most significant bit.
[0158] As Figure 9B shown, for Figure 9B the data 0 in, the combined data stored in register R4 can be shifted left by 12 bits to adjust the position of sub-data 1 of data 0 to the high position, that is, the 16th to 13th bits of the target data 0. The combined data stored in register R5 is shifted left by 4 bits to expand the bit positions, and 0 is filled in the low positions. Use an AND module to extract the sign bit s and the 8 - 10 bits of the data part of data 0 in R4, extract the 0 - 7 bits of the data part of data 0 in register R5, and use an OR module to splice the extracted bits to obtain the target data 0, and the target data 0 is 16 bits.
[0159] As Figure 9B shown, for Figure 9B the data 1 in, the combined data stored in register R4 can be shifted left by 12 bits to adjust the position of sub-data 1 of data 1 to the high position, that is, the 16th to 13th bits of the target data 1. The combined data stored in register R5 is shifted left by 4 bits to expand the bit positions, and 0 is filled in the low positions. Use an AND module to extract the sign bit s and the 8 - 10 bits of the data part of data 1 in R4, extract the 0 - 7 bits of the data part of data 1 in register R5, and use an OR module to splice the extracted bits to obtain the target data 1, and the target data 1 is 16 bits.
[0160] The target data 0 and the target data 1 can be stored in register R0.
[0161] For example, the above process can be expressed by the following formula:
[0162] R0 = ((R4 << 12) & 0xF000F000) | ((R5 << 4) & 0x0FF00FF0)
[0163] Here, R0 represents a storage unit for storing the target data 0 and the target data 1. Adopting the arrangement method as Figure 7B shown, each target data can be obtained through shift operations and Boolean operations during decoding.
[0164] Adopting the above-mentioned arrangement method, the target data 0 and the target data 1 can be obtained in one clock cycle. The throughput of the conversion module is M double words output per cycle, and M is the number of threads. Moreover, the acquisition of different target data can all use the process expressed by the above formula, and the same hardware logic or hardware module can be reused to obtain the best throughput and conversion efficiency. And only shift operations and Boolean operations are used in the conversion process, and the conversion process is easy to implement and time-consuming is short, which can further improve the conversion efficiency and achieve the balance of calculation and storage.
[0165] At least one embodiment of the present disclosure provides a computing core, which includes a conversion module. When data compression is required and the data bit width of the target format of data compression is not an integer multiple of 8, the conversion module can perform data compression through an encoding operation and store the combined data obtained by compression in a register. Thus, without changing the hardware architecture, the hardware can support the storage of non-standard data formats. The compressed data can be transmitted to other components through a network, and the transmission bandwidth can be saved during the transmission process. The compressed data can also be stored in a storage space such as a memory to save storage space, which can not only ensure the computing accuracy but also save storage space, achieving a balance between accuracy and storage space.
[0166] In at least one embodiment, a special compressed storage arrangement method is provided. This method can achieve data decompression and restoration through simple Boolean operations and shift operations during decoding, improving the decoding efficiency, improving the read and write efficiency, and further reducing the additional read and write overhead caused by compression.
[0167] At least one embodiment of the present disclosure further provides a data processing method. Figure 10 It is a schematic flowchart of the data processing method provided by at least one embodiment of the present disclosure.
[0168] As Figure 10 shown, the data processing method provided by at least one embodiment of the present disclosure at least includes step S10.
[0169] In step S10, in response to data compression of the received data in the first data format, and the data bit width m of the compressed data indicated by the data compression not being an integer multiple of 8, read the data to be compressed in the first data format from the register for an encoding operation, and store the multiple combined compressed data obtained through the encoding operation in the register.
[0170] The data bit width n indicated by the first data format is an integer multiple of 8 or a factor of 8, m and n are positive integers, and m is less than n.
[0171] For example, the encoding operation includes: converting the data to be compressed into the second data format indicated by the data compression, decomposing each data to be compressed after format conversion to obtain multiple sub-data, and compressing and combining the decomposed sub-data into multiple combined data in the first data format for storage in the register, where the bit width of each sub-data is an integer multiple of 8 or a factor of 8.
[0172] For the relevant descriptions of the first data format and the register, reference can be made to the relevant content of the aforementioned computing core 100, which will not be elaborated here.
[0173] For example, in some embodiments, step S10 may include: sending a first read address to a register and receiving a plurality of first data stored in T1 registers indicated by the first read address as the data to be compressed; converting the plurality of first data from a first data format to a second data format to obtain a plurality of second data respectively corresponding to the plurality of first data, and for each second data, decomposing the second data into L sub-data, where the sum of the bit widths of the L sub-data is equal to m, and L is a positive integer greater than 1; compressing and merging the plurality of sub-data obtained by decomposing the plurality of second data in the first data format and storing them in T2 registers, where T1 / T2 = n / m, and T1 and T2 are positive integers.
[0174] For the related description of "sending a first read address to a register and receiving a plurality of first data stored in T1 registers indicated by the first read address as the data to be compressed", reference may be made to the content of the aforementioned receiving unit 1011, which will not be elaborated here.
[0175] For the related description of "converting the plurality of first data from a first data format to a second data format to obtain a plurality of second data respectively corresponding to the plurality of first data, and for each second data, decomposing the second data into L sub-data", reference may be made to the content of the aforementioned decomposition unit, which will not be elaborated here.
[0176] For example, in some embodiments, compressing and merging the plurality of sub-data obtained by decomposing the plurality of second data in the first data format and storing them in T2 registers may include: splicing a sub-data into a merged data in the first data format and storing the merged data in the corresponding storage location in the T2 registers, where the a sub-data come from different a second data, and the a sub-data are in the same bit range in their respective second data, and a is a positive integer greater than 1 and equal to n / wb, where wb represents the bit width of any one of the a sub-data.
[0177] The compressed data can be transmitted to other computing cores, streaming processor clusters, etc. to save the transmission bandwidth during transmission, or the compressed data can be stored in local disks, hard disks and other memory or external memory to save storage space.
[0178] For example, a second data come from different vector registers and are stored in storage units controlled by the same thread, so that the data can flow within the thread, avoiding cross-thread interaction, reducing read / write overhead, and improving compression efficiency.
[0179] For example, the plurality of sub-data obtained by decomposing each second data includes a first sub-data, and the first sub-data includes the bits from the b1-th bit to the b2-th bit in the second data, where b1 and b2 are positive integers.
[0180] For any one of the T1 registers, b first data are stored in any storage unit in the any one register, where b is a positive integer, and b first sub-data obtained based on the b second data corresponding to the b first data are stored in the first storage unit of the T2 registers. The relative position relationship of the b first sub-data in the first storage unit is the same as the relative position relationship of the b first data in the any one storage unit. Among the b first sub-data, the least significant bits of two adjacent first sub-data differ by 16 bits.
[0181] For the merged data satisfying the above arrangement relationship, its compression efficiency is the highest. When reading, simple Boolean operations and shift operations can be used to complete data extraction. The implementation method is simple and time-consuming, and the logic for extracting different second data is consistent. Relevant modules or logic can be reused to achieve the highest throughput and improve the compression efficiency and decompression efficiency.
[0182] For the specific process of compression and merging, reference can be made to the relevant description of the aforementioned compression storage unit, which will not be elaborated here.
[0183] For example, as Figure 10 shown, the data processing method provided by at least one embodiment of the present disclosure further includes step S20.
[0184] In step S20, in response to data decompression, a decoding operation is performed on the data directly read from the register to convert the directly read data into the first data format and store it in the register.
[0185] For example, the decoding operation includes: splitting and reorganizing the directly read data according to the position relationship during storage, and performing bit expansion to obtain the data in the first data format.
[0186] For example, in some embodiments, step S20 may include: sending a second read address to the register and receiving a plurality of merged data stored in the T2 registers indicated by the second read address; splitting and reorganizing the plurality of merged data to obtain a plurality of second data, expanding the bit width of each second data to n bits to convert it into the first data format, where the data bit width of each second data is m; and sending the data after format conversion to the T1 registers, where T2 and T1 are positive integers, and T1 / T2 = n / m.
[0187] For example, in some embodiments, the splitting and recombination process is implemented through Boolean logic and shift logic. For example, splitting and recombining multiple merged data to obtain multiple second data, and expanding the bit width of each second data to n bits to convert to the first data format may include: for multiple merged data used to obtain target data, shifting the received multiple merged data to expand the bit positions and adjust the positions of the bit positions according to the positional relationship; extracting corresponding bit positions from the multiple merged data through Boolean logic according to the positional relationship; splicing the extracted corresponding bit positions through Boolean logic to obtain the target data, where the data format of the target data is the first data format.
[0188] In the above embodiments, the decoding process can reuse the same logic to obtain the best throughput and conversion efficiency, and only displacement operations and Boolean operations are used in the conversion process. The conversion process is easy to implement and takes a short time, which can further improve the conversion efficiency and achieve the balance of calculation and storage.
[0189] Figure 11 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 11 shown, the electronic device 300 is, for example, suitable for implementing the data processing method provided by the embodiments of the present disclosure. It should be noted that Figure 11 the components of the electronic device 300 shown are exemplary and not restrictive. According to actual application needs, the electronic device 300 may also have other components.
[0190] As Figure 11 shown, the electronic device 300 may include a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to non-temporary computer-readable instructions stored in the memory to implement various functions.
[0191] For example, when the computer-readable instructions are run by the processing device 301, one or more steps in the data processing method according to any of the above embodiments may be executed. It should be noted that the detailed description of the processing process of the data processing method can refer to the relevant descriptions in the embodiments of the above data processing method.
[0192] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 303 and / or cache memory, etc. For example, computer-readable instructions may be loaded from the storage device 308 into the random access memory (RAM) 303 to run the computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 302, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. Various application programs and various data may also be stored in the computer-readable storage media, such as style images, and various data used and / or generated by the application programs, etc.
[0193] For example, the processing device 301, read-only memory (ROM) 302, and random access memory (RAM) 303 are connected to each other via the bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0194] Generally, the following devices may be connected to the input / output (I / O) interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, a flash memory, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 11 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the electronic device 300 may alternatively implement or have more or fewer devices. For example, the processing device 301 may control other components in the electronic device 300 to perform desired functions. The processing device 301 may be a central processing unit (CPU), a tensor processing unit (TPU), or a graphics processing unit GPU, etc., which has instruction optimization capabilities and / or program execution capabilities. The central processing unit (CPU) may be of X86, ARM, RISC-V architecture, etc. The GPU may be directly integrated into the SOC, directly integrated onto the motherboard, or built into the northbridge chip of the motherboard.
[0195] Figure 12 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 12As shown, the storage medium 400 can be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 401 can be non-temporarily stored on the storage medium 400. For example, when the computer-readable instructions 401 are executed by a processor, one or more steps of the data processing method described above can be performed.
[0196] For example, the storage medium 400 can be applied to the electronic device 300. For example, the storage medium 400 can include the storage device 308 in the electronic device 300.
[0197] For example, the storage device can include any combination of one or more computer program products. The computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions can be stored on the computer-readable storage medium, and the processor can run the computer-readable instructions to implement various functions of the processor. Various application programs and various data can also be stored in the storage medium.
[0198] For example, the storage medium can include the memory card of a smart phone, the cache component of a tablet computer, the hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, and can also be other applicable storage media.
[0199] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0200] The units involved in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases.
[0201] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0202] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0203] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0204] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms for implementing the claims.
[0205] Regarding the present disclosure, the following points need to be noted:
[0206] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.
[0207] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0208] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the said claims.
Claims
1. A computing core includes a conversion module, The conversion module is configured to, in response to data compression being performed on data in a first data format obtained by processing by the computing core, and when the bit width m of the compressed data indicated by the data compression is not an integer multiple of 8, read the data to be compressed from a register to perform an encoding operation, and store multiple merged compressed data obtained through the encoding operation into the register, where wherein the data bit width n indicated by the first data format is an integer multiple of 8 or a factor of 8, m and n are positive integers and m is less than n; wherein the encoding operation includes: converting the data to be compressed into the second data format indicated by the data compression, decomposing each piece of data to be compressed after format conversion to obtain a plurality of sub-data, and compressing and merging all the decomposed sub-data into a plurality of merged data in the first data format for storage in the register, wherein the bit width of each sub-data is an integer multiple of 8 or a factor of 8.
2. The computing core according to claim 1, wherein, The conversion module includes a receiving unit, a decomposing unit, and a compression and storage unit; The receiving unit is configured to send a first read address to the register and receive a plurality of first data stored in T1 registers indicated by the first read address as the data to be compressed; The decomposing unit is configured to convert the plurality of first data from the first data format to the second data format, obtain a plurality of second data corresponding to the plurality of first data respectively, and for each second data, decompose the second data into L sub-data, wherein the sum of the bit widths of the L sub-data is equal to m, and L is a positive integer greater than 1; The compression and storage unit is configured to compress and merge the plurality of sub-data obtained by decomposing the plurality of second data in the first data format and store them in T2 registers, wherein T1 / T2 = n / m, and T1 and T2 are positive integers.
3. The computing core according to claim 2, wherein The compression and storage unit executes compressing and merging the plurality of sub-data obtained by decomposing the plurality of second data in the first data format and storing them in T2 registers, including performing the following operations: concatenating a sub-data into a merged data in the first data format and storing the merged data in the corresponding storage location in the T2 registers, wherein the a sub-data come from different a second data, and the a sub-data are in the same bit range in their respective second data, a is a positive integer greater than 1 and equal to n / wb, and wb represents the bit width of any one of the a sub-data.
4. The computing core according to claim 3, wherein, Both the T1 registers and the T2 registers are vector registers, each thread can operate multiple vector registers simultaneously, and multiple storage units in one vector register are controlled by multiple threads respectively, the a second data come from different vector registers and are stored in the storage units controlled by the same thread.
5. The computing core according to claim 2, wherein, The plurality of sub-data obtained by decomposing each second data includes a first sub-data, and the first sub-data includes the bits from the b1-th bit to the b2-th bit in the second data, where b1 and b2 are positive integers, for any one of the T1 registers, b first data are stored in any one of the storage units in the any one of the registers, where b is a positive integer, the b first sub-data obtained based on the b second data corresponding to the b first data are stored in the first storage unit in the T2 registers. The relative position relationship of the b first sub-data in the first storage unit is the same as the relative position relationship of the b first data in any storage unit. Among the b first sub-data, the least significant bits of two adjacent first sub-data differ by 16 bits.
6. The computing core according to claim 2, wherein The decomposition unit performs converting the multiple first data from the first data format to the second data format to obtain multiple second data respectively corresponding to the multiple first data, including performing the following operations: For each first data: Taking the least significant bit of the first data as the 1st bit, and retaining the nth bit to the (n - m + 1)th bit of the first data as the second data corresponding to the first data; or Performing a rounding operation on the (n - m + 2)th bit to the 1st bit of the first data based on a preset rounding rule, adjusting the nth bit to the (n - m + 1)th bit of the first data according to the rounding operation result, and retaining the adjusted nth bit to the (n - m + 1)th bit of the first data as the second data corresponding to the first data.
7. The computing core according to claim 2 is further configured to: after the receiving unit receives the multiple first data in the T1 registers, clear the data stored in the T1 registers.
8. The computing core according to claim 2, wherein, T1 × t = n, T2 × t = m, where t is a positive integer and represents the product of all common divisors of n and m.
9. The computing core according to claim 2, wherein, When L = 2, in response to m = 10, the bit widths of the 2 sub-data decomposed from the second data are 2 bits and 8 bits respectively. In response to m = 12, the bit widths of the 2 sub-data decomposed from the second data are 4 bits and 8 bits respectively.
10. The computing core according to claim 2 further includes a denormalization processing unit. The denormalization processing unit is configured to: In response to the number of bits of the exponent part indicated by the first data format being different from the number of bits of the exponent part indicated by the second data format, perform a denormalization processing operation on the multiple first data, and send the multiple first data after the denormalization processing operation to the decomposition unit for subsequent operations.
11. The computing core according to any one of claims 1-10, wherein, The conversion module is further configured to, in response to data decompression, perform a decoding operation on the data directly read from the register to convert the directly read data into the first data format and store it in the register for the computing core to read and use. Wherein, the decoding operation includes: Splitting and reorganizing the directly read data according to the position relationship during storage, and performing bit expansion to obtain the data in the first data format.
12. The computing core according to claim 11, wherein, The conversion module includes a receiving unit and a decoding unit. The receiving unit is configured to send a second read address to the register and receive multiple merged data stored in T2 registers indicated by the second read address. The decoding unit is configured to split and reorganize the multiple merged data to obtain multiple second data, expand the bit width of each second data to n bits to convert it into the first data format, wherein the data bit width of each second data is m. The receiving unit is further configured to send the data after format conversion to T1 registers, where T2 and T1 are positive integers, and T1 / T2 = n / m.
13. The computing core according to claim 12, wherein, The decoding unit is implemented by an OR module, an AND module, and a displacement module; For multiple merged data for obtaining target data, the displacement module is configured to shift the received multiple merged data to expand bit positions and adjust the positions of bit positions according to the position relationship; The AND module is configured to extract corresponding bit positions from the multiple merged data according to the position relationship; The OR module is configured to splice the extracted corresponding bit positions to obtain the target data, where the data format of the target data is the first data format.
14. The computing core according to claim 13, wherein, The OR module, the AND module, and the displacement module are multiplexed when different target data are obtained.
15. A data processing method, comprising: In response to data compression of the received data in the first data format, and when the data bit width m indicated by the data compression is not an integer multiple of 8, reading data to be compressed in the first data format from a register for an encoding operation, and storing multiple compressed merged data obtained through the encoding operation into the register, where the data bit width n indicated by the first data format is an integer multiple of 8 or a factor of 8, and m and n are positive integers and m is less than n; Wherein, the encoding operation comprises: Converting the data to be compressed into the second data format indicated by the data compression, decomposing each data to be compressed after format conversion to obtain multiple sub-data, and compressing and merging all the decomposed sub-data into multiple merged data in the first data format for storage in the register, where the bit width of each sub-data is an integer multiple of 8 or a factor of 8.
16. The data processing method according to claim 15, further comprising: In response to data decompression, performing a decoding operation on the data directly read from the register to convert the directly read data into the first data format and storing it in the register; Wherein, the decoding operation comprises: Splitting and reorganizing the directly read data according to the position relationship during storage and performing bit position expansion to obtain the data in the first data format.
17. An electronic device, comprising: A memory that stores computer-executable instructions non-transiently; A processor configured to run the computer-executable instructions, Wherein, when the computer-executable instructions are run by the processor, the data processing method according to claim 15 or 16 is implemented.
18. A non-transitory computer-readable storage medium, wherein, The non-transient computer-readable storage medium stores computer-executable instructions, When the computer-executable instructions are executed by a processor, the data processing method according to claim 15 or 16 is implemented.
Citation Information
Patent Citations
Data compression, transmission, receiving and uncompressing method and corresponding device
CN102790999A
Self-adaptive compression processing method and device for time series data
CN118868954A