Data processing method and device, chip and computer equipment
By converting low-precision data views into high-precision data, utilizing the address continuity of high-precision data, and adapting efficient memory access instructions, the problem of low efficiency in low-precision data memory access is solved, and computing performance and efficiency are improved.
Patent Information
- Application Number
- CN202511141056.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-15
AI Technical Summary
The data access method for low-precision data types affects the performance ceiling, resulting in low memory access efficiency and inability to fully utilize the potential of the hardware.
By converting the low-precision data block view into a high-precision data block, taking advantage of the address continuity of the high-precision data type, adapting to more efficient memory access instructions, and performing logical data processing during the calculation process, the result is finally redescribed as low-precision data, keeping the original data storage method unchanged.
It improves the memory access efficiency and computing performance of low-precision data, solves the memory access efficiency bottleneck in non-continuous memory address scenarios, and achieves a balance between memory access efficiency and versatility.
Smart Images

Figure CN120631450A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence chip technology, and in particular to a data processing method, apparatus, chip, and computer equipment. Background Art
[0002] In AI network computing, low-precision data types (such as BF16 and INT8) are widely used due to their high computational efficiency and small memory footprint. In the hardware architecture of related technologies, the data granularity of high-precision data types matches the hardware memory layout, so dedicated memory access instructions in related technologies can ensure efficient memory access for high-precision data types (such as FP32).
[0003] However, the data access method for low-precision data types affects the performance upper limit. Therefore, an effective data processing method is needed to optimize the memory access method for low-precision data. Summary of the Invention
[0004] The present application provides a data processing method, apparatus, chip and computer equipment, which solves the technical problem in the related art that the data access method for low-precision data types affects the performance upper limit, can improve the memory access efficiency for low-precision data, and achieve a balance between memory access efficiency and versatility.
[0005] In order to achieve the above objectives, the main technical solutions adopted in this application include: In a first aspect, an embodiment of the present application provides a data processing method, the method comprising: When it is determined that the address range corresponding to the low-precision data block is discontinuous in the memory according to the memory access granularity, data transmission mode, and data arrangement mode of the source tensor data, the low-precision data block using the first data type is described as a high-precision data block using the second data type; wherein the precision of the second data type is higher than the precision of the first data type; Performing a data processing-related operation based on the low-precision data block in a manner corresponding to the second data type to obtain a data processing result using the second data type; wherein the data processing-related operation includes at least one of data movement and data calculation; The data processing result is re-described as low-precision result data using the first data type.
[0006] Optionally, performing data processing-related operations based on the low-precision data block in a manner corresponding to the second data type to obtain a data processing result using the second data type includes: Reading the low-precision data block from the memory into a register using a loading method corresponding to the second data type to obtain high-precision equivalent data using the second data type; Data processing related operations are performed based on the high-precision equivalent data in the register to obtain a data processing result using the second data type.
[0007] Optionally, performing data processing related operations based on the high-precision equivalent data in the register to obtain a data processing result using the second data type includes: The high-precision equivalent data is stored from the register to the memory using a storage method corresponding to the second data type, and is used as a data processing result using the second data type.
[0008] Optionally, performing data processing related operations based on the high-precision equivalent data in the register to obtain a data processing result using the second data type includes: Describing the high-precision equivalent data as a combination of N low-precision equivalent data using the first data type; wherein N is an integer greater than or equal to 2; Performing a first type of calculation using the N low-precision equivalent data to obtain N first intermediate calculation results using the first data type; Based on the N first intermediate calculation results using the first data type, they are re-described and stored to obtain a data processing result using the second data type.
[0009] Optionally, the re-describing and storing the N first intermediate calculation results using the first data type to obtain a data processing result using the second data type includes: Describing the N first intermediate calculation results as a whole as a first calculation result using the second data type; The first calculation result is stored in the memory using a storage method corresponding to the second data type to obtain a data processing result using the second data type.
[0010] Optionally, performing data processing-related operations based on the low-precision data block in a manner corresponding to the second data type to obtain a data processing result using the second data type includes: Using a loading method corresponding to the second data type, the low-precision data block is read from the memory into a cache to obtain a high-precision data block using the second data type; describing the high-precision data block in the cache as a low-precision data block using the first data type; Using a loading method corresponding to the first data type, the low-precision data block is read from the cache into a register and subjected to a first type conversion to obtain high-precision data using a third data type; performing a second type of calculation and a second type of conversion based on the high-precision data in the register to obtain a second intermediate calculation result using the first data type; Using the cache as a transfer, re-description and storage are performed based on the second intermediate calculation result to obtain a data processing result using the second data type.
[0011] Optionally, the precision of the third data type is higher than the precision of the first data type, and the precision of the third data type is lower than the precision of the second data type.
[0012] Optionally, the second type of calculation is a high-precision calculation with accuracy requirements; performing the second type of calculation and the second type of conversion based on the high-precision data in the register to obtain a second intermediate calculation result using the first data type includes: performing high-precision calculation in the register using the high-precision data of the third data type to obtain a high-precision calculation result of the third data type; The high-precision calculation result is converted from the third data type to the first data type to obtain a second intermediate calculation result using the first data type.
[0013] Optionally, the re-describing and storing based on the second intermediate calculation result using the cache as a transfer to obtain a data processing result using the second data type includes: Using a storage method corresponding to the first data type, storing the second intermediate calculation result from the register to the cache; re-describe the second intermediate calculation result as a second calculation result using the second data type; The second calculation result is stored from the cache to the memory using a storage method corresponding to the second data type, thereby obtaining a data processing result using the second data type.
[0014] In a second aspect, an embodiment of the present application provides a data processing device, the device comprising: a data block description module, configured to, when it is determined based on the memory access granularity, data transmission mode, and data arrangement mode of the source tensor data that an address range corresponding to the low-precision data block is discontinuous in the memory, describe the low-precision data block of a first data type as a high-precision data block of a second data type; wherein the second data type has a higher precision than the first data type; a data processing operation module, configured to perform data processing-related operations based on the low-precision data block in a manner corresponding to the second data type, to obtain a data processing result using the second data type; wherein the data processing-related operations include at least one of data movement and data calculation; A result redescription module is configured to redescribe the data processing result as low-precision result data using the first data type.
[0015] In a third aspect, an embodiment of the present application provides a chip, including: processor; a memory for storing processor-executable instructions; The processor is configured to perform any of the above-mentioned data processing methods when executing the instructions stored in the memory.
[0016] In a fourth aspect, an embodiment of the present application provides a computer device comprising a memory, an artificial intelligence chip, and a computer program stored in the memory and executable on the artificial intelligence chip, wherein the artificial intelligence chip implements any of the above-described data processing methods when executing the computer program.
[0017] In an embodiment of the present application, first, it is determined whether the address of the low-precision data block in the memory is continuous based on the memory access granularity, data transmission method and data arrangement method of the source tensor data; if the address is not continuous, the low-precision data block is described as a high-precision data block through the view conversion mechanism, thereby utilizing the address continuity advantage of the high-precision data type to adapt to more efficient memory access instructions; then, data processing related operations are performed based on the low-precision data block to obtain the data processing result of the high-precision data type; finally, the data processing result of the high-precision data type is redescribed as low-precision result data, thereby solving the memory access efficiency bottleneck problem of low-precision tensor data in the non-continuous memory address scenario without changing the original data storage method, thereby improving memory access efficiency and computing performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 A schematic structural diagram of a general-purpose graphics processor provided in an embodiment of the present application; Figure 2A schematic diagram of the data arrangement of tensor data provided in an embodiment of the present application; Figure 3 A flowchart of a data processing method provided in an embodiment of the present application; Figure 4 A flowchart of a data processing method provided in one embodiment of the present application; Figure 5 The data processing process in the memory access scenario provided in the embodiment of the present application; Figure 6 A flowchart of a data processing method provided in yet another embodiment of the present application; Figure 7 A flowchart of a data processing method provided in another embodiment of the present application; Figure 8 The data processing process in the low-precision computing scenario provided by the embodiment of this application; Figure 9 A flowchart of the data calculation process provided in the embodiment of the present application; Figure 10 The data processing process in the high-precision computing scenario provided by the embodiments of this application; Figure 11 A schematic diagram of the framework of a data processing device provided in an embodiment of the present application; Figure 12 A schematic structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0021] The data processing method of this application is based on artificial intelligence (AI). Artificial intelligence is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence.
[0022] The data processing method of this application can be applied to an artificial intelligence processor, which can be any of a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), or a GPGPU (General-Purpose Graphics Processing Unit). A general-purpose graphics processor (GPGPU) is used as an example for illustration.
[0023] Figure 1 The following is a schematic diagram of the structure of a general-purpose graphics processing unit (GPGPU). Figure 1 ,A general graphics processor is actually an array of programmable multiprocessors. For example, the programmable multiprocessor can be a streaming processor cluster (SPC), including Figure 1 The stream processor clusters shown are 1, M-1, and M, where M is a positive integer greater than 1. In general-purpose graphics processors, one stream processor cluster processes one computational task, or multiple stream processor clusters can process one computational task. Taking stream processor cluster 1 as an example, a stream processor cluster includes multiple execution units (EUs). Each processing unit includes an arithmetic logic unit (ALU) and a floating-point unit (FPU), and is responsible for executing specific computational tasks. Each processing unit includes thread local registers (TLRs) for storing source and destination data related to the computational task. The global memory buffer (GMB) connects the EUs to high-bandwidth memory (HBM) for data transfer, improving data transfer efficiency between the EUs and HBM. The HBM provides large-capacity, low-latency storage to support the high-performance computing requirements of GPUs. Instructions enter the SPC from the outside and are dispatched to each EU for execution. Input data is read from the HBM, transferred to the SPC via the GMB, and then distributed to the EUs for processing. The results (output data) processed by EU are returned to GMB through SPC and finally written to HBM.
[0024] See also Figure 2 , the following explanation is given by taking tensor data under the column-major layout as an example. Figure 2 The data layout of FP32, BF16, S8 / U8 and other data types under the colmajor layout is shown in the figure. Figure 2The darker areas in the diagram represent data blocks accessed with 256B of memory at a time. For example, for the FP32 data type, two 128B blocks are arranged consecutively; for the BF16 data type, the first 128B (64B × 2) and the last 128B (64B × 2) are arranged alternately; and for the S8 / U8 data type, the first 128B (32B × 4) and the last 128B (32B × 4) are arranged alternately.
[0025] In some cases, memory access at a higher granularity, such as at least 256 bytes, is required to improve hardware performance. This is because cachelines exist on both the L2 cache and HBM on the bus, providing the minimum granularity for reads. It is worth noting that for memory access operations with a matrix colmajor layout, the bursting methods and data arrangement characteristics of different data types directly affect performance. Hardware requires a data access granularity of at least 256 bytes to improve efficiency, but non-contiguous address intervals limit the actual performance ceiling. For details, please refer to Table 1.
[0026] Table 1 Regarding the FP32 data type, the data is densely arranged and is fully adapted to hardware instructions, that is, there is no need to cross 2048 bytes, and the memory access efficiency is high.
[0027] For the BF16 data type, the interval between the first 128 bytes (64 bytes × 2) and the last 128 bytes (64 bytes × 2) is 2048 bytes. The entire data must span this interval, so its usage scenarios are limited.
[0028] For data types S8 / U8, the interval between the first 128B (32B×4) and the last 128B (32B×4) is 2048 bytes, and the entire data must also span this interval, limiting its versatility.
[0029] Analysis revealed that, under the colmajor layout, 256-byte burst reads of low-precision data (BF16 / S8 / U8) require a 2048-byte gap, preventing the hardware from efficiently loading data using sequential memory access instructions (such as burst2 / burst4). For example, using BF16 as an example, a 256-byte read requires accessing four 64-byte data blocks at addresses (0,0), (1,0), (32,0), and (33,0). The first two blocks (0,0) and (1,0) are separated by a 2048-byte gap from the last two blocks ((32,0) and (33,0). For example, taking S8 / U8 as an example, a 256-byte read requires accessing 8 32-B data blocks with addresses (0,0), (1,0), (2,0), (3,0), (64,0), (65,0), (66,0), and (67,0). The first four segments (0,0), (1,0), (2,0), and (3,0) are separated by 2048 bytes from the last four segments (64,0), (65,0), (66,0), and (67,0).
[0030] It should be noted that in different chips, the memory layout and arrangement of data of different data types stored on the chip are different, which means that different data types have limited transmission modes for data input and output. Burst mode is a data transmission method. In Burst mode, after the starting address and concurrent length (Burstlengths) are specified, the data transmission process will automatically start from the starting address and perform continuous read / write operations on the same number of storage units. In Burst mode, Burstm can indicate loading m data sub-blocks at a time. For example, Burst4 means loading 4 data sub-blocks at a time, and Burst8 means loading 8 data sub-blocks at a time. The Burst modes that can be used for different data types are also restricted accordingly.
[0031] When the aforementioned gap exists for data types BF16 or S8 / U8, the overall data size must also span 2048 bytes during the segmentation process, limiting its use in certain scenarios. Furthermore, this address gap limit affects performance during actual data access.
[0032] Based on this, an embodiment of the present application provides a data processing method that is suitable for non-continuous address scenarios of low-precision data under some data layout methods (such as column-major arrangement). Specifically, low-precision data is described as high-precision data through view operations, and the address continuity advantage of high-precision data is utilized to be compatible with the dedicated memory access instructions of the relevant hardware architecture, thereby reducing performance loss. It should be noted that the view operation in the embodiment of the present application can be regarded as a typecast view, which can perform transfer operations or calculation operations with higher-precision data types by updating the interpretation method or description of the data, and does not involve actual data copying or physical data conversion; or, the view operation in the embodiment of the present application can be understood as logically describing low-precision data as high-precision data without changing the actual physical arrangement and content of the low-precision data.
[0033] The embodiment of the present application processes low-precision data as another data type (a data type with higher precision) without any change in the underlying data, thereby logically using memory access instructions that are more suitable for high-precision data types to move data, and can take advantage of the address continuity of high-precision data to achieve efficient memory access. After the calculation is completed, the calculation results are re-described or reinterpreted, and the actual description of the low-precision data is restored and stored back in the memory, thereby achieving a dual improvement in memory access efficiency and versatility.
[0034] According to an embodiment of the present application, a data processing method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here. Figure 3 , Figure 3 The data processing method in the embodiment of the present application is shown as a flow chart. The data processing method includes the following steps: S110. When it is determined that the address range corresponding to the low-precision data block is discontinuous in the memory based on the memory access granularity, data transmission method and data arrangement method of the source tensor data, the low-precision data block using the first data type is described as a high-precision data block using the second data type.
[0035] The precision of the second data type is higher than that of the first data type. The first data type refers to a data type used to represent a low-precision data block. For example, the first data type may be BF16, FP16, INT8, or UINT8. The second data type refers to a data type used to logically reinterpret the low-precision data block, and its precision is higher than the first data type. For example, the second data type may be FP32, which has higher numerical precision than BF16 or FP16. By describing the low-precision data block as a high-precision data type, the address continuity advantage of the high-precision data type can be utilized, thereby adapting to more efficient memory access instructions.
[0036] Source tensor data can be a collection of tensor-type data input to a data processing pipeline, typically used in scenarios such as deep learning, image processing, and matrix operations. For example, source tensor data can be weight parameters or activation value tensors in a convolutional neural network, or pixel data tensors in image processing. Source tensor data can be stored in memory using various data layouts.
[0037] Data layout refers to how tensor data is organized and stored in memory. Examples include row-major and column-major layouts. Different data layouts affect data continuity in memory, which in turn affects memory access efficiency. In column-major layout, when the memory access granularity is large, cross-column accesses are prone to occur, resulting in discontinuous data addresses.
[0038] A low-precision data block refers to a local block of source tensor data stored using a low-precision data type. For example, a low-precision data block can be a block of data types such as BF16, FP16, or INT8 (signed 8-bit integer). Low-precision data blocks are often used to save memory bandwidth and improve computational throughput, but they can lead to address discontinuity under certain data arrangements.
[0039] Specifically, whether the addresses of low-precision data blocks in memory are continuous is determined based on the memory access granularity, data transmission method, and data arrangement method of the source tensor data; if the addresses are not continuous, the low-precision data blocks are described as high-precision data blocks through view operations, thereby taking advantage of the address continuity of high-precision data types.
[0040] It is understandable that the description operation (view) in this embodiment can be understood as a logical reinterpretation or description of a low-precision data block through view conversion to adapt to memory access instructions of a high-precision data type, rather than a physical data operation. In other words, a data type with discontinuous addresses during memory access is described as a data type with continuous addresses during memory access. By way of example, a low-precision data block using a first data type is described as a high-precision data block using a second data type, and a BF16 data block is viewed as an FP32 data block, an FP16 data block is viewed as an FP32 data block, an INT8 data block is viewed as an FP32 data block, a Uint8 data block is viewed as an FP32 data block, and an FP8 data block is viewed as an FP32 data block.
[0041] It should be noted that the view operation in the embodiments of the present application and the view operator in the related art both modify the description of the data without changing the actual content of the data, but the two are different. In the related art, the view operator is an operator within the computational graph, while the view operation in the embodiments of the present application uses a non-existent typecastview to optimize performance, that is, the process of converting a view or representation of a certain data type into a view or representation of another data type. The view operation in the embodiments of the present application can essentially be understood as creating a new descriptor to describe or interpret low-precision data as high-precision data. Logical type conversion is achieved through the DATA_TYPE field of the descriptor, redefining the access rules for low-precision data blocks.
[0042] It is understandable that in deep learning and hardware acceleration scenarios, the tensor descriptor (Tensor Desc) is a specific implementation of the descriptor, which is used to define the key information of the tensor (Tensor), such as: data type (such as BF16, FP32, INT8), arrangement layout (such as row-first arrangement, column-first arrangement), dimension information (number of channels, width, height, depth, number of batches), stride (Stride), padding (padValue), etc.
[0043] S120 . Perform data processing-related operations based on the low-precision data block in a manner corresponding to the second data type to obtain a data processing result using the second data type.
[0044] The data processing related operations include at least one of data handling and data calculation, and may include but are not limited to matrix operations, data handling, etc.
[0045] The data processing result refers to the result obtained after the data processing operation, and its data type is the second data type. For example, after using FP32 precision for data transfer or calculation, the data processing result is also FP32 precision data, which has higher numerical precision than the original BF16 or INT8 data.
[0046] Specifically, after the low-precision data block is logically described as a high-precision data block, the data is calculated or processed using a data processing method corresponding to the high-precision data type to obtain a data processing result using the high-precision data type.
[0047] S130: Re-describe the data processing result as low-precision result data using the first data type.
[0048] The low-precision result data may be low-precision output data of the first data type. Specifically, the data processing result of the high-precision data type is reinterpreted as a low-precision data type through view conversion. For example, the FP32 precision output data is reinterpreted as BF16 or INT8 precision data through a view operation, and written back to the memory using the processing method corresponding to the first data type, thereby ensuring compatibility with subsequent processes.
[0049] In the above embodiment, first, it is determined whether the address of the low-precision data block in the memory is continuous based on the memory access granularity, data transmission method and data arrangement method of the source tensor data; if the address is not continuous, the low-precision data block is described as a high-precision data block through the view conversion mechanism, thereby taking advantage of the address continuity advantage of the high-precision data type and adapting to more efficient memory access instructions; then, data processing related operations are performed based on the low-precision data block to obtain the data processing result of the high-precision data type; finally, the data processing result of the high-precision data type is redescribed as low-precision result data, thereby solving the memory access efficiency bottleneck problem of low-precision tensor data in the non-continuous memory address scenario without changing the original data storage method, thereby improving memory access efficiency and computing performance.
[0050] In some embodiments, see Figure 4 , performing data processing-related operations based on the low-precision data block in a manner corresponding to the second data type to obtain a data processing result using the second data type, including: S310 , using a loading method corresponding to the second data type, to read the low-precision data block from the memory into the register, and obtain high-precision equivalent data using the second data type.
[0051] S320: Perform data processing related operations based on the high-precision equivalent data in the register to obtain a data processing result using the second data type.
[0052] Among them, the loading method corresponding to the second data type can be to select the memory access granularity and data transmission instructions suitable for the second data type according to the data access characteristics of the second data type, so as to continuously load low-precision data blocks from the memory into the register.
[0053] High-precision equivalent data can be the data format presented in registers after logically interpreting or describing a low-precision data block as a high-precision data type through view conversion. It should be noted that although the underlying data remains unchanged, it is processed as high-precision data in the register, allowing it to be read using read instructions for high-precision data types.
[0054] Specifically, first, a loading method corresponding to the second data type is used to read the low-precision data block from the memory into the register to obtain high-precision equivalent data of the second data type; then, data processing-related operations are performed based on the high-precision equivalent data in the register, and finally a data processing result of the second data type is obtained. By introducing a view conversion mechanism in the data loading stage, low-precision data blocks that originally had discontinuous addresses due to data arrangement can be efficiently loaded in an equivalent form of high-precision data type, thereby improving the overall data memory access efficiency. For example, when the first data type is BF16 and the second data type is FP32, the burst2 memory access granularity (256Bytes read each time) and the FP32-specific load instruction can be selected to load the data blocks originally stored in the BF16 format with discontinuous addresses into the register in the FP32 format.
[0055] Furthermore, data processing related operations are performed based on the high-precision equivalent data in the register to obtain a data processing result using the second data type, including: using a storage method corresponding to the second data type to store the high-precision equivalent data from the register to the memory, and using it as the data processing result using the second data type.
[0056] Among them, the storage method corresponding to the second data type refers to selecting a memory access granularity and data transfer instruction suitable for the data type according to the data storage characteristics of the second data type, for writing high-precision equivalent data in the register to the memory. Specifically, the high-precision equivalent data is transferred from the register to the memory through the memory access instruction. For example, in a GPU or AI acceleration chip, the high-precision equivalent data is written back from the register to the memory of the chip. In this embodiment, the data write back process is executed based on the storage method of the second data type, thereby ensuring that the data can adapt to the optimized instruction set of the hardware platform when it is written back to the memory.
[0057] In the above embodiment, by introducing a storage method that matches the second data type after data processing is completed, and by adopting a storage method corresponding to the second data type, the problem of not being able to use efficient storage instructions when the address is discontinuous due to data arrangement is solved; by storing the results in memory with a high-precision data type, it is ensured that high-precision equivalent data can be written back to the memory in an efficient and compatible manner, thereby completing the entire data processing process.
[0058] See also Figure 5 This section illustrates the data processing process in a memory access scenario. When only data needs to be moved, using BF16 as an example, 1) view the low-precision BF16 data into high-precision FP32 data; 2) read the high-precision FP32 data into a register using the corresponding FP32 load method; 3) write the data to the HBM using the corresponding FP32 store method; and 4) re-view the high-precision data in the HBM back to its original data type (i.e., BF16). This completes the data move.
[0059] In some embodiments, see Figure 6 , performing data processing related operations based on the high-precision equivalent data in the register to obtain a data processing result using the second data type, including: S610: Describe the high-precision equivalent data as a combination of N low-precision equivalent data of the first data type.
[0060] Wherein, N is an integer greater than or equal to 2. N can be the number of low-precision equivalent data into which the high-precision equivalent data is split. The value of N depends on the byte alignment relationship between the high-precision data type (second data type) and the low-precision data type (first data type). For example, when the second data type is FP32 (4 Byte) and the first data type is BF16 (2 Byte), N=2; when the second data type is FP32 (4 Byte) and the first data type is INT8 (1 Byte), N=4. The selection of N must ensure that the combination of low-precision equivalent data after splitting can fully cover the numerical range of the high-precision equivalent data, while adapting to the computing unit characteristics of the hardware platform.
[0061] Specifically, through the view conversion mechanism, the high-precision equivalent data in the register is logically split into a combination of multiple low-precision data blocks. For example, when the high-precision equivalent data is of FP32 type and the first data type is of BF16 type, each FP32 data can be split into a combination of two BF16 data (i.e., N=2). Although the underlying data does not physically change, it is interpreted as a combination of low-precision data in the register, allowing low-precision calculation instructions to be called for processing.
[0062] S620: Perform a first type of calculation using N low-precision equivalent data to obtain N first intermediate calculation results using the first data type.
[0063] Among them, the first type of calculation can be to calculate the low-precision equivalent data after splitting through the calculation instructions of the low-precision data type. For example, the low-precision equivalent data after splitting can be calculated in parallel by calling the addition instruction of BF16 or INT8. The first type of calculation can be understood as a low-precision calculation without accuracy requirements. The first intermediate calculation result can be a calculation result that is processed by the low-precision calculation instruction and is still stored in the first data type.
[0064] Specifically, the calculation instructions corresponding to the first data type are used to calculate the split low-precision equivalent data, and the resulting first intermediate result is still of the first data type. For example, if the BF16 calculation instructions are used to calculate the split BF16 equivalent data, the resulting intermediate calculation result is still of the BF16 type. This process ensures the low-precision characteristics of the calculation stage, thereby fully leveraging the hardware platform's optimization capabilities for low-precision calculations.
[0065] S630: Re-describe and store the N first intermediate calculation results using the first data type to obtain a data processing result using the second data type.
[0066] Specifically, the low-precision first intermediate calculation result is logically reinterpreted as a high-precision data type through the view conversion mechanism, and written to the memory according to the storage method of the second data type, thereby obtaining a data processing result using the second data type. For example, the intermediate calculation results of two BF16s are combined into an FP32 data processing result, and written to the memory through the FP32 storage instruction. This process achieves a seamless connection between low-precision calculations and high-precision results through the view conversion mechanism, while maintaining the compatibility of the data processing flow.
[0067] Further, see Figure 7 , re-describing and storing the N first intermediate calculation results using the first data type to obtain a data processing result using the second data type, including: S710: Describe the N first intermediate calculation results as a whole as a first calculation result using a second data type.
[0068] S720: Use a storage method corresponding to the second data type to store the first calculation result in the memory to obtain a data processing result using the second data type.
[0069] Specifically, first, N low-precision first intermediate calculation results are logically combined into a high-precision first calculation result through view conversion; secondly, a storage method corresponding to the second data type is selected, and the high-precision first calculation result in the register is written into the memory to obtain a high-precision data type evil data processing result.
[0070] In the above embodiment, by logically merging the low-precision first intermediate calculation result into the first calculation result of the high-precision data type through view conversion and adapting the storage method of the high-precision data type, the problem of not being able to use efficient storage instructions when the address is discontinuous due to data arrangement is solved, thereby achieving both low-precision calculation efficiency and high-precision data processing result output.
[0071] See also Figure 8 This section illustrates the data processing process in low-precision computing scenarios. When low-precision computing is required, using BF16 as an example, 1) the low-precision BF16 data in the HBM is viewed as high-precision FP32 data; 2) the FP32 high-precision data is read from the HBM into a register using the load method corresponding to FP32; 3) the FP32 high-precision data is viewed as a combination of two BF16 low-precision data, and low-precision computing is performed to obtain two BF16 low-precision computing results; 4) the two BF16 low-precision computing results are viewed as high-precision FP32 results; 5) the FP32 high-precision computing result is written back to the HBM using the store method corresponding to FP32; 6) the FP32 high-precision computing result is viewed back to the original data type (i.e., BF16). At this point, the low-precision computing is completed.
[0072] In some embodiments, see Figure 9 , performing data processing-related operations based on the low-precision data block in a manner corresponding to the second data type to obtain a data processing result using the second data type, including: S910 . Read the low-precision data block from the memory into the cache using a loading method corresponding to the second data type to obtain a high-precision data block using the second data type.
[0073] The cache refers to an intermediate storage unit located between the memory and the register, which is usually used to temporarily store frequently accessed data to reduce memory access latency. The cache can be GMB.
[0074] In some cases, since the data precision of the register processing of the low-precision data block is higher than the precision of the first data type, if the low-precision data block is directly loaded into the register, the calculation precision will be lost, so it is necessary to use the cache for transit operation. Specifically, according to the data access characteristics of the second data type, the memory access granularity and data transfer instructions suitable for the second data type are selected to load the low-precision data block from the memory into the cache. For example, when the second data type is FP32, the burst2 memory access granularity (256Bytes read each time) and the FP32-specific load instructions can be selected to load the low-precision data block originally stored in the BF16 format and with discontinuous addresses into the cache in FP32 format. This process uses the view conversion mechanism to logically interpret the low-precision data as high-precision data, thereby adapting to the load instructions of the high-precision data type. S920, describe the high-precision data block in the cache as a low-precision data block using the first data type.
[0075] Specifically, through the view conversion mechanism, high-precision data blocks in the cache are logically reinterpreted as low-precision data types. For example, an FP32 data block in the cache is logically interpreted as a BF16 data block. Although the underlying data remains unchanged, it is treated as a low-precision data block in the cache, thereby adapting to load instructions of the low-precision data type. This process provides the basis for subsequent loading of data from the cache into registers and performing type conversion.
[0076] S930 . Read the low-precision data block from the cache into the register using a loading method corresponding to the first data type and perform a first type conversion to obtain high-precision data using a third data type.
[0077] The precision of the third data type is higher than that of the first data type, and the precision of the third data type is lower than that of the second data type. The first type conversion is a conversion operation that converts a low-precision data block into high-precision data of the third data type through a hardware mechanism after loading it into a register.
[0078] In some cases, in order to ensure calculation accuracy, the cache is used for a transfer operation. Further, after the transfer operation, a first type conversion is performed in the process of loading the low-precision data block from the cache to the register. Specifically, according to the data access characteristics of the first data type, a memory access granularity and data transfer instruction suitable for the first data type are selected to load the low-precision data block in the cache into the register. In the register, the low-precision data is type-converted to obtain high-precision data using a third data type. Exemplarily, the third data type can be BF20, which has higher precision than BF16 or INT8, and lower precision than the second data type (such as FP32).
[0079] S940 , performing a second type of calculation and a second type of conversion based on the high-precision data in the register to obtain a second intermediate calculation result using the first data type.
[0080] The second type of calculation may be a high-precision calculation. Specifically, a high-precision calculation operation is performed on high-precision data of the third data type in a register and converted to the first data type to obtain a second intermediate calculation result using the first data type. For example, accumulation is performed on BF20 data in a register and converted to BF16 or INT8 precision data through a hardware mechanism.
[0081] Furthermore, the second type of calculation is a high-precision calculation with precision requirements. Performing the second type of calculation and the second type of conversion based on the high-precision data in the register to obtain a second intermediate calculation result using the first data type includes: performing the high-precision calculation in the register using the high-precision data using the third data type to obtain a high-precision calculation result using the third data type; and converting the high-precision calculation result from the third data type to the first data type to obtain a second intermediate calculation result using the first data type.
[0082] S950: Re-describe and store the second intermediate calculation result using the cache as a transfer to obtain a data processing result using the second data type.
[0083] Specifically, through the view conversion mechanism, the low-precision second intermediate calculation result is logically reinterpreted as data of the second data type, and is written into the memory through cache transfer to obtain a data processing result of the second data type.
[0084] Furthermore, with the cache as a transit, the second intermediate calculation result is re-described and stored to obtain a data processing result using the second data type, including: using a storage method corresponding to the first data type to store the second intermediate calculation result from the register to the cache; re-describing the second intermediate calculation result as a second calculation result using the second data type; using a storage method corresponding to the second data type to store the second calculation result from the cache to the memory to obtain a data processing result using the second data type.
[0085] In the above embodiment, first, a low-precision data block is loaded into the cache as a high-precision data type. Subsequently, the high-precision data block in the cache is logically interpreted as a low-precision data block and loaded into a register for a first type conversion. Next, a second type of calculation and a second type of conversion are performed in the register to generate a low-precision second intermediate calculation result. Finally, the second intermediate calculation result is redescribed as a high-precision data type through cache transfer and written to memory. This process, through multi-level type conversion and cache transfer, achieves efficient loading of low-precision data, high-precision calculation, and result reorganization, while maintaining the compatibility of the data processing flow.
[0086] See also Figure 10 , exemplarily illustrates the data processing process in high-precision computing scenarios. When data requires high-precision computing, taking BF16 as an example, 1) view the low-precision data of type BF16 in HBM into high-precision data of type FP32; 2) use the corresponding load method of FP32 to load the high-precision data of type FP32 from HBM into GMB; 3) in GMB, view the high-precision data of FP32 into a combination of two low-precision data of type BF16; 4) use the corresponding load method of BF16 to load the low-precision data of type BF16 into the register (register), and convert the data from type BF16 to type BF20. 5) In registers, high-precision calculations are performed using BF20 data to obtain high-precision results. 6) The high-precision results are converted to BF16 and stored in the GMB using the store method corresponding to the BF16 type. 7) The BF16 results in the GMB are described as FP32 results. 8) The FP32 results are stored in the HBM using the store method corresponding to FP32. 9) The FP32 results are viewed back to the original data type (i.e., BF16). This completes the high-precision calculation. It should be noted that since the low-precision BF16 data in registers is processed at BF20 precision, directly loading it into the register would result in precision loss. Therefore, the GMB is used for intermediate transfer. After this transfer, the data is converted from BF16 to BF20 during the register load process.
[0087] The present application provides a data processing device. Figure 11 12 is a schematic diagram of a data processing apparatus according to an embodiment of the present invention. The data processing apparatus 1200 may include a data block description module 1210 , a data processing operation module 1220 , and a result redescription module 1230 .
[0088] The data block description module 1210 is configured to describe a low-precision data block of a first data type as a high-precision data block of a second data type, when it is determined that an address range corresponding to the low-precision data block is discontinuous in memory based on a memory access granularity, a data transmission mode, and a data arrangement mode of the source tensor data; wherein the second data type has a higher precision than the first data type; a data processing operation module 1220 configured to perform data processing-related operations based on the low-precision data block in a manner corresponding to the second data type, to obtain a data processing result of the second data type; wherein the data processing-related operations include at least one of data movement and data calculation; The result redescription module 1230 is configured to redescribe the data processing result as low-precision result data of the first data type.
[0089] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0090] An embodiment of the present application provides a chip, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0091] The processor involved in the embodiments of the present application can be any one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), an NPU (Neural Network Processing Unit), a DPU (Deep Learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit), and is determined when the embodiments of the present application are applied to a specific product or technology.
[0092] Embodiments of the present application provide a non-transitory computer-readable storage medium. The storage medium may be a non-transitory computer-readable storage medium that can non-transitorily store one or more computer-readable instructions. For example, when the computer-readable instructions are executed by a processor, one or more steps of the aforementioned data processing method may be performed. The storage medium may be applied to an electronic device, for example, the storage medium may include a storage device in the electronic device.
[0093] The storage device may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, a flash memory, etc. One or more computer-readable instructions may be stored on the computer-readable storage medium, and the processor may execute the computer-readable instructions to implement various functions of the processor. The storage medium may also store various application programs and various data.
[0094] The storage medium may include a memory card of a smart phone, a cache component of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a flash memory, or any combination of the above storage media, or other applicable storage media.
[0095] The present application embodiment provides a computer device, such as Figure 12 As shown, it includes at least one artificial intelligence chip 1510 and a memory 1520 connected to the at least one artificial intelligence chip 1510. The specific connection medium between the artificial intelligence chip 1510 and the memory 1520 is not limited in the embodiment of the present application. Figure 12 For example, the artificial intelligence chip 1510 and the memory 1520 are connected via a bus. Buses can be divided into address buses, data buses, control buses, etc. In the embodiment of the present application, the memory 1520 stores instructions that can be executed by at least one artificial intelligence chip 1510. By executing the instructions stored in the memory 1520, the at least one artificial intelligence chip 1510 can perform the steps of the above-mentioned data processing method.
[0096] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0097] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0098] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0099] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0100] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
[0101] Although the embodiments of the present application have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the appended claims.
Claims
1. A data processing method, characterized in that: The method comprises: When it is determined that the address range corresponding to the low-precision data block is discontinuous in the memory according to the memory access granularity, data transmission mode, and data arrangement mode of the source tensor data, the low-precision data block using the first data type is described as a high-precision data block using the second data type; wherein the precision of the second data type is higher than the precision of the first data type; Performing a data processing-related operation based on the low-precision data block in a manner corresponding to the second data type to obtain a data processing result using the second data type; wherein the data processing-related operation includes at least one of data movement and data calculation; The data processing result is re-described as low-precision result data using the first data type.
2. The method according to claim 1, characterized in that The performing data processing-related operations based on the low-precision data block in a manner corresponding to the second data type to obtain a data processing result using the second data type includes: Reading the low-precision data block from the memory into a register using a loading method corresponding to the second data type to obtain high-precision equivalent data using the second data type; Data processing related operations are performed based on the high-precision equivalent data in the register to obtain a data processing result using the second data type.
3. The method according to claim 2, characterized in that The performing data processing related operations based on the high-precision equivalent data in the register to obtain a data processing result using the second data type includes: The high-precision equivalent data is stored from the register to the memory using a storage method corresponding to the second data type, and is used as a data processing result using the second data type.
4. The method according to claim 2, characterized in that The performing data processing related operations based on the high-precision equivalent data in the register to obtain a data processing result using the second data type includes: Describing the high-precision equivalent data as a combination of N low-precision equivalent data using the first data type; wherein N is an integer greater than or equal to 2; Performing a first type of calculation using the N low-precision equivalent data to obtain N first intermediate calculation results using the first data type; Based on the N first intermediate calculation results using the first data type, they are re-described and stored to obtain a data processing result using the second data type.
5. The method according to claim 4, characterized in that The re-describing and storing the N first intermediate calculation results using the first data type to obtain a data processing result using the second data type includes: Describing the N first intermediate calculation results as a whole as a first calculation result using the second data type; The first calculation result is stored in the memory using a storage method corresponding to the second data type to obtain a data processing result using the second data type.
6. The method according to claim 1, wherein The performing data processing-related operations based on the low-precision data block in a manner corresponding to the second data type to obtain a data processing result using the second data type includes: Using a loading method corresponding to the second data type, the low-precision data block is read from the memory into a cache to obtain a high-precision data block using the second data type; describing the high-precision data block in the cache as a low-precision data block using the first data type; Using a loading method corresponding to the first data type, the low-precision data block is read from the cache into a register and subjected to a first type conversion to obtain high-precision data using a third data type; performing a second type of calculation and a second type of conversion based on the high-precision data in the register to obtain a second intermediate calculation result using the first data type; Using the cache as a transfer, re-description and storage are performed based on the second intermediate calculation result to obtain a data processing result using the second data type.
7. The method according to claim 6, characterized in that The precision of the third data type is higher than the precision of the first data type, and the precision of the third data type is lower than the precision of the second data type.
8. The method according to claim 6, characterized in that The second type of calculation is a high-precision calculation with accuracy requirements; performing the second type of calculation and the second type of conversion based on the high-precision data in the register to obtain a second intermediate calculation result using the first data type includes: performing high-precision calculation in the register using the high-precision data of the third data type to obtain a high-precision calculation result of the third data type; The high-precision calculation result is converted from the third data type to the first data type to obtain a second intermediate calculation result using the first data type.
9. The method according to claim 6, characterized in that The re-describing and storing the second intermediate calculation result based on the cache to obtain a data processing result using the second data type includes: Using a storage method corresponding to the first data type, storing the second intermediate calculation result from the register to the cache; re-describe the second intermediate calculation result as a second calculation result using the second data type; The second calculation result is stored from the cache to the memory using a storage method corresponding to the second data type, thereby obtaining a data processing result using the second data type.
10. A data processing device, characterized in that: The device comprises: a data block description module, configured to, when it is determined based on the memory access granularity, data transmission mode, and data arrangement mode of the source tensor data that an address range corresponding to the low-precision data block is discontinuous in the memory, describe the low-precision data block of a first data type as a high-precision data block of a second data type; wherein the second data type has a higher precision than the first data type; a data processing operation module, configured to perform data processing-related operations based on the low-precision data block in a manner corresponding to the second data type, to obtain a data processing result using the second data type; wherein the data processing-related operations include at least one of data movement and data calculation; A result redescription module is configured to redescribe the data processing result as low-precision result data using the first data type.
11. A chip, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the data processing method according to any one of claims 1 to 9 when executing the instructions stored in the memory.
12. A computer device, characterized in that: The method comprises a memory, an artificial intelligence chip, and a computer program stored in the memory and executable on the artificial intelligence chip, wherein the artificial intelligence chip implements the data processing method according to any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Data processing method, device and equipment
CN115470235A
Data processing method, computing device and related product
CN116185274A
Improving accuracy of machine learning operations by compensating for lower precision with scaled transforms
CN119790409A
Inference acceleration method and device, electronic equipment and storage medium
CN120450040A
Computer processor for higher precision computations using a mixed-precision decomposition of operations
US20190042244A1
Cited By
Data processing method and device, computing system, storage medium, program product and computer equipment
CN121210149A