Data format conversion device and method, electronic equipment and storage medium
By setting up a data format conversion device inside the GPU to convert high-precision data into low-precision data, the problem of increased computing and storage requirements in AI models is solved, and data processing efficiency and storage efficiency are improved.
Patent Information
- Application Number
- CN202511212015.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-09-26
AI Technical Summary
As the scale of AI models expands, the use of FP32 data format leads to a significant increase in computing power and storage capacity. There is an urgent need for a solution to convert high-precision data into low-precision data to reduce computing and storage requirements.
A data format conversion device is set up inside the GPU, including a data input module, a scaling factor determination module and a format conversion module. By determining the target scaling factor, high-precision source data is converted into low-precision target data and sent to the matrix operation module for processing.
It effectively improves the data processing efficiency of the GPU, reduces the data storage and transmission bandwidth requirements, and improves the processing efficiency of the matrix operation module.
Smart Images

Figure CN120704739A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a data format conversion device and method, an electronic device, and a storage medium. Background Art
[0002] Traditional machine learning models use the FP32 data format for computation and storage. However, as the model scales, the required computing power and storage capacity also increase significantly. Currently, to reduce the computational power and storage capacity of models, lower-precision data formats such as FP16, BF16, FP8, and INT8 are being widely used. Hardware for artificial intelligence (AI) models, such as graphics processing units (GPUs), now widely support lower-precision data formats. Therefore, a solution for converting high-precision data to lower-precision data is urgently needed. Summary of the Invention
[0003] The present disclosure provides a data format conversion device and method, an electronic device, and a technical solution for a storage medium.
[0004] According to one aspect of the present disclosure, a data format conversion device is provided, comprising: the data format conversion device is located in a graphics processing unit (GPU) and is connected to a data memory and a matrix operation module in the GPU; the data format conversion device comprises: a data input module, a scaling factor determination module, a format conversion module, and a data output module; the data input module is used to read a source data group from the data memory and send the source data group to the scaling factor determination module and the format conversion module, wherein the source data group includes N source data in a first floating-point data format, where N is a positive integer greater than 1; the scaling factor determination module is used to determine a target scaling factor shared by the N source data and send the target scaling factor to the format conversion module; the format conversion module is used to scale each source data according to the target scaling factor to obtain a target scaling factor. The target data group is prepared and the target data group is sent to the data output module, wherein the target data group includes the target scaling factor and N target data in a second floating-point data format, and the precision of the second floating-point data format is lower than the precision of the first floating-point data format; the data output module is used to send the target data group to the matrix operation module; the scaling factor determination module includes: a data selector and a scaling factor determination submodule; the data selector is used to compare the exponential data of the N source data, select the target source data, and send the target source data to the scaling factor determination submodule, wherein the target source data is the source data with the largest exponential data among the N source data; the scaling factor determination submodule is used to determine the target scaling factor according to the target source data, and send the target scaling factor to the format conversion module.
[0005] In one possible implementation, the data format of the target scaling factor is a W-bit binary number; the scaling factor determination submodule is specifically configured to: determine that the decimal value of the target scaling factor is 2 when all the exponent data of the target source data are 1. W -1.
[0006] In a possible implementation, the data format of the target scaling factor is a W-bit binary number; the scaling factor determination submodule is specifically configured to: determine an initial scaling factor according to the second floating-point data format and the target source data when the exponent data of the target source data are not all 1; and determine the initial scaling factor in the decimal value between -(2 W-1 -1) to 2 W-1 -1, the decimal value of the target scaling factor is determined to be the sum of the decimal value of the initial scaling factor and the target offset, wherein the target offset is 2 W-1-1; the decimal value of the initial scaling factor is greater than 2 W-1 In the case of -1, the decimal value of the target scaling factor is determined to be 2 W -1 -1 and the sum of the target offset; when the decimal value of the initial scaling factor is less than -(2 W-1 -1), the decimal value of the target scaling factor is determined to be -(2 W-1 -1) and the target offset.
[0007] In one possible implementation, the scaling factor determination submodule is specifically used to: determine a first scaling parameter based on the target source data; determine a second scaling parameter based on the maximum exponent data corresponding to the second floating-point data format; and subtract the second scaling parameter from the first scaling parameter to obtain the initial scaling factor.
[0008] In a possible implementation, the value of N is determined according to a data group size parameter corresponding to the target data group.
[0009] In one possible implementation, the data memory is a global data memory in the GPU and / or a local data memory in the GPU; the matrix operation module is an arithmetic logic unit ALU array in the GPU.
[0010] In a possible implementation, the matrix operation module is configured to perform data processing on the target data group to obtain a processed data group, and send the processed data group to the data storage for storage.
[0011] In one possible implementation, the source data group and the processed data group are image data involved in an artificial intelligence (AI) model.
[0012] According to one aspect of the present disclosure, a data format conversion method is provided, which is applied to a data format conversion device, wherein the data format conversion device is located in a GPU and is connected to a data memory and a matrix operation module in the GPU; the data format conversion device comprises: a data input module, a scaling factor determination module, a format conversion module, and a data output module; the method comprises: using the data input module to read a source data group from the data memory, and sending the source data group to the scaling factor determination module and the format conversion module, wherein the source data group comprises N source data in a first floating-point data format, where N is greater than 1. A positive integer; using the scaling factor determination module to determine a target scaling factor shared by the N source data, and sending the target scaling factor to the format conversion module; using the format conversion module to scale each source data according to the target scaling factor to obtain a target data group, and sending the target data group to the data output module, wherein the target data group includes the target scaling factor and N target data in a second floating-point data format, and the precision of the second floating-point data format is lower than the precision of the first floating-point data format; using the data output module, sending the target data group to the matrix operation module.
[0013] According to one aspect of the present disclosure, an electronic device is provided, including: the above-mentioned data format conversion device.
[0014] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the above method.
[0015] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above method is implemented.
[0016] In an embodiment of the present disclosure, a data format conversion device is provided inside a GPU, and the data format conversion device is connected to a data memory and a matrix operation module in the GPU. The data format conversion device includes: a data input module, a scaling factor determination module, a format conversion module, and a data output module; the data input module reads a source data group from the data memory, and sends the source data group to the scaling factor determination module and the format conversion module, wherein the source data group includes N source data in a first floating-point data format, where N is a positive integer greater than 1; the scaling factor determination module determines a target scaling factor shared by the N source data, and sends the target scaling factor to the format conversion module; the format conversion module performs scaling processing on each source data according to the target scaling factor. Processing, obtaining a target data group, and sending the target data group to a data output module, the target data group includes a target scaling factor and N target data in a second floating-point data format, the precision of the second floating-point data format is lower than the precision of the first floating-point data format; the data output module sends the target data group to a matrix operation module; the scaling factor determination module includes: a data selector and a scaling factor determination submodule; the data selector compares the exponential data of the N source data, selects the source data with the largest exponential data among the N source data as the target source data, and sends the target source data to the scaling factor determination submodule; the scaling factor determination submodule determines the target scaling factor according to the target source data, and sends the target scaling factor to the format conversion module. Based on the data format conversion device, the high-precision source data group read from the data storage device is effectively converted into a low-precision target data group and then sent to the matrix operation module for data processing, so as to improve the data processing efficiency of the subsequent matrix operation module.
[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, rather than limiting the present disclosure. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0019] Figure 1 A block diagram of a data format conversion device according to an embodiment of the present disclosure is shown.
[0020] Figure 2 A schematic diagram illustrating the MX data format according to an embodiment of the present disclosure.
[0021] Figure 3 A schematic diagram of a data format conversion device according to an embodiment of the present disclosure is shown.
[0022] Figure 4 A flow chart of a data format conversion method according to an embodiment of the present disclosure is shown.
[0023] Figure 5 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0024] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0025] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0026] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of three situations: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.
[0027] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.
[0028] Currently, AI models use the FP32 data format for computation. However, as models scale, this data format poses challenges to GPU computing power, bandwidth requirements, and data storage capabilities. To balance computing power, data storage, and ultimate model accuracy, AI models are increasingly using data formats such as FP16, BF16, FP8, and INT8. To further reduce bandwidth requirements and data storage, data can be further optimized.
[0029] The present disclosure provides a data format conversion device that effectively converts a high-precision source data set read from a GPU's data memory into a low-precision target data set, which is then sent to a matrix operation module in the GPU for data processing, thereby improving the GPU's data processing efficiency. The data format conversion device provided by the present disclosure is described in detail below.
[0030] Figure 1 FIG. 1 is a block diagram of a data format conversion device according to an embodiment of the present disclosure. Figure 1 As shown, the data format conversion device is located in the GPU and is connected to the data storage and matrix operation module in the GPU; the data format conversion device includes: a data input module, a scaling factor determination module, a format conversion module, and a data output module.
[0031] The data input module is used to read the source data group from the data storage and send the source data group to the scaling factor determination module and the format conversion module, wherein the source data group includes N source data in the first floating point data format, and N is a positive integer greater than 1; the scaling factor determination module is used to determine the target scaling factor shared by the N source data and send the target scaling factor to the format conversion module; the format conversion module is used to scale each source data according to the target scaling factor to obtain a target data group, and send the target data group to the data output module, wherein the target data group includes the target scaling factor, and N target data in a second floating-point data format, the precision of the second floating-point data format is lower than the precision of the first floating-point data format; a data output module, used to send the target data group to the matrix operation module; the scaling factor determination module includes: a data selector, and a scaling factor determination submodule; the data selector compares the exponential data of the N source data, selects the source data with the largest exponential data among the N source data as the target source data, and sends the target source data to the scaling factor determination submodule; the scaling factor determination submodule determines the target scaling factor according to the target source data, and sends the target scaling factor to the format conversion module.
[0032] In an embodiment of the present disclosure, a data format conversion device is provided inside a GPU, and the data format conversion device is connected to a data memory and a matrix operation module in the GPU. The data format conversion device includes: a data input module, a scaling factor determination module, a format conversion module, and a data output module; the data input module reads a source data group from the data memory, and sends the source data group to the scaling factor determination module and the format conversion module, wherein the source data group includes N source data in a first floating-point data format, where N is a positive integer greater than 1; the scaling factor determination module determines a target scaling factor shared by the N source data, and sends the target scaling factor to the format conversion module; the format conversion module performs scaling processing on each source data according to the target scaling factor. Processing, obtaining a target data group, and sending the target data group to a data output module, the target data group includes a target scaling factor and N target data in a second floating-point data format, the precision of the second floating-point data format is lower than the precision of the first floating-point data format; the data output module sends the target data group to a matrix operation module; the scaling factor determination module includes: a data selector and a scaling factor determination submodule; the data selector compares the exponential data of the N source data, selects the source data with the largest exponential data among the N source data as the target source data, and sends the target source data to the scaling factor determination submodule; the scaling factor determination submodule determines the target scaling factor according to the target source data, and sends the target scaling factor to the format conversion module. Based on the data format conversion device, the high-precision source data group read from the data storage device is effectively converted into a low-precision target data group and then sent to the matrix operation module for data processing, so as to improve the data processing efficiency of the subsequent matrix operation module.
[0033] The floating-point data format is used in computer science to represent real numbers. It allows for the representation of very large or very small values, as well as values in between, within a limited storage space. A floating-point number typically consists of three parts: a sign bit, an exponent bit, and a mantissa bit. For example, a single-precision floating-point number (FP32 data format) uses 32 bits (32 bits, 4 bytes) to store a floating-point number, including 1 sign bit, 8 exponent bits, and 23 mantissa bits.
[0034] The MX data format is a microscaling specification designed to enable AI model training and inference by reducing computational bit width and improving hardware performance and efficiency. The MX data format enables computation with lower bit widths and a smaller memory footprint, reducing additional fees and operating costs.
[0035] The MX data format consists of two parts: a shared scale factor and k data items of the same data format (lower precision). The scale factor is shared by the k data items of the same data format and is formatted as an 8-bit unsigned integer (E8M0). k represents the data group size parameter corresponding to the MX data format. The specific value of k can be adjusted flexibly based on the actual application scenario, for example, k = 32.
[0036] The data formats supported by the MX data format are generally data formats with less than 8 bits, such as FP8_E5M2, FP8_E4M3, FP6_E2M3, FP6_E3M2, FP4, INT8, etc.
[0037] Figure 2 Schematic diagram showing the MX data format according to an embodiment of the present disclosure. Figure 2 As shown, the MX data format includes: scaling factor X, k data of the same data type: P1 to P k .
[0038] In the disclosed embodiments, the first and second floating-point data formats are pre-set in the upper software layer based on actual application scenarios. For example, the first floating-point data format is a higher-precision floating-point data format commonly used in actual AI models, such as FP32 or FP16; the second floating-point data format is a lower-precision floating-point data format, such as one of FP8_E5M2, FP8_E4M3, FP6_E2M3, FP6_E3M2, and FP4.
[0039] In the disclosed embodiment, the data group size parameter k corresponding to the MX data format is also pre-set in the upper software layer based on the actual application scenario. For example, considering the actual processing capability of the GPU, which can process 32 data at a time, the data group size parameter k corresponding to the MX data format is set to 32. Based on the pre-set second floating-point data format and the data group size parameter k corresponding to the MX data format, the specific form of the target MX data format can be determined.
[0040] In a possible implementation, the value of N is determined according to a data group size parameter corresponding to the target data group.
[0041] Since the target data set is in the target MX data format, the size of the source data set that the data conversion device reads from the GPU's data memory in a single pass can be determined based on the data set size parameter k corresponding to the target MX data format. In other words, k=N.
[0042] For example, the data group size parameter k corresponding to the target MX data format is 32, so the data input module reads the source data group from the data storage, and the source data group includes N=32 source data in the first floating-point data format.
[0043] Figure 3 FIG. 1 is a schematic diagram showing a data format conversion device according to an embodiment of the present disclosure. Figure 3 As shown, the data format conversion device includes a data input module, a scaling factor determination module, and a format conversion module. The data input module reads a source data set from the GPU's data memory. The source data set includes N=32 source data (source data 0 to source data 31) in a first floating-point data format. The data input module then sends the source data set (source data 0 to source data 31) to the scaling factor determination module and the format conversion module, respectively.
[0044] The data selector in the scaling factor determination module can be a comparison module that compares the exponential data of N source data, selects the target source data with the largest exponential data, and sends the target source data to the scaling factor determination submodule in the scaling factor determination module for subsequent processing.
[0045] like Figure 3 As shown, the scaling factor determination module includes: a data selector and a scaling factor determination submodule. The data selector selects target source data and sends it to the scaling factor determination submodule.
[0046] In one possible implementation, the data format of the target scaling factor is a W-bit binary number; the scaling factor determination submodule is specifically configured to: when the exponential data of the target source data are all 1, determine that the decimal value of the target scaling factor is 2 W -1.
[0047] After the scaling factor determination submodule receives the target source data sent by the data selector, the scaling factor determination submodule first determines whether the exponential data of the target source data are all 1, that is, whether the target source data is a special value NaN or INF.
[0048] The data format of the target scaling factor in the target MX data format is a W-bit unsigned binary integer, that is, the decimal value range of the target scaling factor is 0 to 2 W -1. When the scaling factor determination submodule determines that the exponential data of the target source data are all 1, that is, when the target source data is a special value NaN or INF, the scaling factor determination submodule directly determines the decimal value of the target scaling factor to be 2. W -1.
[0049] When the target MX data format is the standard MX data format, the data format of the target scaling factor is an 8-bit unsigned binary integer, that is, the decimal value range of the target scaling factor is 0 to 255. When the scaling factor determination submodule determines that the exponential data of the target source data are all 1s, that is, the target source data is the special value NaN or INF, the scaling factor determination submodule directly determines that the decimal value of the target scaling factor is 255, that is, the target scaling factor is 11111111.
[0050] In a possible implementation, the data format of the target scaling factor is a W-bit binary number; the scaling factor determination submodule is specifically configured to: determine the initial scaling factor according to the second floating-point data format and the target source data when the exponent data of the target source data are not all 1; and determine the initial scaling factor when the decimal value of the initial scaling factor is between -(2 W-1 -1) to 2 W-1 -1, the decimal value of the target scaling factor is determined to be the sum of the decimal value of the initial scaling factor and the target offset, where the target offset is 2 W-1 -1; the decimal value of the initial scaling factor is greater than 2 W-1 In the case of -1, the decimal value of the target scaling factor is determined to be 2 W-1 The sum of -1 and the target offset; the decimal value of the initial scaling factor is less than -(2 W-1 -1), the decimal value of the target scaling factor is determined to be -(2 W-1 -1) and the target offset.
[0051] The scaling factor determination submodule determines an initial scaling factor according to the second floating-point data format and the target source data when it is determined that the exponent data of the target source data are not all 1, that is, the target source data is not a special value NaN or INF.
[0052] In one possible implementation, the scaling factor determination submodule is specifically used to: determine a first scaling parameter based on target source data; determine a second scaling parameter based on maximum exponent data corresponding to the second floating-point data format; and subtract the second scaling parameter from the first scaling parameter to obtain an initial scaling factor.
[0053] According to the target data source Value max , the first scaling parameter x1 is determined using the following formula (1).
[0054] x1=floor(log2(Value max )) (1).
[0055] The second scaling parameter x2 is determined based on the maximum exponent data corresponding to the second floating-point data format. For example, when the second floating-point data format is FP8_E5M2, the maximum exponent data corresponding to the second floating-point data format FP8_E5M2 is 14. Therefore, the second scaling parameter x2 is determined to be 14.
[0056] After determining the first scaling parameter x1 and the second scaling parameter x2, the initial scaling factor X is determined using the following formula (2).
[0057] X= x1- x2 (2).
[0058] Since the data format of the target scaling factor in the target MX data format is a W-bit unsigned binary integer, that is, the decimal value range of the target scaling factor is 0 to 2 W -1. At this time, the decimal value range of the initial scaling factor is determined to be -(2 W-1 -1) to 2 W-1 -1, then set the target offset to 2 W-1 -1, so that the sum of the initial scaling factor and the target offset is the target scaling factor, and the value range can meet the data format of the target scaling factor in the target MX data format.
[0059] The decimal value of the initial scaling factor is -(2 W-1 -1) to 2 W-1 -1, the decimal value of the target scaling factor is determined to be the sum of the decimal value of the initial scaling factor and the target offset.
[0060] The decimal value of the initial scaling factor is greater than 2 W-1 In the case of -1, the decimal value of the target scaling factor is determined to be 2 W-1 The sum of -1 and the target offset, that is, the decimal value of the target scaling factor is 2 W -1.
[0061] When the decimal value of the initial scaling factor is less than -(2 W-1 -1), the decimal value of the target scaling factor is determined to be -(2 W-1 -1) and the target offset, that is, the decimal value of the target scaling factor is 0.
[0062] When the target MX data format is a standard MX data format, the data format of the target scaling factor is an 8-bit unsigned integer binary number; when the target MX data format is a non-standard MX data format, the data format of the target scaling factor can be a 4-bit unsigned integer binary number, or a 16-bit unsigned integer binary number, etc., and the present disclosure does not make specific limitations on this.
[0063] When the target MX data format is the standard MX data format, the data format of the target scaling factor is an 8-bit unsigned binary integer, that is, the decimal value range of the target scaling factor is 0 to 255. In this case, the decimal value range of the initial scaling factor is determined to be -127 to 127, and the target offset is set to 127, so that the sum of the initial scaling factor and the target offset is the target scaling factor, and the value range can meet the data format of the target scaling factor in the target MX data format.
[0064] When the decimal value of the initial scaling factor is between -127 and 127, the decimal value of the target scaling factor is determined to be the sum of the decimal value of the initial scaling factor and the target offset.
[0065] When the decimal value of the initial scaling factor is greater than 127, the decimal value of the target scaling factor is determined to be the sum of 127 and the target offset, that is, the decimal value of the target scaling factor is 255.
[0066] When the decimal value of the initial scaling factor is less than -127, the decimal value of the target scaling factor is determined to be the sum of -127 and the target offset, that is, the decimal value of the target scaling factor is 0.
[0067] After the scaling factor determination submodule determines the target scaling factor, it sends the target scaling factor to the format conversion module. Figure 3 As shown, the scaling factor determination submodule sends the target scaling factor to the format conversion module.
[0068] After receiving the target scaling factor, the format conversion module performs scaling processing on each source data in the first floating point data format according to the target scaling factor to obtain target data in the second floating point data format corresponding to each source data.
[0069] For example, for the source data V in the first floating point data format i , according to the target scaling factor X, the source data V i Perform scaling processing to obtain the source data V i The corresponding target data V in the second floating point data format i '=V i / X.
[0070] like Figure 3 As shown, the format conversion module performs scaling processing on source data 0 to source data 31 in the first floating point data format based on the target scaling factor to obtain target data 0 to target data 31 in the second floating point data format.
[0071] The format conversion module constructs a target data group in the target MX data format according to the target scaling factor and the N target data in the second floating point data format, and then sends the target data group to the matrix operation module. Figure 3 As shown, the format conversion module sends the target data group (target scaling factor, target data 0 to target data 31) to the matrix operation module.
[0072] In one possible implementation, the data memory is a global data memory in the GPU and / or a local data memory in the GPU; and the matrix operation module is an arithmetic and logic unit (ALU) array in the GPU.
[0073] The data format conversion device of the embodiment of the present disclosure can be applied to the data format conversion during the data processing of the ALU array on the higher-precision data stored in the global data memory; it can also be applied to the data format conversion during the data processing of the ALU array on the higher-precision data stored in the local data memory.
[0074] In a possible implementation, the matrix operation module is used to perform data processing on the target data group to obtain a processed data group, and send the processed data group to a data memory for storage.
[0075] After receiving the target data set, the matrix operation module can process the target data set according to the corresponding data processing rules to obtain a processed data set. The data processing rules corresponding to the matrix operation module can be flexibly set according to the actual application scenario, and this disclosure does not make specific limitations on this.
[0076] After the matrix operation module obtains the processed data group, the processed data group can be sent to the data memory for storage. Since the processed data group includes N data in the second floating point data format with lower precision, the storage capacity occupied by the data memory can be reduced, and the bandwidth occupied during the data transmission process can also be reduced. Figure 3 As shown, the matrix operation module sends the processed data group to the data memory for storage.
[0077] In one possible implementation, the source data group and the processed data group are image data involved in the AI model.
[0078] The GPU where the data format conversion device of the embodiment of the present disclosure is located can be a GPU used under an AI model, and the AI model can be a model for performing data processing on images. Therefore, the source data group and processed data group corresponding to the data format conversion device can both be image data involved in the AI model.
[0079] In an embodiment of the present disclosure, a data format conversion device is provided inside a GPU, and the data format conversion device is provided to be connected to a data storage device and a matrix operation module in the GPU, the data format conversion device including: a data input module, a scaling factor determination module, a format conversion module, and a data output module; the data input module reads a source data group from the data storage device, and sends the source data group to the scaling factor determination module and the format conversion module, the source data group including N source data in a first floating-point data format, N being a positive integer greater than 1; the scaling factor determination module determines a target scaling factor shared by the N source data, and sends the target scaling factor to the format conversion module; the format conversion module scales each source data according to the target scaling factor to obtain a target data group, and sends the target data group to the data output module, the target data group including the target scaling factor and N target data in a second floating-point data format, the precision of the second floating-point data format being lower than the precision of the first floating-point data format; the data output module sends the target data group to the matrix operation module. Based on the data format conversion device, the high-precision source data group read from the data storage is effectively converted into a low-precision target data group and then sent to the matrix operation module for data processing, so as to improve the data processing efficiency of the subsequent matrix operation module.
[0080] Figure 4 A flow chart of a data format conversion method according to an embodiment of the present disclosure is shown. Figure 1 or Figure 3 The data format conversion device shown is located in the GPU and is connected to the data storage and matrix operation module in the GPU. The data format conversion device includes: a data input module, a scaling factor determination module, a format conversion module, and a data output module. Figure 4 As shown, the method includes:
[0081] In step S41, a data input module is used to read a source data group from a data storage device, and the source data group is sent to a scaling factor determination module and a format conversion module, wherein the source data group includes N source data in a first floating-point data format, where N is a positive integer greater than 1.
[0082] In step S42 , a scaling factor determination module is used to determine a target scaling factor shared by the N source data, and the target scaling factor is sent to a format conversion module.
[0083] In step S43, the format conversion module is used to scale each source data according to the target scaling factor to obtain a target data group, and the target data group is sent to the data output module, wherein the target data group includes the target scaling factor and N target data in the second floating-point data format, and the precision of the second floating-point data format is lower than the precision of the first floating-point data format.
[0084] In step S44, the target data group is sent to the matrix operation module using the data output module.
[0085] The scaling factor determination module includes: a data selector and a scaling factor determination submodule;
[0086] Step S42 specifically includes:
[0087] In step S421, a data selector is used to compare the exponential data of the N source data, select the target source data, and send the target source data to the scaling factor determination submodule, wherein the target source data is the source data with the largest exponential data among the N source data;
[0088] In step S422 , the scaling factor determination submodule is used to determine a target scaling factor according to the target source data, and the target scaling factor is sent to the format conversion module.
[0089] Apply the above Figure 1 or Figure 3 The specific process of the data format conversion device shown in the figure performing data format conversion can be referred to the above Figures 1 to 3 The detailed description of the related embodiments is omitted here.
[0090] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0091] This method has a specific technical connection with the internal structure of the computer system, and can solve the technical problem of how to improve the hardware computing efficiency or execution effect (including reducing the amount of data storage, reducing the amount of data transmission, increasing the hardware processing speed, etc.), thereby obtaining the technical effect of improving the internal performance of the computer system in accordance with the laws of nature.
[0092] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0093] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.
[0094] The embodiment of the present disclosure further provides an electronic device, including: the above-mentioned data format conversion device.
[0095] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above method.
[0096] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0097] The electronic device may be provided as a terminal, a server, or other forms of devices.
[0098] Figure 5 FIG. 1 is a block diagram of an electronic device according to an embodiment of the present disclosure. Figure 5 , the electronic device 1900 can be provided as a server or a terminal device. Figure 5 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.
[0099] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as a Microsoft Server operating system (Windows Server 2003). TM ), Apple's graphical user interface operating system (Mac OS X TM ), a multi-user, multi-process computer operating system (Unix TM ), a free and open source Unix-like operating system (Linux TM ), an open-source Unix-like operating system (FreeBSD TM ) or similar.
[0100] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.
[0101] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0102] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, (but not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure within a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0103] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0104] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, the state information of the computer-readable program instructions is used to personalize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), so that the electronic circuit can execute the computer-readable program instructions, thereby implementing various aspects of the present disclosure.
[0105] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0106] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0107] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0108] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0109] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).
[0110] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.
[0111] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0112] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0113] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A data format conversion device, characterized in that: The data format conversion device is located in the graphics processing unit (GPU) and is connected to the data storage and matrix operation module in the GPU. The data format conversion device includes: a data input module, a scaling factor determination module, a format conversion module, and a data output module. The data input module is configured to read a source data group from the data storage and send the source data group to the scaling factor determination module and the format conversion module, wherein the source data group includes N source data in a first floating-point data format, where N is a positive integer greater than 1; The scaling factor determination module is configured to determine a target scaling factor shared by the N source data, and send the target scaling factor to the format conversion module; The format conversion module is configured to perform scaling processing on each source data according to the target scaling factor to obtain a target data group, and send the target data group to the data output module, wherein the target data group includes the target scaling factor and N target data in a second floating-point data format, where the precision of the second floating-point data format is lower than the precision of the first floating-point data format; The data output module is used to send the target data group to the matrix operation module; The scaling factor determination module includes: a data selector, a scaling factor determination submodule; The data selector is configured to compare the exponential data of the N source data, select target source data, and send the target source data to the scaling factor determination submodule, wherein the target source data is the source data with the largest exponential data among the N source data; The scaling factor determination submodule is configured to determine the target scaling factor according to the target source data, and send the target scaling factor to the format conversion module.
2. The data format conversion device according to claim 1, wherein: The data format of the target scaling factor is a W-bit binary number; The scaling factor determination submodule is specifically configured to: When all the exponent data of the target source data are 1, the decimal value of the target scaling factor is determined to be 2. W -1.
3. The data format conversion device according to claim 1, wherein: The data format of the target scaling factor is a W-bit binary number; The scaling factor determination submodule is specifically configured to: When the exponent data of the target source data are not all 1, determining an initial scaling factor according to the second floating-point data format and the target source data; The decimal value of the initial scaling factor is -(2 W-1 -1) to 2 W-1 -1, the decimal value of the target scaling factor is determined to be the sum of the decimal value of the initial scaling factor and the target offset, wherein the target offset is 2 W-1 -1; The decimal value of the initial scaling factor is greater than 2 W-1 -1, the decimal value of the target scaling factor is determined to be 2 W-1 The sum of -1 and the target offset; The decimal value of the initial scaling factor is less than -(2 W-1 -1), the decimal value of the target scaling factor is determined to be -(2 W-1 -1) and the target offset.
4. The data format conversion device according to claim 3, characterized in that: The scaling factor determination submodule is specifically configured to: determining a first scaling parameter according to the target source data; determining a second scaling parameter according to maximum exponent data corresponding to the second floating-point data format; The initial scaling factor is obtained by subtracting the second scaling parameter from the first scaling parameter.
5. The data format conversion device according to claim 1, wherein: The value of N is determined according to the data group size parameter corresponding to the target data group.
6. The data format conversion device according to claim 1, wherein: The data memory is a global data memory in the GPU and / or a local data memory in the GPU; The matrix operation module is the arithmetic logic unit ALU array in the GPU.
7. The data format conversion device according to claim 1, wherein: The matrix operation module is used to perform data processing on the target data group to obtain a processed data group, and send the processed data group to the data storage for storage.
8. The data format conversion device according to claim 7, characterized in that: The source data group and the processed data group are image data involved in the artificial intelligence AI model.
9. A data format conversion method, characterized in that: The method is applied to a data format conversion device, which is located in a GPU and connected to a data memory and a matrix operation module in the GPU; The data format conversion device includes: a data input module, a scaling factor determination module, a format conversion module, and a data output module; the method includes: Using the data input module, reading a source data group from the data storage, and sending the source data group to the scaling factor determination module and the format conversion module, wherein the source data group includes N source data in a first floating-point data format, where N is a positive integer greater than 1; Determining a target scaling factor shared by the N source data using the scaling factor determination module, and sending the target scaling factor to the format conversion module; Utilizing the format conversion module, scaling each source data according to the target scaling factor to obtain a target data group, and sending the target data group to the data output module, wherein the target data group includes the target scaling factor and N target data in a second floating-point data format, where the precision of the second floating-point data format is lower than the precision of the first floating-point data format; Using the data output module, sending the target data group to the matrix operation module; The scaling factor determination module includes: a data selector, a scaling factor determination submodule; The utilizing the scaling factor determination module to determine a target scaling factor shared by the N source data, and sending the target scaling factor to the format conversion module, comprises: Using the data selector, comparing the exponential data of the N source data, selecting target source data, and sending the target source data to the scaling factor determination submodule, wherein the target source data is the source data with the largest exponential data among the N source data; The scaling factor determination submodule is used to determine the target scaling factor according to the target source data, and the target scaling factor is sent to the format conversion module.
10. An electronic device, characterized in that: include: The data format conversion device according to any one of claims 1 to 8.
11. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method of claim 9.
12. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to claim 9 is implemented.
Citation Information
Patent Citations
Vector floating point argument reduction
CN102566964A
Data format conversion method and device and matrix processing method and device
CN115237991A
Multiplying and adding operation hardware with mixed precision
CN119645346A
Data processing method, device and computer program product
CN120528432A
Vector scaling instructions for use in an arithmetic logic unit
US20160019027A1