Floating point number processing method, apparatus, computing device and storage medium
By introducing a bit width indication field into floating-point numbers, dynamically adjusting the bit width of the order code and mantissa domain, the problem of insufficient numerical range in AI training and inference is solved, and the numerical range and accuracy are improved without increasing the total bit width and reducing the storage transfer overhead.
Patent Information
- Application Number
- PCT/CN2024/141099
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-03
AI Technical Summary
In AI training and inference, existing floating-point number representations cannot meet the numerical range requirements due to the small bit width of the order code domain, affect performance, and increase data storage and transfer overhead.
By introducing a bit width indication field into floating point numbers, the bit widths of the order code domain and mantissa domain are dynamically adjusted, the order code encoding space is increased, the numerical range is expanded, and the accuracy is maintained to avoid the increase of the total bit width.
Without increasing the total bit width, it meets the numerical range and accuracy requirements of AI training and inference, reduces data storage and transfer overhead, and improves computing efficiency.
Smart Images

Figure CN2024141099_03072025_PF_FP_ABST
Abstract
Description
Floating point number processing method, device, computing device and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 25, 2023, with application number 202311811917.5 and application name "A floating-point processing method, device, computing device and storage medium", the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of computer technology, and in particular to a floating-point number processing method, apparatus, computing device, and storage medium. Background Art
[0004] In a computer system, floating point (FP) is an approximate numerical representation for real numbers, also known as floating-point data representation. Exemplarily, floating-point data representations may include FP8, FP16, and FP32, where FP8 represents an 8-bit floating-point number, FP16 represents a 16-bit floating-point number, and FP32 represents a 32-bit floating-point number. Typically, a floating-point number comprises three domains, namely, a sign domain, an exponent domain, and a mantissa domain. In each of the above floating-point data representations, the bit width of each domain is fixed. For example, FP16 includes a 1-bit sign domain, a 5-bit exponent domain, and a 10-bit mantissa domain; FP32 includes a 1-bit sign domain, an 8-bit exponent domain, and a 23-bit mantissa domain; FP8 may include two types, one of which includes a 1-bit sign domain, a 5-bit exponent domain, and a 2-bit mantissa domain; the other includes a 1-bit sign domain, a 4-bit exponent domain, and a 3-bit mantissa domain. Among them, the bit width of the exponent field determines the numerical range that the floating-point number can represent, and the bit width of the mantissa field determines the numerical precision that the floating-point number can represent.
[0005] With the rapid development of mixed-precision training and inference in artificial intelligence (AI), floating-point numbers are increasingly being used in mixed-precision training. As the size of AI network parameters increases dramatically, using smaller-bit-width floating-point numbers for AI training and inference can save data storage and transfer overhead.
[0006] However, the smaller the overall bit width of the floating-point number, the smaller the bit width of its exponent field, and the smaller the range of values that can be represented. This often cannot meet the needs of AI training and reasoning, affecting the performance of AI training and reasoning. Summary of the Invention
[0007] The embodiments of the present application provide a floating-point number processing method, apparatus, computing device, and storage medium, which can expand the dynamic range of numerical values represented by floating-point numbers to meet the needs of AI training and reasoning.
[0008] In a first aspect, an embodiment of the present application provides a floating-point number processing method, which can be executed by a computing device, or by a chip, chip system, or circuit in the computing device. The floating-point number processing method may include: the computing device obtains a first floating-point number, wherein the first floating-point number may include a first sign field, a bit width indication field, a first exponent field, and a first mantissa field, the first sign field is used to indicate the sign of the first floating-point number, the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field is also used to indicate: the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number. After obtaining the first floating-point number, the computing device may decode the first floating-point number to obtain the above-mentioned sign, exponent, and mantissa.
[0009] In an embodiment of the present application, in addition to the first sign field, the first exponent field and the first mantissa field, the first floating-point number may also include a bit width indication field, which is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the floating-point number. Under the premise that the total bit width is fixed, the bit width of the first exponent field and the bit width of the first mantissa field can change dynamically with the value indicated by the bit width indication field, thereby meeting the requirements for different numerical ranges and precisions of floating-point numbers in different scenarios. The bit width indication field can also indicate whether the first mantissa field represents the mantissa or the exponent offset, that is, the first mantissa field can be used to represent the mantissa or to indicate the exponent. By reusing the first mantissa field, it is possible to ensure that the precision of the floating-point number is less affected and the encoding space of the exponent can be increased, thereby expanding the dynamic range of the numerical value represented by the floating-point number, which can meet the needs of AI training and reasoning.
[0010] In an optional implementation, after obtaining the first floating-point number, the computing device may decode the first floating-point number to obtain a second floating-point number. The second floating-point number may include a second sign field, a second exponent field, and a second mantissa field, wherein the second sign field is used to indicate a sign, the second exponent field is used to indicate an exponent, and the second mantissa field is used to indicate a mantissa.
[0011] In an optional implementation, when the bit width D of the first exponent field of the first floating-point number is not 0, the mantissa may be the value indicated by the first mantissa field, and the exponent may be the value indicated by the first exponent field. When the bit width D of the first exponent field of the first floating-point number is 0, the first exponent field does not exist. In this case, if the first mantissa field represents the mantissa, the exponent may be a first preset exponent. If the first mantissa field represents an offset of the exponent, the exponent is a value obtained by correcting the second preset exponent by the offset, and the mantissa may be a preset mantissa.
[0012] In the above implementation, when the first exponent field does not exist, if the first mantissa field represents the mantissa, the exponent can use the first preset exponent. If the first mantissa field represents the offset of the exponent, the exponent is the value obtained by correcting the second preset exponent using the offset. The value obtained by correcting the second preset exponent using the offset belongs to a different numerical range from the value indicated by the first exponent field. For example, when the total bit width of the first floating-point number is 8, the numerical range to which the value indicated by the first exponent field belongs can be [-15, 15], and the numerical range to which the value obtained by correcting the second preset exponent using the offset belongs can be [-22, -16]. Therefore, the exponent numerical range that the first floating-point number can represent can reach [-22, 15]. Compared with the dynamic range of numerical values that can be represented by the 8-bit floating-point number representation method in the related art, the dynamic range of the numerical values represented by the floating-point number is expanded.
[0013] In one optional implementation, the bit width DW of the bit width indicator field is negatively correlated with the bit width D of the first exponent field. This prevents a jump in the bit width of the first mantissa field caused by a simultaneous increase in the bit widths of the bit width indicator field and the first exponent field when the total bit width of the first floating-point number is fixed. The bit width of the mantissa field determines the precision of floating-point data. Therefore, this implementation allows for smooth changes in the precision of the numerical value represented by the first floating-point number, preventing jumps in the precision of the numerical value represented by the first floating-point number.
[0014] In an optional implementation, when the bit width DW of the bit width indication field is a preset bit width and the value indicated by the bit width indication field is a preset value, the first mantissa field represents an offset. For example, in one embodiment, when the total bit width of the first floating-point number is 8 and the bit width DW of the bit width indication field is 4, if the value of the bit width indication field is "0000", it indicates that the first mantissa field represents an offset; or when the total bit width of the first floating-point number is 8 and the bit width DW of the bit width indication field is 4, if the value of the bit width indication field is "0001", it indicates that the first mantissa field represents an offset.
[0015] In an optional implementation, the computing device may read the first floating-point number from a memory, or obtain the first floating-point number via a communication network, and after decoding the first floating-point number to obtain a sign, an exponent, and a mantissa, the obtained sign, exponent, and mantissa may be used for calculation.
[0016] In the related art, when the total bit width of a floating-point number is determined, the bit widths of its exponent field and mantissa field are always fixed. When the calculation process requires greater precision or a larger numerical range, only a data format with a larger total bit width can be selected. However, the increase in the total bit width means that the exponent field bit width and the mantissa field bit width increase at the same time, which easily leads to the waste of the increased exponent field bit width when only a greater precision is required, and the waste of the increased mantissa field bit width when only a larger numerical range is required, thereby taking up unnecessary storage space and increasing the overhead of data storage and data transfer of the floating-point number. In an embodiment of the present application, the first floating-point number can be used for data storage or data transfer. Since the bit widths of the first exponent field and the first mantissa field of the first floating-point number can dynamically change with the numerical value indicated by the bit width indication field, it is possible to flexibly meet the different requirements for the numerical range and numerical precision of the floating-point number in various scenarios without additionally increasing the total bit width of the floating-point number, that is, without additionally increasing the cost of data storage or data transfer.
[0017] In a second aspect, an embodiment of the present application provides a floating-point number processing method, which can be executed by a computing device, or by a chip, chip system or circuit in the computing device. The floating-point number processing method may include: the computing device obtains a sign, an exponent and a mantissa, and obtains a first floating-point number based on the sign, the exponent and the mantissa. The first floating-point number includes a first sign field, a bit width indication field, a first exponent field and a first mantissa field, the first sign field is used to indicate the sign of the first floating-point number, and the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number. The bit width indication field is also used to indicate: the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number.
[0018] In an optional implementation, when the bit width D of the first exponent field is 0, the first exponent field does not exist, and if the first mantissa field represents a mantissa, the exponent is the first preset exponent.
[0019] In an optional implementation, when the bit width D of the first exponent field is 0, the first exponent field does not exist. If the first mantissa field represents the offset of the exponent, the mantissa is a preset mantissa, and the exponent is the value obtained by correcting the second preset exponent using the offset.
[0020] In an optional implementation, when the bit width D of the first exponent field is not 0, the first mantissa field represents the mantissa, and the exponent is the value indicated by the first exponent field.
[0021] In an optional implementation, the bit width DW of the bit width indication field is negatively correlated with the bit width D.
[0022] In an optional implementation, when the bit width DW of the bit width indication field is a preset bit width, and the value indicated by the bit width indication field is a preset value, the first mantissa field represents an offset of the exponent.
[0023] In a third aspect, an embodiment of the present application provides a floating-point number processing apparatus, which can be applied to a computing device. The floating-point number processing apparatus may include:
[0024] A floating-point number acquisition module is configured to acquire a first floating-point number; the first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field; the first sign field is configured to indicate a sign of the first floating-point number, the bit width indication field is configured to indicate a bit width D occupied by the first exponent field within a total bit width N of the first floating-point number; the bit width indication field is further configured to indicate whether the first mantissa field represents the mantissa of the first floating-point number, or an offset of the exponent of the first floating-point number;
[0025] The decoding module is used to decode the first floating-point number to obtain a sign, an exponent, and a mantissa.
[0026] In an optional implementation, the decoding module may be specifically configured to:
[0027] The first floating-point number is decoded to obtain a second floating-point number; the second floating-point number includes a second sign field, a second exponent field and a second mantissa field, the second sign field is used to indicate the sign, the second exponent field is used to indicate the exponent, and the second mantissa field is used to indicate the mantissa.
[0028] In an optional implementation, when the bit width D is 0, the first exponent field does not exist, and if the first mantissa field represents a mantissa, the exponent is the first preset exponent.
[0029] In an optional implementation, when the bit width D is 0, the first exponent field does not exist. If the first mantissa field represents an offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting the second preset exponent using the offset.
[0030] In an optional implementation, when the bit width D is not 0, the first mantissa field represents the mantissa, and the exponent is the value indicated by the first exponent field.
[0031] In an optional implementation, the bit width DW of the bit width indication field is negatively correlated with the bit width D.
[0032] In an optional implementation, when the bit width DW of the bit width indication field is a preset bit width, and the value indicated by the bit width indication field is a preset value, the first mantissa field represents an offset.
[0033] In an optional implementation, the floating point number acquisition module can be specifically used to:
[0034] The first floating point number is read from a memory, or acquired through a communication network.
[0035] In an optional implementation, the floating-point number processing device may further include a calculation module, which may be configured to perform calculations using a sign, an exponent, and a mantissa.
[0036] In a fourth aspect, an embodiment of the present application provides a floating-point number processing apparatus, which can be applied to a computing device. The floating-point number processing apparatus may include:
[0037] A data acquisition module, used to obtain the sign, exponent and mantissa;
[0038] The encoding module is used to obtain a first floating-point number based on a sign, an exponent and a mantissa; the first floating-point number includes a first sign field, a bit width indication field, a first exponent field and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field is further used to indicate: whether the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number.
[0039] In an optional implementation, when the bit width D is 0, the first exponent field does not exist, and if the first mantissa field represents a mantissa, the exponent is the first preset exponent.
[0040] In an optional implementation, when the bit width D is 0, the first exponent field does not exist. If the first mantissa field represents an offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting the second preset exponent using the offset.
[0041] In an optional implementation, when the bit width D is not 0, the first mantissa field represents the mantissa, and the exponent is the value indicated by the first exponent field.
[0042] In an optional implementation, the bit width DW of the bit width indication field is negatively correlated with the bit width D.
[0043] In an optional implementation, when the bit width DW of the bit width indication field is a preset bit width, and the value indicated by the bit width indication field is a preset value, the first mantissa field represents an offset.
[0044] In a fifth aspect, an embodiment of the present application provides a computing device comprising a processor and a memory; a computer program is stored in the memory; the processor is used to read the computer program stored in the memory and execute any one of the floating-point processing methods provided in the first or second aspect above.
[0045] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to enable a computer to execute any one of the floating-point number processing methods provided in the first or second aspect above.
[0046] In a seventh aspect, an embodiment of the present application provides a computer program product comprising computer-executable instructions, wherein the computer-executable instructions are used to enable a computer to execute any one of the floating-point number processing methods provided in the first or second aspect above.
[0047] The technical effects that can be achieved in any of the second to seventh aspects can refer to the description of the beneficial effects in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] FIG1 is a schematic diagram of an application scenario of an embodiment of the present application;
[0049] FIG2 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0050] FIG3 is a schematic diagram of a significant bit-order distribution of a floating-point number provided in an embodiment of the present application;
[0051] FIG4 is a flow chart of a floating-point number processing method provided in an embodiment of the present application;
[0052] FIG5 is a schematic diagram of a decoder provided by an embodiment of the present application decoding a first floating-point number;
[0053] FIG6 is a flow chart of another floating-point number processing method provided in an embodiment of the present application;
[0054] FIG7 is a schematic diagram of an encoder encoding a first floating-point number provided by an embodiment of the present application;
[0055] FIG8 is a flowchart of another floating-point number processing method provided in an embodiment of the present application;
[0056] FIG9 is a diagram showing the training effect of AI training using floating-point numbers using related technologies;
[0057] FIG10 is a diagram showing the training effect of AI training using floating-point numbers according to an embodiment of the present application;
[0058] FIG11 is a schematic structural diagram of a floating-point number processing device provided in an embodiment of the present application;
[0059] FIG12 is a schematic structural diagram of another floating-point number processing device provided in an embodiment of the present application;
[0060] FIG13 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.
[0062] Before introducing the specific solutions provided by the embodiments of the present application, some of the terms in the present application are explained to facilitate understanding by those skilled in the art, and the terms in the present application are not limited.
[0063] (1) Integer encoding: refers to encoding integers using fixed-length binary. For example, 3 bits are used to encode 0 to 7, 4 bits are used to encode 0 to 15, and so on.
[0064] (2) Prefix code encoding: This can also be called prefix encoding. If any code in a coding method is not a prefix (leftmost substring) of any other code, then the coding method can be called prefix encoding. For example, unequal-length codes: 1, 01, 001, 0101; or 00, 01, 10, 1100, 1101; and equal-length codes: 00, 01, 10, 11, etc., are all prefix code encodings. Prefix encoding can ensure that there is no ambiguity when decoding compressed files, ensuring correct decoding.
[0065] Prefix code encoding may include conventional prefix code encoding and unconventional prefix code encoding. Conventional prefix code encoding uses a shorter bit width to encode smaller data and a longer bit width to encode larger data. Unconventional prefix code encoding is the opposite, using a shorter bit width to encode larger data and a longer bit width to encode smaller data. In some embodiments of the present application, the Dot field in the floating point number can be encoded by unconventional prefix code encoding. Please refer to the description in the following embodiments for details.
[0066] In the embodiments of the present application, "multiple" refers to two or more. In view of this, in the embodiments of the present application, "multiple" can also be understood as "at least two". "At least one" can be understood as one or more, for example, one, two or more. For example, including at least one means including one, two or more, and does not limit which ones are included. For example, including at least one of A, B and C, then the included ones may be A, B, C, A and B, A and C, B and C, or A, B and C. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / ", unless otherwise specified, generally indicates that the previous and subsequent associated objects are in an "or" relationship.
[0067] Unless otherwise specified, ordinal numbers such as "first" and "second" in the embodiments of the present application are used to distinguish multiple objects and are not used to limit the order, timing, priority or importance of multiple objects.
[0068] Floating-point data representation is a scientific calculation method. In the Institute of Electrical and Electronics Engineers (IEEE) 754 binary floating-point standard, floating-point numbers can contain three fields: the sign field, the exponent field, and the mantissa field. The value of the exponent field represents the integer power exponent of a certain base, and the value of the mantissa field multiplied by the integer power exponent of a certain base can obtain a data; the sign field is used to indicate the positive or negative of the data. For example, taking the binary data 100.101 (i.e. 4.625 in decimal) as an example, this data can be represented as (-1) 0 ×2 2 ×1.00101, where the exponent 0 of the base -1 is the value of the sign field in the floating-point number, used to indicate the positive or negative sign of the data. The exponent 2 of the base 2 is the value of the exponent field in the floating-point number, indicating the number of decimal point shifts. For example, from 100.101 to 1.00101, the number of decimal point shifts is 2. 1.00101 is the mantissa. Because the number before the decimal point must be 1, the value actually stored in the mantissa field of the floating-point number can only include the data after the decimal point, that is, 00101, saving a binary bit to store more mantissa.
[0069] Exemplarily, commonly used floating-point data representations may include FP16 and FP32, where FP16 represents a 16-bit floating-point number and FP32 represents a 32-bit floating-point number. In each floating-point data representation, the bit width of each field is fixed. For example, FP16 includes a 1-bit sign field, a 5-bit exponent field, and a 10-bit mantissa field; FP32 includes a 1-bit sign field, an 8-bit exponent field, and a 23-bit mantissa field. Among them, the bit width of the exponent field determines the range of values that the floating-point number can represent, and the bit width of the mantissa field determines the numerical precision that the floating-point number can represent.
[0070] As the size of AI network parameters increases dramatically, using smaller bit-width floating-point numbers for AI training and inference can save data storage and transfer overhead. Consequently, FP8 floating-point numbers have emerged. There are two types of FP8 floating-point numbers, as shown in Table 1.
[0071] Table 1
[0072] As shown in Table 1, FP8 floating-point numbers can be of two types: E5M2 and E4M3. E5M2 floating-point numbers consist of a 1-bit sign field, a 5-bit exponent field, and a 2-bit mantissa field; E4M3 floating-point numbers consist of a 1-bit sign field, a 4-bit exponent field, and a 3-bit mantissa field. The width of the exponent field in a floating-point number determines the exponent range that the floating-point number can represent. E5M2 floating-point numbers can represent exponents in the range [-16, 15], while E4M3 floating-point numbers can represent exponents in a smaller range, [-9, 8].
[0073] Since the dynamic range of values that can be represented by the above two FP8 floating-point numbers is relatively small, they often cannot meet the needs of AI training and reasoning, affecting the performance of AI training and reasoning. Using traditional floating-point data representation, only data formats with larger total bit widths can be selected, such as FP16 floating-point numbers or FP32 floating-point numbers. This easily leads to waste of the order field bit width or mantissa field bit width of floating-point numbers with larger total bit widths, causing floating-point numbers to occupy unnecessary storage space and greatly increasing the overhead of floating-point data storage and data transfer.
[0074] Based on this, an embodiment of the present application provides a floating-point number processing method. The floating-point number provided by the embodiment of the present application, in addition to including the first symbol domain, can also include a bit width indication domain (i.e., Dot domain), a first exponent domain, and a first mantissa domain, wherein the Dot domain is used to indicate the bit width D occupied by the first exponent domain in the total bit width N of the floating-point number. When the bit width D is non-zero, the first exponent domain is used to indicate the exponent of the floating-point number, and the first mantissa domain is used to indicate the mantissa of the floating-point number. Under the same total bit width, the bit width of the first exponent domain and the bit width of the first mantissa domain can change dynamically with the numerical value of the Dot domain, meeting the requirements of different numerical ranges and precisions of floating-point numbers in different scenarios. When the bit width D is 0, the Dot domain can also be used to indicate: whether the first mantissa domain represents the mantissa of the floating-point number or the offset of the exponent of the floating-point number. By reusing the first mantissa domain, it is possible to ensure that the precision of the floating-point number is less affected, and the encoding space of the exponent can be increased, thereby expanding the dynamic range of the numerical value represented by the floating-point number. For example, taking a floating-point number with a total bit width N of 8 as an example, the dynamic range of the values represented by the floating-point numbers of the above-mentioned E4M3 is [-9, 8], the dynamic range of the values that can be represented by the floating-point numbers of E5M2 is [-16, 15], and the dynamic range of the values that can be represented by the floating-point numbers of HiF8_DML provided in the embodiment of the present application is [-22, 15], which greatly expands the dynamic range of the values represented by the floating-point numbers and can meet the needs of AI training and reasoning.
[0075] The floating-point number processing method provided in the embodiment of the present application can be applied to computing devices and is widely applicable to various industries such as scientific research, engineering, finance, aerospace, and medical care. These industries have a large amount of data that needs to be stored, calculated, and transmitted every day. Figure 1 exemplarily shows a schematic diagram of an application scenario provided in the embodiment of the present application. As shown in Figure 1, in this application scenario, computing devices 100, computing devices 200, and computing devices 300 can establish a communication connection through a network. Wherein, the network can be a wired network or a wireless network, for example, wireless fidelity (WiFi), Bluetooth, mobile network, etc. Computing devices 100, computing devices 200, and computing devices 300 can be any electronic device such as a computer, a server, a smart wearable device, a smart home, a tablet computer, a laptop computer, a vehicle-mounted terminal, a smart phone, etc. When the computing device is a server, it can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center, etc. A communication connection is established between computing devices 100, computing devices 200, and computing devices 300, and data storage, calculation, and transmission based on floating-point numbers can be performed.
[0076] It should be noted that in actual application scenarios, the system may include more than 3 computing devices or less than 3 computing devices, and this application does not limit this.
[0077] The internal structures of computing devices 100, computing devices 200, and computing devices 300 may be the same. Take computing device 100 as an example for illustration, as shown in FIG2 . The computing device 100 may include a processor 110 and a memory 120. For example, the processor 110 may be, but is not limited to, a central processing unit (CPU), a high performance computing (HPC) business acceleration chip, a graphics processing unit (GPU), or an embedded neural network processing unit (NPU) in the field of AI. The processor 110 may include a decoder 111, an encoder 112, and a computing unit 113. The number of computing units 113 may be one or more.
[0078] Exemplarily, the processor 110 may include some or all of the following computing units: a scalar computing unit, a vector computing unit, a matrix computing unit, or a tensor computing unit.
[0079] The following is an introduction to different computing units.
[0080] 1. Scalar calculation unit: A scalar, also known as a pure quantity, has only size but no direction. A circuit for scalar calculation is called a scalar calculation unit. Scalar calculation is mostly used for general-purpose calculations. In an embodiment of the present application, an arithmetic logic unit (ALU) based on the HiF8_DML data format can be embedded in the execution unit (EXU) part of a CPU multi-stage pipeline, or in the scalar calculation part of other processors with similar functions.
[0081] 2. Vector computing unit: A vector, also known as a vector, usually refers to a one-dimensional array with a length greater than 1. A computing unit specially designed for vector computing with a certain degree of parallelism is called a vector computing unit, such as a single instruction multiple data (SIMD) processor. Vector computing units are mostly used in fields such as HPC high-performance computing and AI machine learning, including solutions to mathematical problems such as linear programming, Fourier transform, filtering calculations, and linear algebra, partial differential equations, and integration. In an embodiment of the present application, an arithmetic execution unit (vector unit) based on the HiF8_DML data format can be embedded in a vector computing acceleration unit or a vector processor.
[0082] 3. Matrix calculation unit: A matrix is a two-dimensional array arranged in a rectangular array. A computing unit specially designed for matrix calculation with corresponding parallelism is called a matrix calculation unit, such as a systolic array processor. Matrix calculation units are mostly used for matrix calculations in fields such as HPC high-performance computing and AI machine learning, including matrix multiplication, matrix inversion, matrix decomposition, etc. In an embodiment of the present application, a matrix unit based on the HiF8_DML data format can be embedded in the matrix calculation acceleration unit.
[0083] 4. Tensor computing unit: A tensor is a multidimensional array with a dimension greater than 2, and a 3-dimensional array is common. A computing unit specially designed for tensor calculation with corresponding parallelism is called a tensor accumulation unit. Tensor computing units are mostly used in the field of AI machine learning, such as convolution operations. In an embodiment of the present application, a tensor unit based on the HiF8_DML data format can be embedded in the tensor computing acceleration unit.
[0084] Exemplarily, when the computing device 100 performs general computing, high-performance computing or AI training, a large amount of floating-point data is needed. At this point, the computing device 100 can decode the obtained floating-point number through the decoder 111 based on a floating-point number processing method provided in an embodiment of the present application to obtain decoded data. Exemplarily, the floating-point number can be read from the local memory 120 of the computing device 100, or obtained from the computing device 200 or the computing device 300 through the network, or obtained from other devices in the network. The decoder 111 decodes the obtained floating-point number, and after obtaining the decoded data, the decoded data can be transmitted to the computing unit 113, and the corresponding calculation is completed by the computing unit 113. The computing unit 113 transmits the calculation result to the encoder 112, and the encoder 112 re-encodes the calculation result into a floating-point number, which can be used for data storage and data transfer. For example, the processor 110 can save the floating-point number encoded by the encoder 112 into the memory 120.
[0085] Optionally, the structure and function of the computing device 200 and the computing device 300 in Figure 1 can be specifically referred to the computing device 100 shown in Figure 2. In some possible embodiments, the computing device 100, the computing device 200 and the computing device 300 may include more or fewer components than those shown in Figure 2, and the embodiments of the present application do not specifically limit this.
[0086] For easier understanding, the data format of the floating point numbers provided in the embodiment of the present application is first introduced below. Table 2 shows the data format of the floating point numbers provided in the embodiment of the present application.
[0087] Table 2
[0088] As shown in Table 2, the floating-point number provided in the embodiment of the present application may include a first sign field, a Dot field, a first exponent field, and a first mantissa field, wherein the Dot field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the floating-point number. Wherein, N is an integer greater than 1, and D is an integer greater than or equal to 0. The following embodiment is described using binary as an example, and the total bit width N of the floating-point number is 8. First, in conjunction with Tables 2 and 3, each field of the floating-point number is described in detail.
[0089] Table 3
[0090] 1. First Sign Field: The first sign field, also known as the sign bit, precedes the Dot field, as shown in Table 3. It occupies 1 bit of the floating-point number's total bit width, N, and is used to indicate the sign of the data. By default, 0 indicates positive and 1 indicates negative. You can also use 0 for negative and 1 for positive depending on your needs; this is not a limitation in this application.
[0091] 2. Dot field: It occupies DW bits in the total bit width N of the floating-point number, and the value of DW is 2 to 4. The numerical value (or encoding value) represented by the Dot field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the floating-point number, that is, the numerical value of the Dot field is D. When D is non-zero, the first exponent field is used to indicate the exponent of the floating-point number; the first mantissa field is used to indicate the mantissa of the floating-point number. When the bit width D is 0, the bit width indication field can also be used to indicate: whether the first mantissa field represents the mantissa of the floating-point number or the offset of the exponent of the floating-point number.
[0092] Optionally, the encoding method of the Dot field can use prefix code encoding. Prefix code encoding can include conventional prefix code encoding and unconventional prefix code encoding. Conventional prefix code encoding uses a shorter bit width to represent a smaller value and a longer bit width to represent a larger value. Unconventional prefix code encoding is the opposite, using a shorter bit width to represent a larger value and a longer bit width to represent a smaller value.
[0093] In some embodiments, the Dot domain can be encoded using an unconventional prefix code encoding method, that is, a longer bit width is used to represent a smaller value, and a shorter bit width is used to represent a larger value. For example, when the bit width DW is the first bit width value, the bit width DW1 occupied by the Dot domain is used to encode D1 values, and the bit width D of the first value domain belongs to any one of the D1 values; when the bit width DW is the second bit width value, the bit width DW2 occupied by the Dot domain is used to encode D2 values, and the bit width D of the first value domain belongs to any one of the D2 values. Among them, DW1 is smaller than DW2, and the smallest value among the D1 values is greater than the largest value among the D2 values. D1 and D2 are integers greater than or equal to 0.
[0094] Exemplarily, the Dot field occupies 2 to 4 bits in the total bit width N of the floating-point number, that is, the bit width (width) of the Dot field is 2 to 4. The specific encoding method can be shown in Table 4.
[0095] Table 4
[0096] As shown in Table 4, when the bit width of the Dot field is 2, a bit width of 2 bits can be used to encode and represent any one of the three values 2, 3, and 4. For example, coding "11" can represent the value "4", coding "10" can represent the value "3", and coding "01" can represent the value "2". When the bit width of the Dot field is 3, a bit width of 3 bits can be used to encode and represent the value 1. For example, coding "001" can represent the value "1". The values 1, 2, 3, and 4 are used to indicate the bit width D of the first-order code field, and the first-order code field is used to represent the exponent of the floating-point number. When the bit width of the Dot field is 4, a bit width of 4 bits can be used to encode and represent the value 0. At this time, the first-order code field does not exist.
[0097] The bit width indicator field uses unconventional prefix encoding. The bit width DW of the bit width indicator field is negatively correlated with the value indicated by the bit width indicator field. In other words, the bit width DW of the bit width indicator field is negatively correlated with the bit width D of the first exponent field. When the total bit width of the floating-point number is fixed, this can prevent jumps in the bit width of the first mantissa field caused by simultaneous increases in the bit widths of the bit width indicator field and the first exponent field. The bit width of the mantissa field determines the precision of the floating-point data. Therefore, the precision of the numerical value represented by the first floating-point number can be smoothly changed, preventing jumps in the precision of the numerical value represented by the first floating-point number. In particular, near the center of the exponent, jumps in the bit width of the mantissa field can be smoothed, that is, jumps in the precision of the numerical value near the center of the exponent can be smoothed.
[0098] When the bit width of the Dot field is 4, the Dot field is also used to indicate whether the first mantissa field represents the mantissa of the floating-point number or the offset of the exponent of the floating-point number. For example, in one embodiment, the code "0001" can represent the value "0". At this time, the bit width D of the first exponent field is 0 bits, the first mantissa field represents the mantissa of the floating-point number, and the code "0000" can indicate that the floating-point number is DML data. DML data refers to a value below a normal number (subnormal) or a value of an abnormal number (denormal). At this time, the first mantissa field represents the offset of the exponent of the floating-point number. That is to say, when the bit width of the Dot field is 4, if the Dot field is the preset value "0000", the Dot field indicates that the floating-point number is DML data, or in other words, indicates that the first mantissa field represents the offset of the exponent of the floating-point number. In another embodiment, the code "0000" can also be used to represent the value "0". In this case, the bit width D of the first exponent field is 0 bits, and the first mantissa field represents the mantissa of the floating-point number. The code "0001" can indicate that the floating-point number is DML data, that is, the Dot field is the preset value "0001", and the Dot field indicates that the floating-point number is DML data, or in other words, indicates that the first mantissa field represents the offset of the exponent of the floating-point number. From the above description, it can be seen that the bit width indication field can be described as Dot = [0, 4] & DML, indicating that: except for the DML flag, the value D represented by the code of the bit width indication field ranges from 0 to 4, and the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the floating-point number, that is, the bit width D of the first exponent field ranges from 0 to 4 bits.
[0099] 3. First exponent field: The exponent used to represent the floating-point number. It occupies a bit width D in the total bit width N of the floating-point number, and the value range of D is 0 to 4 bits.
[0100] For example, the first exponent field Es is used to represent the exponent of a floating point number. Assuming that the exponent of a floating point number has a numerical range of E, the value Ev represented by the first exponent field Es belongs to the numerical range E. The numerical range E can be determined by the following formula 1: E = (-1) Se ×[2 D-1 ,(2 D -1)] Formula 1
[0101] Wherein, Se is the sign bit of the value Ev, which can also be called the exponent sign bit of the floating-point number. Se occupies 1 bit in the bit width D of the exponent field and is used to indicate the positive or negative value of the first exponent field. In one embodiment, Se is 0 to indicate that the value Ev is positive, and Se is 1 to indicate that the value Ev is negative. In other embodiments, Se can also be 0 to indicate that the value Ev is positive, and Se can be 1 to indicate that the value Ev is negative. This application does not specifically limit this.
[0102] Combining Formula 1 and Table 3, when the bit width D of the first exponent field is 0, the value Ev represented by the first exponent field is 0, that is, the numerical range E of the floating-point exponent is 0. When the bit width D of the first exponent field is 1, the value Ev represented by the first exponent field can be 1 or -1, that is, the numerical range E of the floating-point exponent is ±1. When the bit width D of the first exponent field is 2 to 4, the value Ev represented by the first exponent field can be any value between [-2, -15] and [2, 15], that is, the numerical range E of the floating-point exponent is ±[2, 15]. Therefore, the numerical range E of the floating-point exponent that can be represented by the first exponent field is [-15, 15].
[0103] In an optional embodiment, when the bit width D is 0 and not DML, the exponent representing the floating-point number is 0. When the bit width D is non-zero, that is, when the bit width D is any value between 1 and 4, the first exponent field can be encoded using signed data (signed magnitude), which can also be called the encoding method of the sign bit Se following the original code. This representation method adds a sign bit in front of the value, that is, the highest bit of the original code is the sign bit, which is used to indicate the positive or negative of the value, where a sign bit of 0 indicates a positive number and a sign bit of 1 indicates a negative number. The remaining bits in the original code except the sign bit are used to indicate the size of the value, that is, the amplitude of the original code. For example, the original code 1001 represents -1, and 0011 represents +3.
[0104] In an embodiment of the present application, when the bit width D is non-zero, the first order code field in the floating point number can be encoded using an exponent sign bit Se following the original code amplitude, which can be expressed as Es: {Se+Mag[2:end]}, wherein the exponent sign bit Se is the sign bit extracted from the initial original code, used to indicate the positive or negative value of the first order code field; Mag is used to indicate the amplitude of the value of the first order code field, end=D-1, and Mag[2:end] is used to indicate the values from the 2nd to the D-1th bits of the first order code field. For different bit widths D, the highest bit b1 of the amplitude Magg of the value of the first order code field is always 1 (i.e., 1'b1). Therefore, the highest bit 1'b1 does not occupy the bit width during encoding, i.e., the highest bit 1'b1 is hidden and not actually stored. During subsequent decoding, the highest bit 1'b1 can be directly supplemented to obtain the encoding value Ei of the order code field in the normalized floating point number: {Se+1'b1+Mag[2:end]}. Normalized floating-point numbers refer to floating-point numbers that comply with the IEEE 754 binary floating-point standard. Normalized floating-point numbers can be used for floating-point calculations. Because the highest bit of the exponent field's encoded value, Ei, does not occupy the first exponent field and is not actually stored, this significantly saves storage space and reduces data storage and transfer costs.
[0105] For example, as shown in Table 5, when the Dot field is a DML flag, it indicates that the floating-point number is DML data. In this case, the bit width of the first mantissa field is 3, that is, the first mantissa field uses 3 bits to represent the offset of the exponent of the floating-point number. When the Dot field indicates that the bit width D of the first exponent field is 0, the encoding Es of the first exponent field in the floating-point number is None, and the value Ei of the exponent field of the normalized floating-point number obtained by decoding is 0, indicating that the exponent value Ev of the floating-point number is 0. In this case, the bit width of the Dot field is 4, and the bit width of the first mantissa field is 3, that is, the first mantissa field uses 3 bits to represent the mantissa of the floating-point number.
[0106] When the Dot field indicates that the bit width D of the first exponent field is 1, the encoding of the first exponent field in the floating-point number is Es = {Se}, and the decoded value of the first exponent field, that is, the value of the exponent field of the normalized floating-point number is Ei = {Se, 1}, where 1 is the value of the highest bit in Mag. As mentioned above, the highest bit in Mag is always 1. When the bit width D = 1, the exponent Ev of the floating-point number represented by the first exponent field can be 1 or -1, that is, the numerical range of the exponent Ev of the floating-point number can be ±1. At this time, the bit width of the Dot field is 3, and the bit width of the first mantissa field is 3, that is, the first mantissa field uses 3 bits to represent the mantissa of the floating-point number.
[0107] When the Dot field indicates that the bit width of the first exponent field is D=2, the encoding of the first exponent field in the floating-point number is Es={Se, [2]}, where [2] represents the value of the second bit in Mag. The decoded value of the first exponent field, that is, the value of the exponent field of the normalized floating-point number is Ei={Se, 1, [2]}, where 1 is the value of the highest bit in Mag. When the bit width D=2, the exponent Ev of the floating-point number represented by the first exponent field can be any value between [-2, -3] or [2, 3], that is, the numerical range of the exponent Ev of the floating-point number can be ±[2, 3]. At this time, the bit width of the Dot field is 2, and the bit width of the first mantissa field is 3, that is, the first mantissa field uses 3 bits to represent the mantissa of the floating-point number.
[0108] When the Dot field indicates a first exponent field bit width of D = 3, the encoding of the first exponent field in the floating-point number is Es = {Se, [2:3]}, where [2:3] represents the values of bits 2 and 3 in Mag. The decoded value of this first exponent field, i.e., the value of the normalized floating-point exponent field, Ei = {Se, 1, [2:3]}, where 1 represents the value of the most significant bit in Mag. When the bit width D = 3, the exponent Ev of the floating-point number represented by the first exponent field can be any value between [-4, -7] and [4, 7]. In other words, the exponent Ev of the floating-point number can have a numerical range of ±[4, 7]. In this case, the bit width of the Dot field is 2, and the bit width of the first mantissa field is also 2, meaning that the first mantissa field uses 2 bits to represent the mantissa of the floating-point number.
[0109] When the Dot field indicates a first exponent field bit width of D = 4, the encoding of the first exponent field in the floating-point number is Es = {Se, [2:4]}, where [2:4] represents the values of bits 2, 3, and 4 in Mag. The decoded value of this first exponent field, i.e., the normalized exponent field value of the floating-point number, is Ei = {Se, 1, [2:4]}, where 1 represents the value of the most significant bit in Mag. When the bit width D = 4, the exponent Ev of the floating-point number represented by the first exponent field can be any value between [-8, -15] and [8, 15], meaning the exponent Ev of the floating-point number can have a numerical range of ±[8, 15]. In this case, the Dot field bit width is 2, and the first mantissa field bit width is 1, meaning the first mantissa field uses 1 bit to represent the mantissa of the floating-point number.
[0110] It can be seen that when D is greater than 1, the encoding of the first-order code field in the floating-point number is Es={Se+Mag[2:end]}, where Mag[2:end] includes the remaining bits of Mag except the highest bit 1'b1, and the bit width occupied by Mag[2:end] in the first-order code field is D-1.
[0111] Table 5
[0112] 4. First mantissa field: used to represent the mantissa of a floating-point number, or the offset of the exponent of a floating-point number. The bit width occupied by the first mantissa field in a floating-point number is (N-1-DW-D) bits. Exemplarily, the Dot field is used to indicate whether the first mantissa field represents the mantissa of a floating-point number or the offset of the exponent of a floating-point number. In some embodiments, when the bit width DW of the Dot field is 4 and the value of the Dot field is "0001", it is used to indicate that the first mantissa field represents the mantissa of a floating-point number; when the bit width DW of the Dot field is 4 and the value of the Dot field is "0000", it is used to indicate that the first mantissa field represents the offset of the exponent of a floating-point number. In other embodiments, the opposite setting may also be adopted. When the bit width DW of the Dot field is 4 and the value of the Dot field is "0000", it is used to indicate that the first mantissa field represents the mantissa of the floating-point number; when the bit width DW of the Dot field is 4 and the value of the Dot field is "0001", it is used to indicate that the first mantissa field represents the offset of the exponent of the floating-point number.
[0113] When the first mantissa field represents the mantissa of a floating-point number, it is used to store the value after the decimal point. For example, if the decimal digits of 1.xxx are stored, assuming the integer digits 1'b1 are hidden. For example, if the code stored in the first mantissa field is 10011, the decoded value of the first mantissa field is M = 0.10011, and the value represented by the first mantissa field is 1.M, that is, 1.10011.
[0114] As shown in Table 3, when the first mantissa field represents the offset of the exponent of the floating-point number, the bit width of the first mantissa field is 3, that is, the first mantissa field uses 3 bits to encode the integer value M of 0 to 7, which is used to represent the exponent value of M-23, where M can be called the offset of the exponent. It can be seen that the first mantissa field can supplement the numerical range of the exponent [-23, -16]. Among them, when the 3 bits of the first mantissa field are all 0, that is, M=0, M-23=-23, the floating-point number is used to represent a special value (which will be described in detail below). Therefore, the numerical range of the exponent that can be represented by the first mantissa field is [-22, -16]. Combined with the above-mentioned first exponent field (as shown in Table 5), the HiF8_DML floating-point number provided in the embodiment of the present application can represent the exponent range of the floating-point number [-22, 15]. When the exponent of the floating-point number represented by HiF8_DML is 15, the value X of the floating-point number represented is the largest, which can be approximated as X max-pos =2 15 =32768. When the exponent of the floating point number represented by HiF8_DML is -22, the value X of the floating point number represented is the smallest and can be approximated as X min-pos =2 -22 ≈2.38×10 -7 .
[0115] When the first mantissa field represents the offset of the exponent of a floating-point number, the corresponding floating-point number is DML data, and the value X of the floating-point number can be expressed as: X=(-1) s ×2 M-23
[0116] Where S is the value of the sign field of the floating-point number, and M is the value represented by the first mantissa field of the floating-point number. In this case, the mantissa of the floating-point number uses the default value of 1.
[0117] When the first mantissa field represents the mantissa of a floating-point number, the corresponding floating-point number is normal data, and the value X of the floating-point number can be expressed as: X=(-1) s ×2 Ev ×1.M
[0118] Where S is the value of the floating-point number's sign field, Ev is the value represented by the floating-point number's first exponent field, i.e., the exponent of the floating-point number. M is the value represented by the floating-point number's first mantissa field, and 1.M is the mantissa of the floating-point number. The 1 in 1.M is also considered a significant bit. Therefore, the number of significant bits in the floating-point number's mantissa is the bit width of the first mantissa field + 1.
[0119] In summary, the floating point value X can be expressed as:
[0120] The HiF8_DML floating-point number provided in the embodiments of the present application can also support four special values: 0 (Zero), not a number (NaN), positive infinity (+inf), and negative infinity (-inf). These four special values can be expressed by the four boundary values that HiF8_DML can represent.
[0121] (1) Zero: When the value of the sign field S of a floating-point number is 0 and the Dot field is 0000, that is, the Dot field indicates that the floating-point number is DML data. In other words, when the first mantissa field represents the offset of the floating-point number's exponent, and all three bits of the first mantissa field are 0, the corresponding floating-point number can represent ±0, that is, zero. The above description can be summarized as follows: when HiF8 = 8'b 0 0000 000, the value represented is X = Zero.
[0122] (2) NaN: When the value of the sign field S of a floating-point number is 1 and the Dot field is 0000, that is, the Dot field indicates that the floating-point number is DML data, that is, the first mantissa field represents the offset of the floating-point number's exponent, and when all three bits of the first mantissa field are 0, the corresponding floating-point number can represent a non-numeric value, that is, NaN. The above description can be summarized as follows: when HiF8 = 8'b 1 0000 000, it means X = NaN.
[0123] (3) Positive infinity: When the first exponent field of a floating-point number has a bit width of 4, the first mantissa field has a bit width of 1, and the value of the first exponent field is 0111 and the value of the first mantissa field is 1, if the value of the sign field S of the floating-point number is 0, then the corresponding floating-point number can be represented as +inf. The above description can be summarized as follows: when the first exponent field Es = 4'b0111 = 15 and the first mantissa field M = 1'b1, that is, HiF8 = 8'b 0 11 0111 1, it represents X = +inf.
[0124] (4) Negative infinity: When the first exponent field of a floating-point number has a bit width of 4, the first mantissa field has a bit width of 1, and the value of the first exponent field is 0111 and the value of the first mantissa field is 1, if the value of the sign field S of the floating-point number is 1, then the corresponding floating-point number can be represented as -inf. The above description can be summarized as follows: when the first exponent field Es = 4'b0111 = 15 and the first mantissa field M = 1'b1, that is, HiF8 = 8'b 1 11 0111 1, it represents X = -inf.
[0125] FIG3 is a schematic diagram of the significant precision (significant precision)-exponent distribution of a floating-point number of HiF8_DML provided in an embodiment of the present application, wherein the size of the significant bit can represent the precision of the floating-point number. When the first mantissa field is used to represent the mantissa of the floating-point number, the significant bit of the mantissa can be the bit width of the first mantissa field plus 1. As shown in FIG3 , HiF8_DML has a tapered precision feature and can provide up to 4 significant bits of precision. At the same time, HiF8_DML has an exponent range of [-22, 15], which is much larger than the exponent range of [-9, 8] or [-16, 15] of FP8, and is almost equivalent to the exponent range of [-24, 15] of FP16. It can be seen that HiF8_DML can expand the dynamic range of the value represented by the floating-point number. Moreover, the numerical precision near the exponent center is significantly higher than the numerical precision far away from the exponent center. For example, when the exponent value is [-3, 3], the mantissa has the highest number of significant digits, which is 4 significant digits. When the exponent value is between ±[4, 7], the mantissa has 3 significant digits. When the exponent value is between ±[8, 15], the mantissa has 2 significant digits. When the exponent value is less than -15, the mantissa has 1 significant digit. At this time, the mantissa is the default value 1.
[0126] The HiF8_DML provided in the embodiment of the present application can be used for data storage or data transfer. Correspondingly, the embodiment of the present application also provides a floating point number processing method, which can be executed by the processor 110 shown in Figure 2. As shown in Figure 4, the method may include the following steps:
[0127] S401: Obtain a first floating-point number.
[0128] The processor can read the first floating-point number from a local memory of the computing device to which the processor belongs, or can obtain the first floating-point number from another computing device through a network. The first floating-point number may include a first sign field, a bit width indication field, a first exponent field, and a first mantissa field. Among them, the first sign field is used to indicate the sign of the first floating-point number, that is, to indicate the positive or negative of the first floating-point number, and the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field can also be used to indicate: whether the first mantissa field represents the mantissa of the first floating-point number or the offset of the exponent of the first floating-point number.
[0129] S402: Decode the first floating-point number to obtain a sign, an exponent, and a mantissa.
[0130] In some embodiments, the first floating-point number can be decoded to obtain a second floating-point number; the second floating-point number may include a second sign field, a second exponent field, and a second mantissa field, the second sign field is used to indicate the above sign, the second exponent field is used to indicate the above exponent, and the second mantissa field is used to indicate the above mantissa.
[0131] Exemplarily, the value of the first sign field of the first floating-point number can be used as the value of the second sign field of the second floating-point number, and the second exponent field and second mantissa field of the second floating-point number can be determined based on the values of the first exponent field and the first mantissa field of the first floating-point number.
[0132] When the bit width D of the first exponent field is 0, if the bit width indication field indicates that the first mantissa field represents the mantissa of the first floating-point number, the value indicated by the first mantissa field can be used as the value indicated by the second mantissa field of the second floating-point number, and the first preset exponent can be used as the value indicated by the second exponent field of the second floating-point number, wherein the first preset exponent can be -15. If the bit width indication field indicates that the first mantissa field represents the offset of the exponent of the first floating-point number, the preset mantissa can be used as the value indicated by the second mantissa field of the second floating-point number, and the second preset exponent can be corrected using the value of the first mantissa field to obtain the value of the exponent field of the second floating-point number. The preset mantissa can be 1, and the second preset exponent can be -23.
[0133] When the bit width D of the first exponent field is non-zero, the first exponent field is used to represent the exponent of the first floating-point number, and the first mantissa field is used to represent the mantissa of the first floating-point number. In this case, the first exponent field and the first mantissa field can be determined in the first floating-point number based on the bit width indication field, and then the value indicated by the first mantissa field is used as the value indicated by the second mantissa field of the second floating-point number, and the value indicated by the first exponent field is used as the value indicated by the second exponent field of the second floating-point number.
[0134] In some embodiments, a decoder in a processor can be used to convert a first floating-point number into a second floating-point number. Exemplarily, a schematic diagram of the working principle of the decoder can be shown in Figure 5, and the decoder can also be called a HiF8_DML decoder. As shown in Figure 5, the first floating-point number includes a sign field, a Dot field, a first exponent field, and a first mantissa field in sequence, and the total bit width N of the first floating-point number is 8. Among them, the sign field occupies 1 bit and is the highest bit in the first floating-point number. The Dot field occupies 2 to 4 bits, the first exponent field occupies 0 to 4 bits, and the first mantissa field occupies 1 to 3 bits. First, for the sign field in the floating-point number, the decoder can directly read and output its value, S = 0 or 1. Secondly, the decoder can determine the bit width of the Dot field based on the value of the second, third, or fourth bit in the first floating-point number. For example, if the value of the second bit in the first floating-point number is "1", it can be determined that the bit width of the Dot field is 2; if the value of the second bit in the first floating-point number is "0" and the value of the third bit is "1", it can be determined that the bit width of the Dot field is also 2; if the values of the second and third bits in the first floating-point number are both "0", and the value of the fourth bit is "1", it can be determined that the bit width of the Dot field is 3; if the values of the second, third and fourth bits in the first floating-point number are all "0", it can be determined that the bit width of the Dot field is 4.
[0135] If the bit width of the Dot field is 4, it can be determined based on the value of the Dot field whether the first mantissa field represents the mantissa of the first floating-point number or the offset of the exponent of the first floating-point number.
[0136] If the first mantissa field represents the mantissa of the first floating-point number, the first floating-point number is normal data. The decoder can read the value of the first mantissa field, use the value of the first mantissa field as the value of the second mantissa field of the second floating-point number, and use the first preset exponent as the value of the second exponent field of the second floating-point number. The first preset exponent can be -15.
[0137] If the first mantissa field represents an offset from the exponent of the first floating-point number, the first floating-point number is DML data. The decoder can read the value of the first mantissa field and use it to correct the second preset exponent to obtain the value of the exponent field of the second floating-point number. For example, the second preset exponent may be -23. When the value of the first mantissa field is M, the value of the second exponent field of the second floating-point number may be M-23. The decoder can use the preset mantissa 1 as the value represented by the second mantissa field of the second floating-point number.
[0138] If the bit width of the Dot field is 2 or 3, the decoder can extract the Dot field and, by performing a multiplexing (MUX) operation on the Dot field, decode the Dot field based on a preset encoding rule to obtain the value D of the Dot field. The preset encoding rule can be as shown in Table 4 above, where the code "11" can represent the value "4", the code "10" can represent the value "3", the code "01" can represent the value "2", and the code "001" can represent the value "1".
[0139] Based on the value D of the Dot field, the decoder can extract the first exponent field and the first mantissa field located after the first exponent field after the Dot field, and use the value indicated by the first exponent field as the value indicated by the second exponent field of the second floating-point number; and use the value indicated by the first mantissa field as the value indicated by the second mantissa field of the second floating-point number. At this point, the decoder can determine the values of the second exponent field and the second mantissa field of the second floating-point number and the value of the second sign field based on the values of the first exponent field and the first mantissa field in the first floating-point number and the value of the first sign field.
[0140] For example, the first floating-point number input to the decoder is "11010110", with a total bit width N of 8 bits. The first bit "1" is the sign field, and the second to third bits "10" are the Dot field. According to the preset encoding rules, the value of the Dot field obtained by decoding is 3, that is, the bit width of the first exponent field is determined to be 3 bits. Therefore, the decoder can accurately extract the fourth to sixth bits "101" in the first floating-point number as the first exponent field, which is used to represent the exponent of the first floating-point number. The remaining seventh to eighth bits "10" are the first mantissa field, which is used to represent the mantissa of the first floating-point number. Based on this, the value represented by the exponent field of the second floating-point number can be decoded to be 5, and the value represented by the mantissa field of the second floating-point number is 0.10.
[0141] In other embodiments, the first floating-point number may be decoded to directly obtain the sign, exponent, and mantissa. For example, the sign may be determined based on the value of the first sign field of the first floating-point number. For example, if the value of the first sign field is 0, the sign is positive, and if the value of the first sign field is 1, the sign is negative.
[0142] When the first exponent field bit width D is 0, it indicates that the first exponent field does not exist. In this case, if the Dot field indicates that the first mantissa field represents the mantissa of the first floating-point number, the exponent may be the first preset exponent -15. If the Dot field indicates that the first mantissa field represents the offset of the exponent of the first floating-point number, the exponent may be the value obtained by modifying the second preset exponent -23 by the offset, and the mantissa may be the preset mantissa 1.
[0143] When the first exponent field bit width D is not 0, the mantissa is the value indicated by the first mantissa field of the first floating-point number, and the exponent is the value indicated by the first exponent field of the first floating-point number.
[0144] After obtaining the sign, exponent and mantissa, calculations can be performed based on the sign, exponent and mantissa. For example, in other embodiments, as shown in FIG6 , the floating-point number processing method executed by the processor may include the following steps:
[0145] S601: Obtain a first floating-point number.
[0146] S602: Decode the first floating-point number to obtain a sign, an exponent, and a mantissa.
[0147] S603, performing calculation using the sign, exponent and mantissa to obtain a calculation result.
[0148] In some embodiments, the processor includes a decoder, and the processor can decode the first floating-point number through the decoder to obtain a sign, an exponent, and a mantissa. The decoder can transmit the decoded sign, exponent, and mantissa to a calculation unit in the processor, which receives the sign, exponent, and mantissa and performs corresponding calculations to obtain a calculation result. The calculation result can be a floating-point number using the same encoding method as the above-mentioned second floating-point number, that is, the calculation result can include a sign field, an exponent field, and a mantissa field, the sign field is used to indicate the sign of the calculation result, the exponent field is used to indicate the exponent of the calculation result, and the mantissa field is used to indicate the mantissa of the calculation result.
[0149] S604: Convert the calculation result into a third floating-point number based on the exponent and mantissa of the calculation result.
[0150] The third floating-point number is a floating-point number encoded in the same manner as the first floating-point number, and the third floating-point number may include a first sign field, a Dot field, a first exponent field, and a first mantissa field.
[0151] The processor may use the value of the sign field of the calculation result as the value of the first sign field of the third floating-point number, and determine the numerical range to which the calculation result belongs according to the exponent and mantissa of the calculation result.
[0152] If the numerical range of the calculation result falls within the first set range, the value of the first mantissa field of the third floating-point number can be determined based on the exponent field and mantissa field of the calculation result, and the Dot field can be set to the first value; the first value indicates that the first mantissa field represents an offset of the exponent of the first floating-point number. The first set range refers to the range of the DML data.
[0153] If the numerical range of the calculation result falls within the second set range, the values of the first exponent field, first mantissa field, and Dot field of the third floating-point number can be determined based on the exponent field and mantissa field of the calculation result, respectively. The second set range refers to the range of normal data. If, based on the exponent field and mantissa field of the calculation result, the value of the first exponent field is determined to be the first preset exponent, i.e., the bit width D of the first exponent field is 0, the Dot field can be set to the second value, which is used to indicate that the first mantissa field represents the mantissa of the third floating-point number.
[0154] In some embodiments, the processor may further include an encoder. After the calculation unit in the processor obtains the calculation result, the calculation result can be transmitted to the encoder, and the calculation result is converted into a third floating-point number by the encoder. The encoder may also be called a HiF8_DML encoder. Exemplarily, a schematic diagram of the working principle of the encoder may be shown in Figure 7. The calculation result may include a sign domain, an exponent domain, and a mantissa domain. The encoder may use the value of the sign domain of the calculation result as the value of the first sign domain of the third floating-point number. The encoder may determine the numerical range to which the calculation result belongs based on the value of the exponent domain of the calculation result.
[0155] If the numerical range to which the calculation result belongs is the first set range, that is, the calculation result is DML data, the encoder can determine the value of the first mantissa field of the third floating-point number based on the exponent and mantissa of the calculation result. Exemplarily, the encoder can process the mantissa of the calculation result, and adjust the exponent of the calculation result according to the processed mantissa, and obtain the first mantissa field of the third floating-point number based on the adjusted exponent. For example, assuming that the mantissa of the calculation result is 1.XXXXXX, the encoder can perform mantissa shift and rounding operations on the mantissa of the calculation result, round it to an integer bit, and obtain the processed mantissa 1.0, which will cause the value of the exponent to be added by 1, so the exponent of the calculation result can be adjusted, and the value of the adjusted exponent is added by 23 to obtain the value of the first mantissa field of the third floating-point number, and the first mantissa field of the third floating-point number is encoded according to the value of the first mantissa field. The encoder can set the Dot field to a first value, and the first value is used to indicate that the first mantissa field represents the offset of the exponent of the third floating-point number.
[0156] If the numerical range to which the calculation result belongs is the second set range, that is, the calculation result is normal data, the encoder can determine the first exponent field and the first mantissa field of the third floating-point number according to the exponent and mantissa of the calculation result, respectively, and determine the bit width indication field of the third floating-point number based on the first exponent field. For example, the encoder can process the mantissa of the calculation result, and adjust the exponent of the calculation result according to the processed mantissa, obtain the value of the first mantissa field of the third floating-point number based on the processed mantissa, and obtain the value of the first exponent field of the third floating-point number based on the adjusted exponent, determine the bit width D occupied by the first exponent field based on the value of the first exponent field, and determine the value of the bit width indication field based on the bit width D occupied by the first exponent field. When the bit width D occupied by the first exponent field is 0, the bit width indication field is set to the second value; the second value is used to indicate that the first mantissa field represents the mantissa of the first floating-point number. For example, the encoder can encode the Dot field, the first exponent field, and the first mantissa field of the third floating-point number by performing a leading 1 operation on the absolute value of the exponent in the calculation result, and performing operations such as shifting and rounding the mantissa in the calculation result.
[0157] For example, assuming that the sign bit of the calculation result is 1, the exponent is "00011" (i.e., 3), and the mantissa is "1111." The encoder can determine that the sign bit of the third floating-point number is "1." The encoder performs a forward search operation on the absolute value of the exponent "00011" until the leading 1 is found. It can determine that the bit width of the first exponent field of the third floating-point number is 2 bits, and the encoding value of the first exponent field is "01," where "0" is the exponent sign bit, indicating that the exponent is positive, and the highest bit "1" in the exponent amplitude "11" can be hidden and does not occupy the bit width. This has been described above and will not be repeated here. Based on the bit width of 2 bits of the first exponent field, the encoder can determine that the value represented by the Dot field of the third floating-point number is 2. Still taking the encoding rules shown in Table 4 as an example, the value represented by the Dot field is 2, and it can be determined that the encoding value corresponding to the Dot field is "01," which occupies a bit width of 2 bits. The remaining bit width of the first mantissa field that can be encoded is 3 bits. The bit width of the input calculation result's mantissa, "1111," is greater than the bit width of the first mantissa field of the third floating-point number. Therefore, the mantissa "1111" can be rounded to obtain the mantissa "10.000," including the hidden digit to the left of the decimal point. At this time, the decimal point in the mantissa "10.000" must be shifted one position to the left, resulting in the mantissa "1.0000," including the hidden digit to the left of the decimal point. Since the decimal point is shifted one position to the left, the exponent must be increased by 1 (i.e., 3 plus 1 equals 4). This results in an amplitude of 4 for the first exponent field, corresponding to the encoded value of "100." The highest bit, "1," in "100" is hidden, resulting in the final encoded value of the first exponent field being "00," and the final encoded value of the first mantissa field being "000." Ultimately, the encoder encodes the third floating-point number as "10100000."
[0158] The third floating-point number finally obtained is the data of HiF8_DML provided in the embodiment of the present application. The third floating-point number can be used for data storage and data transfer to reduce the resources required for data storage or data transfer.
[0159] Based on the same inventive concept as the above embodiment, the present embodiment further provides a floating point number processing method for obtaining the first floating point number used in the above process. As shown in FIG8 , the method may include the following steps:
[0160] S801, obtain the sign, exponent and mantissa.
[0161] S802: Obtain a first floating-point number based on the sign, the exponent, and the mantissa.
[0162] For example, the processor may obtain data to be encoded. The data to be encoded may be data in a format different from the format of the first floating-point number, or may be an operation result output by the computing unit. The data to be encoded may include a sign, an exponent, and a mantissa. The processor extracts the sign, exponent, and mantissa from the data to be encoded, and obtains the first floating-point number based on the sign, exponent, and mantissa.
[0163] The process of the processor obtaining the first floating-point number based on the sign, the exponent and the mantissa can be performed with reference to the process of obtaining the third floating-point number based on the calculation result above, which will not be repeated here.
[0164] The floating-point number provided by the embodiment of the present application, on the basis of the standard sign field, exponent field and mantissa field, adds an additional Dot field, and indicates the effective bit width of the first exponent field of the floating-point number through the Dot field. The bit width is the bit width D occupied by the exponent of the floating-point number when it is actually stored, so that the bit width of the first exponent field in the floating-point number can change dynamically with the value of the Dot field. Correspondingly, the bit width of the first mantissa field in the floating-point number also changes dynamically, so that the floating-point number has a cone-shaped precision feature, so that the data at the center of the exponent has a higher mantissa bit width, that is, it has higher precision; the farther the data is from the exponent center, the mantissa bit width gradually decreases and the precision decreases. The exponent of the floating-point number can determine the numerical range of the floating-point number. Therefore, the embodiment of the present application can effectively balance the total bit width, numerical range and numerical precision of the floating-point number, and achieve the different requirements for the numerical range and numerical precision of the floating-point number in various scenarios without increasing the total bit width and the data storage or data transfer cost, thereby improving the use effect of the floating-point number.
[0165] Exemplarily, the floating-point number representation method provided by the embodiment of the present application can be applied to the lossy compression of high-precision data and the early approximate solution calculation of the mixed-precision matrix solver. For general computing and high-performance computing, the embodiment of the present application can obtain higher computing task convergence speed and accuracy at the same total bit width, that is, the same data storage or data transfer overhead. For AI neural network training and reasoning, the embodiment of the present application can meet the functional and accuracy requirements of neural network training and reasoning at the same total bit width. When the bit width occupied by the first order code field is 0, the embodiment of the present application provides 3 bits of encoding space for the mantissa. The 3-bit mantissa precision can already meet the needs of AI neural network training and reasoning. Therefore, the embodiment of the present application sets the bit width of the Dot field to 4 bits, of which 1 bit is used to indicate whether the first mantissa field represents the mantissa or the exponent offset. If the first mantissa field represents the exponent offset, it means that the floating point number is DML data. When the floating-point number is DML data, the mantissa can use the default value of 1. The mantissa precision can meet the needs of AI neural network training and inference. At this time, using the first mantissa field to represent the offset of the exponent can expand the dynamic range of the exponent and expand the numerical range of the floating-point number that can be represented.
[0166] For example, as mentioned above, the floating-point representation provided in the embodiment of the present application has a numerical range of [-22, 15], which is much larger than the numerical range of [-9, 8] or [-16, 15] of FP8. The numerical range provided in the embodiment of the present application can also meet the needs of AI neural network training and reasoning. For example, for the end-to-end training of a large language model (LLM), as shown in Figure 9, the traditional floating-point representation FP8 is used. When the LLM is trained end-to-end, since the numerical range of the data that can be represented by the floating-point number is small, when a large amount of data with very small values appears, these data cannot be distinguished and are considered to be the same value, resulting in the inability to obtain correct training results. This is manifested in that when the training reaches more than one hundred times, the loss value suddenly increases, the LLM training collapses, and it cannot converge. When the floating-point number representation method HiF8_DML provided in the embodiment of the present application is used to perform end-to-end training on LLM, as shown in Figure 10, since the numerical range of data that can be represented by floating-point numbers is large, it can meet the needs of LLM training. Therefore, the training accuracy is improved, and LLM can converge normally, thereby improving the performance and stability of LLM training.
[0167] In summary, the embodiments of the present application can ensure that the numerical accuracy of AI neural network training and reasoning is minimal, while freeing up coding space to express DML data, thereby achieving a good balance between numerical accuracy and numerical range.
[0168] Furthermore, the present embodiment also defines the range of values that can be encoded in the first-order code field at different bit widths, as well as the range of values that can be encoded in the first mantissa field when representing the exponent offset. This effectively avoids the problem of overlapping exponent values at different bit widths, ensuring that the data encoding method of the present embodiment is free of information duplication and non-redundant encoding. Based on this, the highest bit of the original code amplitude in the first-order code field can be hidden and not stored, further reducing the data storage or data transfer costs of floating-point numbers.
[0169] Based on the same design concept as the above-mentioned method embodiment, the embodiment of the present application also provides a floating-point number processing device. The floating-point number processing device can be applied to the computing device shown in Figure 1 or Figure 2, or applied to the decoder 111 shown in Figure 2. The floating-point number processing device can be used to implement the functions of the above-mentioned method embodiment, thereby achieving the beneficial effects possessed by the above-mentioned method embodiment. As shown in Figure 11, the floating-point number processing device 1100 can include a floating-point number acquisition module 1101 and a decoding module 1102.
[0170] The floating-point number acquisition module 1101 can be used to acquire a first floating-point number; the first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field is further used to indicate: whether the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number;
[0171] The decoding module 1102 may be configured to decode the first floating-point number to obtain a sign, an exponent, and a mantissa.
[0172] In some embodiments, the decoding module 1102 can be specifically used to: decode the first floating-point number to obtain a second floating-point number; the second floating-point number includes a second sign field, a second exponent field and a second mantissa field, the second sign field is used to indicate the sign, the second exponent field is used to indicate the exponent, and the second mantissa field is used to indicate the mantissa.
[0173] In some embodiments, the floating-point processing device 1100 may further include a calculation module, which may be configured to perform calculations using the sign, exponent, and mantissa. In other embodiments, the calculation module may be located in a calculation unit outside the decoder, and the floating-point processing device 1100 may transmit the sign, exponent, and mantissa to the calculation unit, which then performs calculations using the sign, exponent, and mantissa.
[0174] It should be noted that, in some embodiments, the floating-point number acquisition module 1101 can be used to execute any step in the floating-point data processing method, and the decoding module 1102 can be used to execute any step in the floating-point data processing method. The steps that the floating-point number acquisition module 1101 and the decoding module 1102 are responsible for implementing can be specified as needed. The floating-point number acquisition module 1101 and the decoding module 1102 each implement different steps in the floating-point data processing method to achieve the full functionality of the floating-point number processing device.
[0175] The functional modules in the embodiments of the present application may be integrated into a processor, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional units.
[0176] Based on the same design concept as the above-mentioned method embodiment, the embodiment of the present application also provides a floating-point number processing device. The floating-point number processing device can be applied to the computing device shown in Figure 1 or Figure 2, or applied to the encoder 112 shown in Figure 2. The floating-point number processing device can be used to implement the functions of the above-mentioned method embodiment, thereby achieving the beneficial effects possessed by the above-mentioned method embodiment. As shown in Figure 12, the floating-point number processing device 1200 can include a data acquisition module 1201 and an encoding module 1202.
[0177] The data acquisition module 1201 can be used to obtain a sign, an exponent, and a mantissa; the encoding module 1202 can be used to obtain a first floating-point number based on the sign, the exponent, and the mantissa. The first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field. The first sign field is used to indicate the sign of the first floating-point number, and the bit width indication field is used to indicate the bit width D occupied by the first exponent field within the total bit width N of the first floating-point number. The bit width indication field is also used to indicate whether the first mantissa field represents the mantissa of the first floating-point number or the offset of the exponent of the first floating-point number.
[0178] It should be noted that, in some embodiments, the data acquisition module 1201 can be used to execute any step in the floating-point data processing method, and the encoding module 1202 can be used to execute any step in the floating-point data processing method. The steps that the data acquisition module 1201 and the encoding module 1202 are responsible for implementing can be specified as needed. The data acquisition module 1201 and the encoding module 1202 respectively implement different steps in the floating-point data processing method to achieve the full functionality of the floating-point number processing device.
[0179] The functional modules in the embodiments of the present application may be integrated into a processor, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional units.
[0180] Based on the same technical concept as the above-mentioned method embodiments, the present application also provides a computing device in the embodiments. The computing device can be any of the computing devices shown in Figure 1 or Figure 2. The computing device can be used to implement the functions of the method embodiments shown in Figure 4, Figure 6, or Figure 8, thereby achieving the beneficial effects of the above-mentioned method embodiments.
[0181] In some embodiments, the structure of the computing device 1300 can be as shown in Figure 13, including a processor 1301 and a memory 1302 connected to the processor 1301. The processor 1301 and the memory 1302 can be connected to each other via a bus. The processor 1301 can be a general-purpose processor, such as a microprocessor, or other conventional processor. The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0182] Among them, the memory 1302 can be used to store software programs and modules, and the processor 1301 executes various functional applications and data processing of the terminal device 1300 by running the software programs and modules stored in the memory 1302, such as any floating-point processing method provided in the embodiments of the present application.
[0183] The memory 1302 may primarily include a program storage area and a data storage area. The program storage area may store an operating system, at least one application program, and the like; the data storage area may be used to store user data, etc. Furthermore, the memory 1302 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state memory device.
[0184] The processor 1301 in the computing device 1300 is configured to execute computer instructions or programs stored in the memory 1302 to perform the functions of any of the above-described method embodiments. In some embodiments, the processor 1301 may include one or more processing units, which may be independent devices or integrated into one or more processors. The processor 1301 may also include a controller that generates operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0185] In some embodiments, the processor 1301 of the computing device 1300 may include a decoder, and the processor 1301 may be used to implement the floating-point number processing method shown in Figure 4; in other embodiments, the processor 1301 of the computing device 1300 may include an encoder, and the processor 1301 may be used to implement the floating-point number processing method shown in Figure 8; or, in other embodiments, the processor 1301 of the computing device 1300 may include a decoder and an encoder, and the processor 1301 may be used to implement the floating-point number processing method shown in Figure 6.
[0186] In one embodiment, the computing device 1300 may further include a communication module, which may be used to communicate with a network device.
[0187] It is understood that the structures illustrated in the embodiments of the present application do not constitute specific limitations on the terminal device. In other embodiments of the present application, the terminal device may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0188] The present application also provides a computer program product comprising computer-executable instructions. In one embodiment, the computer-executable instructions are used to enable a computer to perform the functions of the above method embodiment.
[0189] Computer-executable instructions can be stored in a computer-readable storage medium. The present application also provides a computer-readable storage medium having executable instructions stored therein. In one embodiment, the computer-executable instructions are used to cause a computer to perform the functions of the above method embodiment.
[0190] The computer-readable storage medium provided in the embodiments of the present application may be a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a mobile hard disk, a CD-ROM, or any other form of computer-readable storage medium known in the art.
[0191] Computer-executable instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium such as a floppy disk, hard disk, or magnetic tape; an optical medium such as a digital video disc (DVD); or a semiconductor medium such as a solid-state drive.
[0192] In the various embodiments of the present application, if there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, including a series of steps or units. The method, system, product or device is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0193] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations thereof may be made without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are intended to be illustrative only of the solutions defined by the appended claims and are to be construed as covering any and all modifications, variations, combinations or equivalents within the scope of the present application.
[0194] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include these modifications and variations.
Claims
1. A floating-point number processing method, characterized in that, The method includes: Obtaining a first floating-point number, where the first floating-point number includes a first sign field, a bit-width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit-width indication field is used to indicate the bit-width D occupied by the first exponent field in the total bit-width N of the first floating-point number; the bit-width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number. Decoding the first floating-point number to obtain the sign, the exponent, and the mantissa.
2. The method according to claim 1, characterized in that The decoding the first floating-point number to obtain the sign, the exponent, and the mantissa includes: Decoding the first floating-point number to obtain a second floating-point number; the second floating-point number includes a second sign field, a second exponent field, and a second mantissa field, the second sign field is used to indicate the sign, the second exponent field is used to indicate the exponent, and the second mantissa field is used to indicate the mantissa.
3. The method according to claim 1 or 2, characterized in that, When the bit-width D is 0, the first exponent field does not exist. If the first mantissa field represents the mantissa, the exponent is a first preset exponent.
4. The method according to any one of claims 1 to 3, characterized in that When the bit-width D is 0, the first exponent field does not exist. If the first mantissa field represents the offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using the offset.
5. The method according to any one of claims 1 to 4, characterized in that When the bit-width D is not 0, the first mantissa field represents the mantissa, and the exponent is the value indicated by the first exponent field.
6. The method according to any one of claims 1 to 5, characterized in that The bit-width DW of the bit-width indication field is negatively correlated with the bit-width D.
7. The method according to any one of claims 1 to 6, characterized in that When the bit-width DW of the bit-width indication field is a preset bit-width and the value indicated by the bit-width indication field is a preset value, the first mantissa field represents the offset.
8. The method according to any one of claims 1 to 7, characterized in that The obtaining the first floating-point number includes: Reading the first floating-point number from a memory, or obtaining the first floating-point number through a communication network.
9. The method according to any one of claims 1 to 8, characterized in that The method further includes: Performing calculations using the sign, the exponent, and the mantissa.
10. A floating-point number processing method, characterized in that, The method includes: Obtaining a sign, an exponent, and a mantissa; Based on the sign, the exponent, and the mantissa, obtaining a first floating-point number; the first floating-point number includes a first sign field, a bit-width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit-width indication field is used to indicate the bit-width D occupied by the first exponent field in the total bit-width N of the first floating-point number; the bit-width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number.
11. The method according to claim 10, characterized in that, When the bit-width D is 0, the first exponent field does not exist. If the first mantissa field represents the mantissa, the exponent is a first preset exponent.
12. The method according to claim 10 or 11, characterized in that, When the bit-width D is 0, the first exponent field does not exist. If the first mantissa field represents the offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using the offset.
13. The method according to any one of claims 10 to 12, characterized in that, When the bit width D is not 0, the first mantissa field represents the mantissa, and the exponent is the value indicated by the first exponent field.
14. The method according to any one of claims 10 to 13, characterized in that The bit width DW of the bit width indication field is negatively correlated with the bit width D.
15. The method according to any one of claims 10 to 14, characterized in that When the bit width DW of the bit width indication field is a preset bit width and the value indicated by the bit width indication field is a preset value, the first mantissa field represents the offset.
16. A floating-point processing device, characterized in that, The device includes: A floating-point number acquisition module, configured to acquire a first floating-point number, where the first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number; A decoding module, configured to decode the first floating-point number to obtain the sign, the exponent, and the mantissa.
17. The device according to claim 16, characterized in that, Specifically, the decoding module is configured to: Decode the first floating-point number to obtain a second floating-point number; the second floating-point number includes a second sign field, a second exponent field, and a second mantissa field, the second sign field is used to indicate the sign, the second exponent field is used to indicate the exponent, and the second mantissa field is used to indicate the mantissa.
18. A floating-point processing device, characterized in that, The device includes: A data acquisition module, configured to acquire a sign, an exponent, and a mantissa; An encoding module, configured to obtain a first floating-point number based on the sign, the exponent, and the mantissa; the first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number.
19. The device according to claim 18, characterized in that, When the bit width D is 0, the first exponent field does not exist. If the first mantissa field represents the offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using the offset.
20. A computing device, characterized in that, Including a processor and a memory; A computer program is stored on the memory; The processor is configured to read the computer program stored in the memory and execute the following steps: Acquire a first floating-point number, where the first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number; Decode the first floating-point number to obtain the sign, the exponent, and the mantissa.
21. The computing device according to claim 20, wherein Specifically, the processor is configured to: Decode the first floating-point number to obtain a second floating-point number; the second floating-point number includes a second sign field, a second exponent field, and a second mantissa field, where the second sign field is used to indicate the sign, the second exponent field is used to indicate the exponent, and the second mantissa field is used to indicate the mantissa.
22. A computing device, characterized in that, Comprising a processor and a memory; A computer program is stored on the memory; The processor is configured to read the computer program stored in the memory and perform the following steps: Obtain a sign, an exponent, and a mantissa; Based on the sign, the exponent, and the mantissa, obtain a first floating-point number; the first floating-point number includes a first sign field, a bit-width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit-width indication field is used to indicate the bit-width D occupied by the first exponent field in the total bit-width N of the first floating-point number; the bit-width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number.
23. The computing device according to claim 22, wherein When the bit-width D is 0, the first exponent field does not exist. If the first mantissa field represents the offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using the offset.
24. A computer-readable storage medium, characterized in that, Stored with computer-executable instructions for causing a computer to execute the method according to any one of claims 1 to 9; or, the method according to any one of claims 10 to 15.
25. A computer program product, characterized in that, Containing computer-executable instructions for causing a computer to execute the method according to any one of claims 1 to 9; or, the method according to any one of claims 10 to 15.
Citation Information
Patent Citations
Floating-point number processing method and device, computing equipment and storage medium
CN120215872A
Floating-point number processing device
CN106990937A
Arithmetic unit, floating-point number calculation method and device, chip and calculation equipment
CN114327360A
Floating-point number recoding and decoding method, system and device
CN115202617A
Floating-point number processing method and related equipment
CN116841500A