Floating-point number processing method and device, computing equipment and storage medium

By introducing a bit width indication field into floating-point numbers, dynamically adjusting the bit width of the order code domain and mantissa domain, the problem of insufficient numerical range and accuracy in AI training and inference is solved, and a wider numerical dynamic range and higher accuracy are achieved.

CN120215872APending Publication Date: 2025-06-27HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311811917.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing floating-point number representations are difficult to meet the requirements of numerical range and accuracy in AI training and inference, resulting in performance impact.

Method used

By introducing a bit width indication field into a floating point number, the bit widths of the order code domain and the mantissa domain are dynamically adjusted, the numerical dynamic range of the floating point number is expanded, and the encoding space of the order code is increased by multiplexing the mantissa domain.

Benefits of technology

Without increasing the total bit width of floating point numbers, the numerical dynamic range and accuracy of floating point numbers are improved, meeting the needs of AI training and inference, and reducing the cost of data storage and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215872A_ABST
    Figure CN120215872A_ABST
Patent Text Reader

Abstract

The invention discloses a floating-point number processing method and device, computing equipment and a storage medium, and relates to the technical field of computers. The computing device obtains a first floating-point number, and decodes the first floating-point number to obtain a symbol, an order code and a mantissa. The first floating-point number may include a first symbol field, a bit width indication field, a first order code field, and a first mantissa field. The bit width indication field is used for indicating the bit width D occupied by the first-order code field in the total bit width N of the floating-point number, and when the total bit width is fixed, the bit width of the first-order code field and the bit width of the first mantissa field can dynamically change along with the numerical value indicated by the bit width indication field. In addition, the bit width indication field can also indicate whether the first mantissa field represents the mantissa or the offset of the order code, and by multiplexing the first mantissa field, the influence on the precision of the floating-point number is small, and the coding space of the order code can be increased, so that the dynamic range of the numerical value represented by the floating-point number is expanded, and the requirements of AI training and reasoning can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a floating-point number processing method, apparatus, computing device, and storage medium. Background Art

[0002] In a computer system, floating point (FP) is an approximate numerical representation method for real numbers, also known as floating-point data representation. Exemplarily, floating-point data representation may include FP8, FP16, and FP32. FP8 represents an 8-bit floating point number, FP16 represents a 16-bit floating point number, and FP32 represents a 32-bit floating point number. Generally, a floating point number includes three fields, namely a sign field, an exponent field, and a mantissa field. In each of the above floating-point data representation methods, the bit widths of the respective fields are fixed. For example, FP16 includes a 1-bit (bit) sign field, a 5-bit exponent field, and a 10-bit mantissa field; FP32 includes a 1-bit sign field, an 8-bit exponent field, and a 23-bit mantissa field; FP8 may include two types, one of which includes a 1-bit sign field, a 5-bit exponent field, and a 2-bit mantissa field; the other includes a 1-bit sign field, a 4-bit exponent field, and a 3-bit mantissa field. Among them, the bit width of the exponent field determines the numerical range that the floating point number can represent, and the bit width of the mantissa field determines the numerical accuracy that the floating point number can represent.

[0003] With the rapid development of artificial intelligence (AI) mixed-precision training and inference, floating point numbers are increasingly applied to mixed-precision training. As the scale of AI network parameters has increased sharply, using floating point numbers with a smaller bit width for AI training and inference can save data storage and data transfer overhead.

[0004] However, the smaller the overall bit width of the floating point number, the smaller the bit width of its exponent field, and the smaller the numerical range that can be represented, which often cannot meet the needs of AI training and inference and affects the performance of AI training and inference. Summary of the Invention

[0005] Embodiments of this application provide a floating-point number processing method, apparatus, computing device, and storage medium, which can expand the dynamic range of the numerical values represented by floating point numbers and meet the needs of AI training and inference.

[0006] In a first aspect, an embodiment of the present application provides a floating-point number processing method, which can be executed by a computing device, or by a chip, a chip system, or a circuit in the computing device. The floating-point number processing method may include: The computing device obtains a first floating-point number, where the first floating-point number may include a first sign field, a bit-width indication field, a first exponent field, and a first mantissa field. The first sign field is used to indicate the sign of the first floating-point number, and the bit-width indication field is used to indicate the bit-width D occupied by the first exponent field in the total bit-width N of the first floating-point number; the bit-width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number. After obtaining the first floating-point number, the computing device may decode the first floating-point number to obtain the above-mentioned sign, exponent, and mantissa.

[0007] In the embodiment of the present application, in addition to including the first sign field, the first exponent field, and the first mantissa field, the first floating-point number may further include a bit-width indication field. The bit-width indication field is used to indicate the bit-width D occupied by the first exponent field in the total bit-width N of the floating-point number. On the premise that the total bit-width is fixed, the bit-width of the first exponent field and the bit-width of the first mantissa field may vary dynamically according to the value indicated by the bit-width indication field, so as to meet the requirements for different numerical ranges and precisions of floating-point numbers in different scenarios. The bit-width indication field may further indicate whether the first mantissa field represents the mantissa or the offset of the exponent. That is to say, the first mantissa field may be used to represent the mantissa or to indicate the offset of the exponent. By multiplexing the first mantissa field, it is possible to ensure a relatively small impact on the precision of the floating-point number and increase the coding space of the exponent, thereby expanding the dynamic range of the values represented by the floating-point number and meeting the needs of AI training and inference.

[0008] In an optional implementation manner, after obtaining the first floating-point number, the computing device may decode the first floating-point number to obtain a second floating-point number. The second floating-point number may include a second sign field, a second exponent field, and a second mantissa field. The second sign field is used to indicate the sign, the second exponent field is used to indicate the exponent, and the second mantissa field is used to indicate the mantissa.

[0009] In an optional implementation manner, when the bit-width D of the first exponent field of the first floating-point number is not 0, the mantissa may be the value indicated by the first mantissa field, and the exponent may be the value indicated by the first exponent field; when the bit-width D of the first exponent field of the first floating-point number is 0, the first exponent field does not exist. At this time, if the first mantissa field represents the mantissa, the exponent may be a first preset exponent. If the first mantissa field represents the offset of the exponent, the exponent is a value obtained by correcting a second preset exponent using the offset, and the mantissa may be a preset mantissa.

[0010] In the above implementation, when the first exponent field does not exist, if the first mantissa field represents a mantissa, the exponent can adopt a first preset exponent; if the first mantissa field represents an exponent offset, the exponent is a value obtained by correcting a second preset exponent using the offset. The value obtained by correcting the second preset exponent using the offset and the value indicated by the first exponent field belong to different numerical ranges. For example, when the total bit width of the first floating-point number is 8, the numerical range to which the value indicated by the first exponent field belongs can be [-15, 15], and the numerical range to which the value obtained by correcting the second preset exponent using the offset belongs can be [-22, -16]. Therefore, the numerical range of exponents that the first floating-point number can represent can reach [-22, 15], which expands the dynamic range of the values represented by the floating-point number compared with the dynamic range of the values that can be represented by the 8-bit floating-point number representation method in the related art.

[0011] In an alternative implementation, the bit width DW of the bit width indication field is negatively correlated with the bit width D of the first exponent field. When the total bit width of the first floating-point number is fixed, it is possible to avoid the jump in the bit width of the first mantissa field caused by the simultaneous increase in the bit widths of the bit width indication field and the first exponent field. The bit width of the mantissa field determines the precision of the floating-point data. Therefore, this implementation can make the precision of the value represented by the first floating-point number change smoothly and prevent the precision of the value represented by the first floating-point number from jumping.

[0012] In an alternative implementation, when the bit width DW of the bit width indication field is a preset bit width and the value indicated by the bit width indication field is a preset value, the first mantissa field represents an offset. For example, in an embodiment, when the total bit width of the first floating-point number is 8 and the bit width DW of the bit width indication field is 4, if the value of the bit width indication field is "0000", it indicates that the first mantissa field represents an offset; or, when the total bit width of the first floating-point number is 8 and the bit width DW of the bit width indication field is 4, if the value of the bit width indication field is "0001", it indicates that the first mantissa field represents an offset.

[0013] In an alternative implementation, the computing device can read the first floating-point number from the memory, or obtain the first floating-point number through a communication network. After decoding the first floating-point number to obtain the sign, exponent, and mantissa, the obtained sign, exponent, and mantissa can be used for calculation.

[0014] In the related art, when the total bit width of a floating-point number is determined, the bit widths of its exponent field and mantissa field are always fixed. When a larger precision or a larger numerical range is required during the calculation process, only a data format with a larger total bit width can be selected. However, an increase in the total bit width means that the bit widths of both the exponent field and the mantissa field increase simultaneously. This easily leads to waste of the increased bit width of the exponent field when only a larger precision is needed, and waste of the increased bit width of the mantissa field when only a larger numerical range is needed, thus occupying unnecessary storage space and increasing the overhead of data storage and data transfer of the floating-point number. In the embodiments of the present application, the first floating-point number can be used for data storage or data transfer. Since the bit widths of the first exponent field and the first mantissa field of the first floating-point number can dynamically change according to the value indicated by the bit width indication field, various different requirements for the numerical range and numerical precision of the floating-point number in different scenarios can be flexibly met without additionally increasing the total bit width of the floating-point number, that is, without additionally increasing the cost of data storage or data transfer.

[0015] In a second aspect, embodiments of the present application provide a floating-point number processing method. This method can be executed by a computing device, or by a chip, a chip system, or a circuit in the computing device. The floating-point number processing method may include: The computing device obtains a sign, an exponent, and a mantissa, and based on the sign, the exponent, and the mantissa, obtains a first floating-point number. The first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field. The first sign field is used to indicate the sign of the first floating-point number, and the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number. The bit width indication field is further used to indicate that: the first mantissa field represents the mantissa of the first floating-point number, or, the offset of the exponent of the first floating-point number.

[0016] In an optional implementation manner, when the bit width D of the first exponent field is 0, the first exponent field does not exist. If the first mantissa field represents the mantissa, the exponent is a first preset exponent.

[0017] In an optional implementation manner, when the bit width D of the first exponent field is 0, the first exponent field does not exist. If the first mantissa field represents the offset of the exponent, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using this offset.

[0018] In an optional implementation manner, when the bit width D of the first exponent field is not 0, the first mantissa field represents the mantissa, and the exponent is the value indicated by the first exponent field.

[0019] In an optional implementation manner, the bit width DW of the bit width indication field is negatively correlated with the bit width D.

[0020] In an alternative implementation, when the bit width DW of the bit width indication field is a preset bit width and the value indicated by the bit width indication field is a preset value, the first mantissa field represents the offset of the exponent.

[0021] In a third aspect, an embodiment of the present application provides a floating-point processing device, which can be applied to a computing device. The floating-point processing device may include:

[0022] A floating-point acquisition module, configured to acquire a first floating point number; the first floating point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating point number, the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating point number; the bit width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating point number, or the offset of the exponent of the first floating point number;

[0023] A decoding module, configured to decode the first floating point number to obtain a sign, an exponent, and a mantissa.

[0024] In an alternative implementation, the decoding module may specifically be configured to:

[0025] Decode the first floating point number to obtain a second floating point number; the second floating point number includes a second sign field, a second exponent field, and a second mantissa field, the second sign field is used to indicate the sign, the second exponent field is used to indicate the exponent, and the second mantissa field is used to indicate the mantissa.

[0026] In an alternative implementation, when the bit width D is 0, the first exponent field does not exist. If the first mantissa field represents the mantissa, the exponent is a first preset exponent.

[0027] In an alternative implementation, when the bit width D is 0, the first exponent field does not exist. If the first mantissa field represents the offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using the offset.

[0028] In an alternative implementation, when the bit width D is not 0, the first mantissa field represents the mantissa, and the exponent is the value indicated by the first exponent field.

[0029] In an alternative implementation, the bit width DW of the bit width indication field is negatively correlated with the bit width D.

[0030] In an alternative implementation, when the bit width DW of the bit width indication field is a preset bit width and the value indicated by the bit width indication field is a preset value, the first mantissa field represents the offset.

[0031] In an alternative implementation, the floating-point acquisition module may specifically be configured to:

[0032] Read a first floating-point number from a memory, or obtain a first floating-point number through a communication network.

[0033] In an alternative implementation, the above floating-point processing device may further include a calculation module, and the calculation module may be used for: performing calculations using a sign, an exponent, and a mantissa.

[0034] In a fourth aspect, an embodiment of the present application provides a floating-point processing device, which can be applied to a computing device. The floating-point processing device may include:

[0035] A data acquisition module, configured to acquire a sign, an exponent, and a mantissa;

[0036] An encoding module, configured to obtain a first floating-point number based on the sign, the exponent, and the mantissa; the first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field is further used to indicate that: the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number.

[0037] In an alternative implementation, when the bit width D is 0, the first exponent field does not exist. If the first mantissa field represents the mantissa, the exponent is a first preset exponent.

[0038] In an alternative implementation, when the bit width D is 0, the first exponent field does not exist. If the first mantissa field represents the offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using the offset.

[0039] In an alternative implementation, when the bit width D is not 0, the first mantissa field represents the mantissa, and the exponent is the value indicated by the first exponent field.

[0040] In an alternative implementation, the bit width DW of the bit width indication field is negatively correlated with the bit width D.

[0041] In an alternative implementation, when the bit width DW of the bit width indication field is a preset bit width and the value indicated by the bit width indication field is a preset value, the first mantissa field represents the offset.

[0042] In a fifth aspect, an embodiment of the present application provides a computing device, which includes a processor and a memory; a computer program is stored on the memory; the processor is configured to read the computer program stored in the memory and execute any one of the floating-point processing methods provided in the first aspect or the second aspect above.

[0043] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute any one of the floating-point processing methods provided in the first aspect or the second aspect above.

[0044] In a seventh aspect, an embodiment of the present application provides a computer program product including computer-executable instructions for causing a computer to execute any one of the floating-point processing methods provided in the first aspect or the second aspect above.

[0045] The technical effects that can be achieved by any one of the second aspect to the seventh aspect above can be referred to the description of the beneficial effects in the first aspect above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic diagram of an application scenario of an embodiment of the present application;

[0047] Figure 2 It is a schematic diagram of the structure of a computing device provided by an embodiment of the present application;

[0048] Figure 3 It is a schematic diagram of the significant bit-exponent distribution of a floating-point number provided by an embodiment of the present application;

[0049] Figure 4 It is a flowchart of a floating-point processing method provided by an embodiment of the present application;

[0050] Figure 5 It is a schematic diagram of a decoder decoding a first floating-point number provided by an embodiment of the present application;

[0051] Figure 6 It is a flowchart of another floating-point processing method provided by an embodiment of the present application;

[0052] Figure 7 It is a schematic diagram of an encoder encoding a first floating-point number provided by an embodiment of the present application;

[0053] Figure 8 It is a flowchart of another floating-point processing method provided by an embodiment of the present application;

[0054] Figure 9 It is a training effect diagram of AI training using floating-point numbers in related technologies;

[0055] Figure 10 It is a training effect diagram of AI training using the floating-point numbers of an embodiment of the present application;

[0056] Figure 11 It is a schematic diagram of the structure of a floating-point processing device provided by an embodiment of the present application;

[0057] Figure 12 It is a schematic structural diagram of another floating-point processing device provided by an embodiment of the present application;

[0058] Figure 13 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0059] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The terms used in the implementation manners part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0060] Before introducing the specific solutions provided by the embodiments of the present application, some terms in the present application are explained to facilitate the understanding of those skilled in the art, and the terms in the present application are not limited.

[0061] (1) Integer encoding: It refers to encoding an integer using a fixed-length binary. For example, using 3-bit binary to encode 0 to 7, using 4-bit binary to encode 0 to 15, and so on.

[0062] (2) Prefix code encoding: It can also be called prefix encoding. If in a coding method, any code is not the prefix (leftmost substring) of any other code, then this coding method can be called prefix encoding. For example, variable-length encodings: 1, 01, 001, 0101; or 00, 01, 10, 1100, 1101; and fixed-length encodings: 00, 01, 10, 11, etc. all belong to prefix code encoding. Prefix encoding can ensure that there is no ambiguity when decoding a compressed file and ensure correct decoding.

[0063] Prefix code encoding can include conventional prefix code encoding and unconventional prefix code encoding. Conventional prefix code encoding uses a shorter bit width to encode smaller data and a longer bit width to encode larger data. Unconventional prefix code encoding is the opposite, which uses a shorter bit width to encode larger data and a longer bit width to encode smaller data. In some embodiments of the present application, the Dot field in the floating-point number can be encoded by unconventional prefix code encoding. For specific details, please refer to the description in the following embodiments.

[0064] In the embodiments of the present application, "a plurality of" means two or more. In view of this, in the embodiments of the present application, "a plurality of" can also be understood as "at least two". "At least one" can be understood as one or more, for example, understood as one, two or more. For example, including at least one means including one, two or more, and it does not limit which ones are included. For example, including at least one of A, B, and C, then what can be included are A, B, C, A and B, A and C, B and C, or A, B, and C. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / ", unless otherwise specified, generally represents an "or" relationship between the front and back associated objects.

[0065] Unless otherwise stated, the ordinal numbers such as "first" and "second" mentioned in the embodiments of the present application are used to distinguish multiple objects and are not used to limit the order, time sequence, priority or importance of multiple objects.

[0066] The floating-point data representation method is a scientific calculation method. In the Institute of Electrical and Electronics Engineers (IEEE) 754 binary floating-point standard, a floating-point number can include three fields, namely the sign field, the exponent field, and the mantissa field. Among them, the value of the exponent field represents an integer exponent of a certain base, and the value of the mantissa field multiplied by an integer exponent of a certain base can obtain a data; the sign field is used to indicate the positive or negative of the data. For example, taking the binary data 100.101 (i.e., 4.625 in decimal) as an example, this data can be expressed as (-1) 0 ×2 2 ×1.00101. Among them, the exponent 0 of the base -1 is the value of the sign field in the floating-point number and is used to represent the positive or negative of the data. The exponent 2 of the base 2 is the value of the exponent field in the floating-point number and represents the number of decimal point shifts. For example, from 100.101 to 1.00101, the number of decimal point shifts is 2. 1.00101 is the mantissa. Since there must be a 1 before the decimal point, the value actually stored in the mantissa field of the floating-point number can only include the data after the decimal point, that is, 00101, so that a binary bit can be saved to store more mantissas.

[0067] Exemplarily, common floating-point data representation methods may include FP16 and FP32. FP16 represents a 16-bit floating-point number, and FP32 represents a 32-bit floating-point number. In each floating-point data representation method, the bit widths of each field are fixed. For example, FP16 includes a 1-bit sign field, a 5-bit exponent field, and a 10-bit mantissa field; FP32 includes a 1-bit sign field, an 8-bit exponent field, and a 23-bit mantissa field. Among them, the bit width of the exponent field determines the numerical range that the floating-point number can represent, and the bit width of the mantissa field determines the numerical precision that the floating-point number can represent.

[0068] As the scale of AI network parameters has increased sharply, using floating-point numbers with a smaller bit width for AI training and inference can save data storage and data transfer overhead. Subsequently, FP8 floating-point numbers emerged. There are two types of FP8 floating-point numbers, as shown in Table 1 for details.

[0069] Table 1

[0070] Sign Exponent mantissa E5M2 1bit 5bits 2bits E4M3 1bit 4bits 3bits

[0071] As shown in Table 1, FP8 floating-point numbers can include two types, one is E5M2, and the other is E4M3. The floating-point number of E5M2 includes a 1-bit sign field, a 5-bit exponent field, and a 2-bit mantissa field; the floating-point number of E4M3 includes a 1-bit sign field, a 4-bit exponent field, and a 3-bit mantissa field. The bit width of the exponent field in the floating-point number determines the exponent range that the floating-point number can represent. The exponent range that the floating-point number of E5M2 can represent is [-16, 15], while the exponent range that the floating-point number of E4M3 can represent is smaller, which is [-9, 8].

[0072] Since the dynamic ranges of the numerical values that the above two types of FP8 floating-point numbers can represent are both small, they often cannot meet the needs of AI training and inference, affecting the performance of AI training and inference. When using traditional floating-point data representation methods, only data formats with a larger total bit width can be selected. For example, select FP16 floating-point numbers or FP32 floating-point numbers. This is likely to cause waste in the bit width of the exponent field or the mantissa field in floating-point numbers with a larger total bit width, making the floating-point numbers occupy unnecessary storage space and greatly increasing the data storage and data transfer overhead of the floating-point numbers.

[0073] Based on this, an embodiment of the present application provides a floating-point number processing method. The floating-point number provided by the embodiment of the present application may include, in addition to the first sign field, a bit-width indication field (i.e., the Dot field), a first exponent field, and a first mantissa field. Among them, the Dot field is used to indicate the bit-width D occupied by the first exponent field in the total bit-width N of the floating-point number. When the bit-width D is non-zero, the first exponent field is used to indicate the exponent of the floating-point number, and the first mantissa field is used to indicate the mantissa of the floating-point number. Under the same total bit-width, the bit-widths of the first exponent field and the first mantissa field can dynamically change with the value of the Dot field to meet the requirements for different numerical ranges and precisions of floating-point numbers in different scenarios. When the bit-width D is 0, the Dot field can also be used to indicate whether the first mantissa field represents the mantissa of the floating-point number or the offset of the exponent of the floating-point number. By reusing the first mantissa field, it is possible to ensure a relatively small impact on the precision of the floating-point number and increase the encoding space of the exponent, thereby expanding the dynamic range of the values represented by the floating-point number. For example, taking a floating-point number with a total bit-width N of 8 as an example, the dynamic range of the values represented by the above-mentioned E4M3 floating-point number is [-9, 8], the dynamic range of the values that the E5M2 floating-point number can represent is [-16, 15], while the dynamic range of the values that the HiF8_DML floating-point number provided by the embodiment of the present application can represent is [-22, 15], greatly expanding the dynamic range of the values represented by the floating-point number and meeting the needs of AI training and inference.

[0074] The floating-point number processing method provided by the embodiment of the present application can be applied to computing devices and is widely applicable to various industries such as scientific research, engineering, finance, aerospace, and healthcare. These industries have a large amount of data that needs to be stored, calculated, and transmitted every day. Figure 1 Exemplarily, a schematic diagram of an application scenario provided by the embodiment of the present application is shown, as Figure 1 shown. In this application scenario, computing device 100, computing device 200, and computing device 300 can establish a communication connection through a network. Among them, the network can be a wired network or a wireless network. For example, wireless fidelity (WiFi), Bluetooth, mobile network, etc. Computing device 100, computing device 200, and computing device 300 can be any electronic devices such as a computer, a server, a smart wearable device, a smart home, a tablet computer, a laptop computer, a vehicle-mounted terminal, a smart phone, etc. When the computing device is a server, it can be a single server, a server cluster composed of multiple servers, or a cloud computing service center, etc. A communication connection is established between computing device 100, computing device 200, and computing device 300, and data storage, calculation, and transmission based on floating-point numbers can be performed.

[0075] It should be noted that in actual application scenarios, the system may include more than 3 computing devices or less than 3 computing devices, and the present application does not limit this.

[0076] The internal structures of computing device 100, computing device 200, and computing device 300 may be the same. Taking computing device 100 as an example for illustration, as Figure 2 shown. Inside computing device 100, there may be a processor 110 and a memory 120. Exemplarily, the processor 110 may be, but is not limited to, a central processing unit (CPU), a high performance computing (HPC) service acceleration chip, a graphics processing unit (GPU), or a neural network processing unit (NPU) in the field of AI, etc. The processor 110 may include a decoder 111, an encoder 112, and a computing unit 113, and the number of computing units 113 may be one or multiple.

[0077] Exemplarily, the processor 110 may include some or all of the following computing units: a scalar computing unit, a vector computing unit, a matrix computing unit, or a tensor computing unit.

[0078] The different computing units will be introduced separately below.

[0079] I. Scalar computing unit: A scalar, also known as a pure quantity, has only magnitude and no direction. The circuit for scalar computing is called a scalar computing unit. Scalar computing is mostly used for general computing. In the embodiments of the present application, an arithmetic logic unit ALU (Arithmetic Logic Unit) based on the HiF8_DML data format may be embedded in the execution unit (EXU) part of the CPU multi-stage pipeline or the scalar computing part of other processors with similar functions.

[0080] II. Vector calculation unit: A vector, also known as a vector quantity, usually refers to a one-dimensional array with a length greater than 1. A calculation unit with a certain degree of parallelism specially designed for vector calculation is called a vector calculation unit, such as a single instruction multiple data (SIMD) processor. Vector calculation units are mostly used in fields such as HPC high-performance computing and AI machine learning, including the solution of mathematical problems such as linear programming, Fourier transform, filtering calculation, as well as linear algebra, partial differential equations, and integrals. In the embodiments of the present application, an arithmetic execution unit (vector unit) based on the HiF8_DML data format can be embedded in the vector calculation acceleration unit or the vector processor.

[0081] III. Matrix calculation unit: A matrix is a two-dimensional array arranged in a rectangular array. A calculation unit with a corresponding degree of parallelism specially designed for matrix calculation is called a matrix calculation unit, such as a systolic array processor. Matrix calculation units are mostly used in matrix calculations in fields such as HPC high-performance computing and AI machine learning, including matrix multiplication, matrix inversion, matrix decomposition, etc. In the embodiments of the present application, a matrix unit (matrix unit) based on the HiF8_DML data format can be embedded in the matrix calculation acceleration unit.

[0082] IV. Tensor calculation unit: A tensor is a multi-dimensional array with a dimension exceeding 2, and a common one is a three-dimensional array. A calculation unit with a corresponding degree of parallelism specially designed for tensor calculation is called a tensor product calculation unit. Tensor calculation units are mostly used in the field of AI machine learning, such as convolution operations. In the embodiments of the present application, a tensor unit (tensor unit) based on the HiF8_DML data format can be embedded in the tensor calculation acceleration unit.

[0083] Exemplarily, when the computing device 100 performs general computing, high-performance computing, or AI training, a large amount of floating-point data is required. At this time, the computing device 100 can use the decoder 111 to decode the obtained floating-point numbers based on a floating-point processing method provided by an embodiment of the present application to obtain the decoded data. Exemplarily, the floating-point numbers can be read from the local memory 120 of the computing device 100, or obtained from the computing device 200 or the computing device 300 through the network, or obtained from other devices in the network. After the decoder 111 decodes the obtained floating-point numbers to obtain the decoded data, the decoded data can be transmitted to the computing unit 113, and the corresponding calculations are completed by the computing unit 113. The computing unit 113 transmits the calculation result to the encoder 112, and the encoder 112 re-encodes the calculation result into a floating-point number, which can be used for data storage and data transfer. For example, the processor 110 can save the floating-point number encoded by the encoder 112 to the memory 120.

[0084] Optionally, Figure 1 For the specific structures and functions of the computing device 200 and the computing device 300, reference can specifically be made to Figure 2 the computing device 100 shown. In some possible embodiments, the computing device 100, the computing device 200, and the computing device 300 may include more or fewer components than Figure 2 those shown, and the embodiments of the present application do not make specific limitations thereon.

[0085] For easier understanding, the data format of the floating-point numbers provided by the embodiments of the present application will be introduced first below. Table 2 shows the data format of the floating-point numbers provided by the embodiments of the present application.

[0086] Table 2

[0087] Region First Sign Field Dot Field First Exponent Field First Mantissa Field Width / bit 1 DW D N-1-DW-D

[0088] As shown in Table 2, the floating-point numbers provided by the embodiments of the present application may include a first sign field, a Dot field, a first exponent field, and a first mantissa field. Among them, the Dot field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the floating-point number. Among them, N is an integer greater than 1, and D is an integer greater than or equal to 0. The following embodiments will be described by taking binary and the total bit width N of the floating-point number being 8 as an example. First, each field of the floating-point number will be introduced in detail in combination with Table 2 and Table 3.

[0089] Table 3

[0090]

[0091] I. First Symbol Field: The first symbol field can also be referred to as the sign bit. As shown in Table 3, the symbol field is located before the Dot field and occupies 1 bit in the total bit width N of the floating-point number, used to represent the positive or negative of the data. By default, 0 represents positive and 1 represents negative. It can also be set that 0 represents negative and 1 represents positive according to actual needs. This application does not make a limitation on this.

[0092] II. Dot Field: It occupies DW bits in the total bit width N of the floating-point number, and the value of DW ranges from 2 to 4. The value (or encoded value) represented by the Dot field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the floating-point number, that is, the value of the Dot field is D. When D is non-zero, the first exponent field is used to indicate the exponent of the floating-point number; the first mantissa field is used to indicate the mantissa of the floating-point number. When the bit width D is 0, the bit width indication field can also be used to indicate whether the first mantissa field represents the mantissa of the floating-point number or the offset of the exponent of the floating-point number.

[0093] Optionally, the encoding method of the Dot field can adopt prefix code encoding. Prefix code encoding can include regular prefix code encoding and non-regular prefix code encoding. Regular prefix code encoding uses a shorter bit width to represent a smaller value and a longer bit width to represent a larger value. Non-regular prefix code encoding is the opposite, using a shorter bit width to represent a larger value and a longer bit width to represent a smaller value.

[0094] In some embodiments, the non-regular prefix code encoding method can be adopted to encode the Dot field, that is, using a longer bit width to represent a smaller value and a shorter bit width to represent a larger value. Exemplarily, when the bit width DW is the first bit width value, the bit width DW1 occupied by the Dot field encodes D1 values, and the bit width D of the first value field belongs to any one of the D1 values; when the bit width DW is the second bit width value, the bit width DW2 occupied by the Dot field encodes D2 values, and the bit width D of the first value field belongs to any one of the D2 values. Among them, DW1 is less than DW2, and the smallest value among the D1 values is greater than the largest value among the D2 values. D1 and D2 are integers greater than or equal to 0.

[0095] Exemplarily, the Dot field occupies 2 - 4 bits in the total bit width N of the floating-point number, that is, the bit width of the Dot field (width) is 2 - 4. The specific encoding method can be as shown in Table 4.

[0096] Table 4

[0097]

[0098] As shown in Table 4, when the bit width of the Dot field is 2, a 2-bit width can be used to encode and represent any one of the three numerical values 2, 3, and 4. For example, the coding "11" can represent the numerical value "4", the coding "10" can represent the numerical value "3", and the coding "01" can represent the numerical value "2". When the bit width of the Dot field is 3, a 3-bit width can be used to encode and represent the numerical value 1. For example, the coding "001" can represent the numerical value "1". The numerical values 1, 2, 3, and 4 are used to indicate the bit width D of the first exponent field, and the first exponent field is used to characterize the exponent of the floating-point number. When the bit width of the Dot field is 4, a 4-bit width can be used to encode and represent the numerical value 0. At this time, the first exponent field does not exist.

[0099] The bit width indication field uses an unconventional prefix coding. The bit width DW of the bit width indication field is negatively correlated with the value indicated by the bit width indication field. That is to say, the bit width DW of the bit width indication field is negatively correlated with the bit width D of the first exponent field. When the total bit width of the floating-point number is fixed, it is possible to avoid the jump of the bit width of the first mantissa field caused by the simultaneous increase of the bit widths of the bit width indication field and the first exponent field. The bit width of the mantissa field determines the precision of the floating-point data. Therefore, it is possible to make the precision of the numerical value represented by the first floating-point number change smoothly and prevent the precision of the numerical value represented by the first floating-point number from jumping; especially near the exponent center, it is possible to smooth the jump of the bit width of the mantissa field, that is, to smooth the precision jump of the numerical value near the exponent center.

[0100] When the bit width of the Dot field is 4, the Dot field is also used to indicate whether the first mantissa field represents the mantissa of a floating-point number or the offset of the exponent of the floating-point number. For example, in one embodiment, the encoding "0001" can represent the value "0". At this time, the bit width D of the first exponent field is 0 bits, and the first mantissa field represents the mantissa of the floating-point number. The encoding "0000" can represent that the floating-point number is DML data. DML data refers to subnormal or denormal values. At this time, the first mantissa field represents the offset of the exponent of the floating-point number. That is to say, when the bit width of the Dot field is 4, if the Dot field is the preset value "0000", then the Dot field indicates that the floating-point number is DML data, or in other words, indicates that the first mantissa field represents the offset of the exponent of the floating-point number. In another embodiment, it is also possible to use the encoding "0000" to represent the value "0". At this time, the bit width D of the first exponent field is 0 bits, and the first mantissa field represents the mantissa of the floating-point number. The encoding "0001" can represent that the floating-point number is DML data, that is, when the Dot field is the preset value "0001", then the Dot field indicates that the floating-point number is DML data, or in other words, indicates that the first mantissa field represents the offset of the exponent of the floating-point number. From the above description, the bit width indication field can be described as Dot = [0,4] & DML, which means that: except for the DML flag, the value D represented by the encoding of the bit width indication field ranges from 0 to 4. The bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the floating-point number. That is to say, the bit width D of the first exponent field ranges from 0 to 4 bits.

[0101] III. First exponent field: used to represent the exponent of the floating-point number, occupying a bit width of D in the total bit width N of the floating-point number, and the value range of D is from 0 to 4 bits.

[0102] Exemplarily, the first exponent field Es is used to represent the exponent of the floating-point number. Assume that the value range of the exponent of the floating-point number is E, and the value Ev represented by the first exponent field Es belongs to this value range E. This value range E can be determined by the following formula 1:

[0103] E = (-1) Se × [2 D-1 , (2 D - 1)] Formula 1

[0104] Where Se is the sign bit of the value Ev, which can also be called the sign bit of the exponent of the floating-point number. Se occupies 1 bit in the bit width D of the exponent field and is used to represent the positive or negative of the value of the first exponent field. In one embodiment, Se being 0 represents that the value Ev is positive, and Se being 1 represents that the value Ev is negative. In other embodiments, it is also possible to use Se being 0 to represent that the value Ev is positive and Se being 1 to represent that the value Ev is negative. The present application does not make specific limitations on this.

[0105] As shown in combination with Formula 1 and Table 3, when the bit width D of the first exponent field is 0, the value Ev represented by the first exponent field is 0, that is, the value range E of the exponent of the floating-point number is 0. When the bit width D of the first exponent field is 1, the value Ev represented by the first exponent field can be 1 or -1, that is, the value range E of the exponent of the floating-point number is ±1. When the bit width D of the first exponent field is 2 to 4, the value Ev represented by the first exponent field can be any value in [-2, -15] or [2, 15], that is, the value range E of the exponent of the floating-point number is ±[2, 15]. It can be seen that the value range E of the exponent of the floating-point number that the first exponent field can represent is [-15, 15].

[0106] In an alternative embodiment, when the bit width D is 0 and it is not DML, the exponent of the floating-point number is represented as 0. When the bit width D is non-zero, that is, when the bit width D is any value from 1 to 4, the first exponent field can be encoded using signed magnitude encoding, which can also be referred to as the encoding method where the sign bit Se follows the original code. This representation method adds a sign bit in front of the numerical value, that is, the highest bit of the original code is the sign bit, which is used to represent the positive or negative of the numerical value. Among them, the sign bit being 0 represents a positive number, and the sign bit being 1 represents a negative number. The remaining bits in the original code except the sign bit are used to represent the magnitude of the numerical value, that is, the magnitude of the original code. For example, the original code 1001 represents -1, and 0011 represents +3.

[0107] In the embodiment of the present application, when the bit width D is non-zero, the first exponent field in the floating-point number can adopt the encoding method where the exponent sign bit Se follows the original code magnitude, and can be represented as Es: {Se + Mag[2:end]}, where the exponent sign bit Se is the sign bit extracted from the initial original code, which is used to represent the positive or negative of the value of the first exponent field; Mag is used to represent the magnitude of the value of the first exponent field, end = D - 1, and Mag[2:end] is used to represent the value of the 2nd to the (D - 1)th bit of the first exponent field. For different bit widths D, the highest bit b1 of the magnitude Mag of the value of the first exponent field is always 1 (i.e., 1’b1). Therefore, the highest bit 1’b1 can not occupy the bit width during encoding, that is, the highest bit 1’b1 is hidden and not actually stored. During subsequent decoding, the highest bit 1’b1 can be directly supplemented to obtain the encoded value Ei of the exponent field in the normalized floating-point number: {Se + 1’b1 + Mag[2:end]}. Among them, the normalized floating-point number refers to the floating-point number that conforms to the IEEE754 binary floating-point standard, and the normalized floating-point number can be used for floating-point calculations. Since the highest bit of the encoded value Ei of the exponent field does not need to occupy the bit width in the first exponent field and is not actually stored, the storage space can be greatly saved, and the cost of data storage and data transfer can be reduced.

[0108] Exemplarily, as shown in Table 5, when the Dot field is the DML flag, it indicates that the floating-point number is DML data. At this time, the bit width of the first mantissa field is 3, that is, the first mantissa field uses 3 bits to represent the offset of the exponent of the floating-point number. When the Dot field indicates that the bit width D of the first exponent field is 0, the encoding Es of the first exponent field in the floating-point number is None, and the decoded value of the exponent field of the normalized floating-point number Ei is 0, indicating that the exponent value Ev of the floating-point number is 0. At this time, the bit width of the Dot field is 4, and the bit width of the first mantissa field is 3, that is, the first mantissa field uses 3 bits to represent the mantissa of the floating-point number.

[0109] When the Dot field indicates that the bit width D of the first exponent field is 1, the encoding Es of the first exponent field in the floating-point number is {Se}, and the decoded value of this first exponent field, that is, the value Ei of the exponent field of the normalized floating-point number is {Se, 1}, where 1 is the value of the highest bit in Mag. As described above, the highest bit in Mag is always 1. When the bit width D = 1, the exponent Ev of the floating-point number represented by the first exponent field can be 1 or -1, that is, the numerical range of the exponent Ev of the floating-point number can be ±1. At this time, the bit width of the Dot field is 3, and the bit width of the first mantissa field is 3, that is, the first mantissa field uses 3 bits to represent the mantissa of the floating-point number.

[0110] When the Dot field indicates that the bit width D of the first exponent field is 2, the encoding Es of the first exponent field in the floating-point number is {Se, [2]}, where [2] represents the value of the second bit in Mag. The decoded value of this first exponent field, that is, the value Ei of the exponent field of the normalized floating-point number is {Se, 1, [2]}, where 1 is the value of the highest bit in Mag. When the bit width D = 2, the exponent Ev of the floating-point number represented by the first exponent field can be any value in [-2, -3] or [2, 3], that is, the numerical range of the exponent Ev of the floating-point number can be ±[2, 3]. At this time, the bit width of the Dot field is 2, and the bit width of the first mantissa field is 3, that is, the first mantissa field uses 3 bits to represent the mantissa of the floating-point number.

[0111] When the Dot field indicates that the bit width D of the first exponent field is 3, the encoding Es of the first exponent field in the floating-point number is {Se, [2:3]}, where [2:3] represents the values of the second and third bits in Mag. The decoded value of this first exponent field, that is, the value Ei of the exponent field of the normalized floating-point number is {Se, 1, [2:3]}, where 1 is the value of the highest bit in Mag. When the bit width D = 3, the exponent Ev of the floating-point number represented by the first exponent field can be any value in [-4, -7] or [4, 7], that is, the numerical range of the exponent Ev of the floating-point number can be ±[4, 7]. At this time, the bit width of the Dot field is 2, and the bit width of the first mantissa field is 2, that is, the first mantissa field uses 2 bits to represent the mantissa of the floating-point number.

[0112] When the Dot field indicates that the bit width D of the first exponent field is 4, the encoding Es of the first exponent field in the floating point number is {Se, [2:4]}, where [2:4] represents the values of the 2nd, 3rd, and 4th bits in Mag. The decoded value of this first exponent field, that is, the value Ei of the exponent field of the normalized floating point number, is {Se, 1, [2:4]}, where 1 is the value of the highest bit in Mag. When the bit width D = 4, the exponent Ev of the floating point number represented by the first exponent field can be any value in [-8, -15] or [8, 15], that is, the numerical range of the exponent Ev of the floating point number can be ±[8, 15]. At this time, the bit width of the Dot field is 2, and the bit width of the first mantissa field is 1, that is, the first mantissa field uses 1 bit to represent the mantissa of the floating point number.

[0113] It can be seen from this that when D is greater than 1, the encoding Es of the first exponent field in the floating point number is {Se + Mag[2:end]}, Mag[2:end] includes the remaining bits in Mag except the highest bit 1’b1, and the bit width occupied by Mag[2:end] in the first exponent field is D - 1.

[0114] Table 5

[0115] D DML 0 1 2 3 4 Es / None Se Se, [2] Se, [2:3] Se, [2:4] Ei / 0 Se, 1 Se, 1, [2] Se, 1, [2:3] Se, 1, [2:4] Ev / 0 ±1 ±[2,3] ±[4,7] ±[8,15] Bit Width of the First Mantissa Field 3 3 3 3 2 1

[0116] IV. First mantissa field: used to represent the mantissa of the floating point number or the offset of the exponent of the floating point number. The bit width occupied by the first mantissa field in the floating point number is (N - 1 - DW - D) bits. Exemplarily, the Dot field is used to indicate whether the first mantissa field represents the mantissa of the floating point number or the offset of the exponent of the floating point number. In some embodiments, when the bit width DW of the Dot field is 4 and the value of the Dot field is "0001", it is used to indicate that the first mantissa field represents the mantissa of the floating point number; when the bit width DW of the Dot field is 4 and the value of the Dot field is "0000", it is used to indicate that the first mantissa field represents the offset of the exponent of the floating point number. In some other embodiments, the opposite setting can also be adopted. When the bit width DW of the Dot field is 4 and the value of the Dot field is "0000", it is used to indicate that the first mantissa field represents the mantissa of the floating point number; when the bit width DW of the Dot field is 4 and the value of the Dot field is "0001", it is used to indicate that the first mantissa field represents the offset of the exponent of the floating point number.

[0117] When the first mantissa field represents the mantissa of the floating point number, the first mantissa field is used to store the value after the decimal point. For example, if storing the decimal part of 1.xxx, assuming the integer part 1’b1 is hidden. For example, if the encoding saved in the first mantissa field is 10011, the decoded value M of the first mantissa field is 0.10011, and the value represented by the first mantissa field is 1.M, that is, 1.10011.

[0118] As shown in Table 3, when the first mantissa field represents the offset of the exponent of the floating-point number, the bit width of the first mantissa field is 3. That is, the first mantissa field uses 3 bits to encode the integer value M from 0 to 7, which is used to represent the exponent value of M - 23. Here, M can be called the offset of the exponent. Thus, it can be seen that the first mantissa field can supplementarily represent the numerical range [-23, -16] of the exponent. Among them, when all 3 bits of the first mantissa field are 0, that is, M = 0 and M - 23 = -23, the floating-point number is used to represent a special value (which will be described in detail below). Therefore, the numerical range of the exponent that the first mantissa field can represent is [-22, -16]. Combined with the above first exponent field (as shown in Table 5), the HiF8_DML floating-point number provided by the embodiment of the present application can represent the exponent range of the floating-point number as [-22, 15]. When the exponent of the floating-point number represented by HiF8_DML is 15, the value X of the represented floating-point number is the largest, which can be approximated as X max-pos = 2 15 = 32768. When the exponent of the floating-point number represented by HiF8_DML is -22, the value X of the represented floating-point number is the smallest, which can be approximated as X min-pos = 2 -22 ≈ 2.38×10 -7 .

[0119] When the first mantissa field represents the offset of the exponent of the floating-point number, the corresponding floating-point number is DML data, and the value X of this floating-point number can be expressed as:

[0120] X = (-1) s × 2 M-23

[0121] where S is the value of the sign field of the floating-point number, and M is the value represented by the first mantissa field of the floating-point number. At this time, the mantissa of the floating-point number adopts the default value 1.

[0122] When the first mantissa field represents the mantissa of the floating-point number, the corresponding floating-point number is normal data, and the value X of this floating-point number can be expressed as:

[0123] X = (-1) s × 2 Ev × 1.M

[0124] where S is the value of the sign field of the floating-point number, Ev is the value represented by the first exponent field of the floating-point number, that is, the exponent of the floating-point number. M is the value represented by the first mantissa field of the floating-point number, and 1.M is the mantissa of the floating-point number. The 1 in 1.M is also regarded as a significant bit. It can be seen that the number of significant bits of the mantissa of the floating-point number is the bit width of the first mantissa field + 1.

[0125] In summary, the value X of the floating-point number can be expressed as:

[0126]

[0127] The HiF8_DML floating-point number provided by the embodiment of this application can also support four special values, namely, zero (Zero) regardless of positive or negative, not a number (NaN), positive infinity (+inf), and negative infinity (-inf). The four special values can be represented by 4 boundary values that HiF8_DML can represent.

[0128] (1) Zero: When the value of the sign field S of the floating-point number is 0 and the Dot field is 0000, that is, the Dot field indicates that the floating-point number is DML data, that is, the first mantissa field represents the offset of the exponent of the floating-point number, and when the 3 bits of the first mantissa field are all 0, the corresponding floating-point number can represent ±0, that is, Zero. The above description can be summarized as: when HiF8 = 8'b 0 0000 000, the represented value X = Zero.

[0129] (2) NaN: When the value of the sign field S of the floating-point number is 1 and the Dot field is 0000, that is, the Dot field indicates that the floating-point number is DML data, that is, the first mantissa field represents the offset of the exponent of the floating-point number, and when the 3 bits of the first mantissa field are all 0, the corresponding floating-point number can represent not a number, that is, NaN. The above description can be summarized as: when HiF8 = 8'b 1 0000 000, it represents X = NaN.

[0130] (3) Positive infinity: When the bit width of the first exponent field of the floating-point number is 4, the bit width of the first mantissa field is 1, and the value of the first exponent field is 0111 and the value of the first mantissa field is 1, if the value of the sign field S of the floating-point number is 0, the corresponding floating-point number can represent +inf. The above description can be summarized as: the first exponent field Es = 4'b0111 = 15, the first mantissa field M = 1'b1, that is, when HiF8 = 8'b 0 11 0111 1, it represents X = +inf.

[0131] (4) Negative infinity: When the bit width of the first exponent field of the floating-point number is 4, the bit width of the first mantissa field is 1, and the value of the first exponent field is 0111 and the value of the first mantissa field is 1, if the value of the sign field S of the floating-point number is 1, the corresponding floating-point number can represent -inf. The above description can be summarized as: the first exponent field Es = 4'b0111 = 15, the first mantissa field M = 1'b1, that is, when HiF8 = 8'b 1 11 0111 1, it represents X = -inf.

[0132] Figure 3It is a schematic diagram of the significant precision - exponent distribution of the floating - point numbers of HiF8_DML provided by an embodiment of the present application. Among them, the size of the significant bits can represent the precision of the floating - point numbers. When the first mantissa field is used to represent the mantissa of the floating - point number, the significant bits of the mantissa can be the bit width of the first mantissa field plus 1. As Figure 3 shown, HiF8_DML has a conical precision characteristic and can provide a precision of up to 4 significant bits at most. At the same time, HiF8_DML has an exponent range of [-22, 15], which is much larger than the exponent range of FP8 [-9, 8] or [-16, 15], and is almost equivalent to the exponent range of FP16 [-24, 15]. It can be seen that HiF8_DML can expand the dynamic range of the values represented by floating - point numbers. Moreover, the numerical precision near the exponent center is significantly higher than that far from the exponent center. For example, when the value of the exponent is [-3, 3], the significant bits of the mantissa are the highest, which is 4 significant bits. When the value of the exponent is between ±[4, 7], the significant bits of the mantissa are 3 bits; when the value of the exponent is between ±[8, 15], the significant bits of the mantissa are 2 bits; when the value of the exponent is less than -15, the significant bits of the mantissa are 1 bit, and at this time, the mantissa is the default value 1.

[0133] The HiF8_DML provided by the embodiment of the present application can be used for data storage or data transfer. Correspondingly, the embodiment of the present application also provides a floating - point number processing method, and this method can be executed by the processor 110 as shown in 2. As Figure 4 shown, this method may include the following steps:

[0134] S401, obtain a first floating - point number.

[0135] The processor can read the first floating - point number from the local memory of the computing device to which the processor belongs, or can obtain the first floating - point number from other computing devices through the network. The first floating - point number may include a first sign field, a bit - width indication field, a first exponent field, and a first mantissa field. Among them, the first sign field is used to indicate the sign of the first floating - point number, that is, to indicate the positive or negative of the first floating - point number. The bit - width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating - point number; the bit - width indication field can also be used to indicate whether the first mantissa field represents the mantissa of the first floating - point number or the offset of the exponent of the first floating - point number.

[0136] S402, decode the first floating - point number to obtain a sign, an exponent, and a mantissa.

[0137] In some embodiments, the first floating - point number can be decoded to obtain a second floating - point number; the second floating - point number may include a second sign field, a second exponent field, and a second mantissa field. The second sign field is used to indicate the above - mentioned sign, the second exponent field is used to indicate the above - mentioned exponent, and the second mantissa field is used to indicate the above - mentioned mantissa.

[0138] Exemplarily, the value of the first sign field of the first floating-point number can be used as the value of the second sign field of the second floating-point number, and based on the values of the first exponent field and the first mantissa field of the first floating-point number, the second exponent field and the second mantissa field of the second floating-point number are determined.

[0139] When the bit width D of the first exponent field is 0, if the bit width indication field indicates that the first mantissa field represents the mantissa of the first floating-point number, the value indicated by the first mantissa field can be used as the value indicated by the second mantissa field of the second floating-point number, and the first preset exponent can be used as the value indicated by the second exponent field of the second floating-point number, where the first preset exponent can be -15. If the bit width indication field indicates that the first mantissa field represents the offset of the exponent of the first floating-point number, the preset mantissa can be used as the value indicated by the second mantissa field of the second floating-point number, and the value of the first mantissa field is used to correct the second preset exponent to obtain the value of the exponent field of the second floating-point number. Wherein, the preset mantissa can be 1, and the second preset exponent can be -23.

[0140] When the bit width D of the first exponent field is non-0, the first exponent field is used to represent the exponent of the first floating-point number, and the first mantissa field is used to represent the mantissa of the first floating-point number. At this time, based on the bit width indication field, the first exponent field and the first mantissa field can be determined in the first floating-point number, and then the value indicated by the first mantissa field is used as the value indicated by the second mantissa field of the second floating-point number, and the value indicated by the first exponent field is used as the value indicated by the second exponent field of the second floating-point number.

[0141] In some embodiments, the decoder in the processor can be used to convert the first floating-point number into the second floating-point number. Exemplarily, the schematic diagram of the working principle of the decoder can be as Figure 5 shown, and this decoder can also be called the HiF8_DML decoder. As Figure 5As shown, the first floating-point number sequentially includes a sign field, a Dot field, a first exponent field, and a first mantissa field. The total bit width N of the first floating-point number is 8. Among them, the sign field occupies 1 bit and is the highest bit in the first floating-point number. The Dot field occupies 2 to 4 bits, the first exponent field occupies 0 to 4 bits, and the first mantissa field occupies 1 to 3 bits. First, for the sign field in the floating-point number, the decoder can directly read and output its value, S = 0 or 1. Secondly, the decoder can determine the bit width of the Dot field based on the value of the second, third, or fourth bit in the first floating-point number. For example, if the value of the second bit in the first floating-point number is "1", the bit width of the Dot field can be determined to be 2; if the value of the second bit in the first floating-point number is "0" and the value of the third bit is "1", the bit width of the Dot field can also be determined to be 2; if the values of the second and third bits in the first floating-point number are both "0" and the value of the fourth bit is "1", the bit width of the Dot field can be determined to be 3; if the values of the second, third, and fourth bits in the first floating-point number are all "0", the bit width of the Dot field can be determined to be 4.

[0142] If the bit width of the Dot field is 4, the value of the first mantissa field can be determined based on the value of the Dot field to represent the mantissa of the first floating-point number or the offset of the exponent of the first floating-point number.

[0143] If the first mantissa field represents the mantissa of the first floating-point number, it indicates that the first floating-point number is normal data. The decoder can read the value of the first mantissa field, use the value of the first mantissa field as the value of the second mantissa field of the second floating-point number, and use the first preset exponent as the value of the second exponent field of the second floating-point number. Among them, the first preset exponent can be -15.

[0144] If the first mantissa field represents the offset of the exponent of the first floating-point number, it indicates that the first floating-point number is DML data. The decoder can read the value of the first mantissa field and use the value of the first mantissa field to correct the second preset exponent to obtain the value of the exponent field of the second floating-point number. For example, the second preset exponent can be -23. When the value of the first mantissa field is M, the value of the second exponent field of the second floating-point number can be M - 23. The decoder can use the preset mantissa 1 as the value represented by the second mantissa field of the second floating-point number.

[0145] If the bit width of the Dot field is 2 or 3, the decoder can extract the Dot field and, through a multiplexer (MUX) operation on the Dot field, decode the Dot field based on the preset coding rule to obtain the value D of the Dot field. The preset coding rule can be as shown in Table 4 above, where the coding "11" can represent the value "4", the coding "10" can represent the value "3", the coding "01" can represent the value "2", and the coding "001" can represent the value "1".

[0146] The decoder can, based on the value D of the Dot field, extract the first exponent field and the first mantissa field located after the first exponent field after the Dot field, and use the value indicated by the first exponent field as the value indicated by the second exponent field of the second floating-point number; use the value indicated by the first mantissa field as the value indicated by the second mantissa field of the second floating-point number. Thus, the decoder can determine the values of the second exponent field, the second mantissa field, and the second sign field of the second floating-point number according to the values of the first exponent field, the first mantissa field, and the first sign field in the above-mentioned first floating-point number.

[0147] For example, the first floating-point number input to the decoder is "11010110", the total bit width N is 8 bits, the first bit "1" is the sign field, and the second bit to the third bit "10" is the Dot field. According to the preset coding rule, the value of this Dot field is decoded to be 3, that is, it is determined that the bit width of the first exponent field is 3 bits. Therefore, the decoder can accurately extract the fourth bit to the sixth bit "101" in the first floating-point number as the first exponent field, and the first exponent field is used to represent the exponent of the first floating-point number. The remaining seventh bit to the eighth bit "10" is the first mantissa field, and the first mantissa field is used to represent the mantissa of the first floating-point number. Based on this, it can be decoded that the value represented by the exponent field of the second floating-point number is 5, and the value represented by the mantissa field of the second floating-point number is 0.10.

[0148] In some other embodiments, the first floating-point number can also be decoded to directly obtain the sign, exponent, and mantissa. Exemplarily, the sign can be determined according to the value of the first sign field of the first floating-point number. For example, if the value of the first sign field is 0, it means the sign is positive, and if the value of the first sign field is 1, it means the sign is negative.

[0149] When the bit width D of the first exponent field is 0, it means the first exponent field does not exist. At this time, if the Dot field indicates that the first mantissa field represents the mantissa of the first floating-point number, the exponent can be the first preset exponent -15. If the Dot field indicates that the first mantissa field represents the offset of the exponent of the first floating-point number, the exponent can be the value obtained by correcting the second preset exponent -23 using this offset, and the mantissa can be the preset mantissa 1.

[0150] When the bit width D of the first exponent field is not 0, the mantissa is the value indicated by the first mantissa field of the first floating-point number, and the exponent is the value indicated by the first exponent field of the first floating-point number.

[0151] After obtaining the sign, exponent, and mantissa, calculations can be performed based on the sign, exponent, and mantissa. Exemplarily, in some other embodiments, as Figure 6 shown, the floating-point processing method executed by the processor may include the following steps:

[0152] S601, obtain the first floating-point number.

[0153] S602, Decode the first floating-point number to obtain the sign, exponent, and mantissa.

[0154] S603, Perform calculations using the sign, exponent, and mantissa to obtain the calculation result.

[0155] In some embodiments, the processor includes a decoder. The processor can decode the first floating-point number through the decoder to obtain the sign, exponent, and mantissa. The decoder can transmit the decoded sign, exponent, and mantissa to the calculation unit in the processor. The calculation unit receives the sign, exponent, and mantissa and performs corresponding calculations to obtain the calculation result. Among them, the calculation result can be a floating-point number with the same encoding method as the above-mentioned second floating-point number. That is to say, the calculation result can include a sign field, an exponent field, and a mantissa field. The sign field is used to indicate the sign of the calculation result, the exponent field is used to indicate the exponent of the calculation result, and the mantissa field is used to indicate the mantissa of the calculation result.

[0156] S604, Based on the exponent and mantissa of the calculation result, convert the calculation result into a third floating-point number.

[0157] The third floating-point number is a floating-point number with the same encoding method as the first floating-point number. The third floating-point number can include a first sign field, a Dot field, a first exponent field, and a first mantissa field.

[0158] The processor can use the value of the sign field of the calculation result as the value of the first sign field of the third floating-point number, and determine the numerical range to which the calculation result belongs according to the exponent and mantissa of the calculation result.

[0159] If the numerical range to which the calculation result belongs is the first set range, the value of the first mantissa field of the third floating-point number can be determined according to the exponent field and mantissa field of the calculation result, and the Dot field can be set to the first value; the first value is used to indicate that the first mantissa field represents the offset of the exponent of the first floating-point number. Among them, the first set range refers to the range of DML data.

[0160] If the numerical range to which the calculation result belongs is the second set range, the values of the first exponent field, the first mantissa field, and the Dot field of the third floating-point number can be determined respectively according to the exponent field and mantissa field of the calculation result. Among them, the second set range refers to the range of normal data. If the value of the first exponent field is determined to be the first preset exponent according to the exponent field and mantissa field of the calculation result, that is, the bit width D of the first exponent field is 0, the Dot field can be set to the second value, and the second value is used to indicate that the first mantissa field represents the mantissa of the third floating-point number.

[0161] In some embodiments, an encoder may also be included in the processor. After the computing unit in the processor obtains the calculation result, the calculation result may be transmitted to the encoder, and the calculation result is converted into a third floating-point number by the encoder. The encoder may also be referred to as a HiF8_DML encoder. Exemplarily, a schematic diagram of the working principle of the encoder may be as shown in Figure 7 shown. The calculation result may include a sign field, an exponent field, and a mantissa field. The encoder may use the value of the sign field of the calculation result as the value of the first sign field of the third floating-point number. The encoder may determine the numerical range to which the calculation result belongs according to the value of the exponent field of the calculation result.

[0162] If the numerical range to which the calculation result belongs is the first set range, that is, the calculation result is DML data, the encoder may determine the value of the first mantissa field of the third floating-point number according to the exponent and mantissa of the calculation result. Exemplarily, the encoder may process the mantissa of the calculation result and adjust the exponent of the calculation result according to the processed mantissa, and obtain the first mantissa field of the third floating-point number based on the adjusted exponent. For example, assuming that the mantissa of the calculation result is 1.XXXXXX, the encoder may perform mantissa shifting and rounding operations on the mantissa of the calculation result, round it to an integer bit, and obtain the processed mantissa 1.0. At this time, the value of the exponent will increase by 1. Therefore, the exponent of the calculation result may be adjusted, and the value of the adjusted exponent is added with 23 to obtain the value of the first mantissa field of the third floating-point number, and the first mantissa field of the third floating-point number is encoded according to the value of the first mantissa field. The encoder may set the Dot field to a first value, and the first value is used to indicate that the first mantissa field represents the offset of the exponent of the third floating-point number.

[0163] If the numerical range to which the calculation result belongs is the second set range, that is, the calculation result is normal data, the encoder may respectively determine the first exponent field and the first mantissa field of the third floating-point number according to the exponent and mantissa of the calculation result, and determine the bit width indication field of the third floating-point number based on the first exponent field. Exemplarily, the encoder may process the mantissa of the calculation result and adjust the exponent of the calculation result according to the processed mantissa, obtain the value of the first mantissa field of the third floating-point number based on the processed mantissa, and obtain the value of the first exponent field of the third floating-point number based on the adjusted exponent. Based on the value of the first exponent field, determine the bit width D occupied by the first exponent field. Based on the bit width D occupied by the first exponent field, determine the value of the bit width indication field. When the bit width D occupied by the first exponent field is 0, set the bit width indication field to a second value; the second value is used to indicate that the first mantissa field represents the mantissa of the first floating-point number. For example, the encoder may respectively implement encoding of the Dot field, the first exponent field, and the first mantissa field of the third floating-point number by performing a leading 1 search operation on the absolute value of the exponent in the calculation result, and performing operations such as mantissa shifting and rounding on the mantissa in the calculation result.

[0164] For example, assume that the sign bit of the calculation result is 1, the exponent is "00011" (i.e., 3), and the mantissa is "1111". The encoder can determine that the sign bit of the third floating-point number is "1". By performing a leading 1 search on the absolute value of the exponent "00011" until the first 1 is found, the encoder can determine that the bit width of the first exponent field of the third floating-point number is 2 bits, and the encoded value of the first exponent field is "01". Among them, "0" is the exponent sign bit, indicating that the exponent is positive, and the highest bit "1" in the exponent magnitude "11" can be hidden and does not occupy the bit width. As described above, it will not be elaborated here. According to the bit width of 2 bits of the first exponent field, the encoder can determine that the value represented by the Dot field of the third floating-point number is 2. Still taking the encoding rule shown in Table 4 as an example, since the value represented by the Dot field is 2, it can be determined that the encoded value corresponding to the Dot field is "01", which occupies a bit width of 2 bits, then the remaining bit width of the first mantissa field that can be encoded is 3 bits. And the bit width of the input mantissa "1111" of the calculation result is greater than the bit width of the first mantissa field of the third floating-point number. Therefore, the mantissa "1111" can be rounded to obtain the mantissa "10.000" including the hidden bit on the left side of the decimal point. At this time, the decimal point in the mantissa "10.000" needs to be shifted one place to the left to obtain the mantissa "1.0000" including the hidden bit on the left side of the decimal point. Since the decimal point is shifted one place to the left, correspondingly, the exponent needs to be incremented by 1 (i.e., 3 plus 1, which is 4), that is, the magnitude of the first exponent field is 4, and the corresponding encoded value is "100". Among them, the highest bit "1" in "100" is hidden, and the final encoded value of the first exponent field is "00", and the final encoded value of the first mantissa field is "000". Finally, the third floating-point number encoded by the encoder is "10100000".

[0165] The finally obtained third floating-point number is the data of HiF8_DML provided by the embodiment of the present application. The third floating-point number can be used for data storage and data transfer to reduce the resources required for data storage or data transfer.

[0166] Based on the same inventive concept as the above embodiment, the embodiment of the present application further provides a floating-point number processing method for obtaining the first floating-point number used in the above process. As Figure 8 shown, the method may include the following steps:

[0167] S801, obtain the sign, exponent, and mantissa.

[0168] S802, based on the sign, exponent, and mantissa, obtain the first floating-point number.

[0169] For example, a processor may obtain data to be encoded. The data to be encoded may be data in a format different from that of the first floating-point number, or may be an operation result output by a computing unit. The data to be encoded may include a sign, an exponent, and a mantissa. The processor extracts the sign, exponent, and mantissa from the data to be encoded, and obtains a first floating-point number based on the sign, exponent, and mantissa.

[0170] The process by which the processor obtains the first floating-point number based on the sign, exponent, and mantissa may be performed by referring to the process of obtaining the third floating-point number based on the calculation result in the foregoing text, and will not be elaborated herein.

[0171] The floating-point number provided in the embodiment of the present application, on the basis of the standard sign field, exponent field, and mantissa field, additionally adds a Dot field, and the Dot field is used to indicate the effective bit width of the first exponent field of the floating-point number. This bit width is the bit width D occupied by the exponent of the floating-point number during actual storage, so that the bit width of the first exponent field in the floating-point number can change dynamically with the value of the Dot field. Correspondingly, the bit width of the first mantissa field in the floating-point number also changes dynamically, so that the floating-point number has a conical precision characteristic, so that the data at the exponent center has a higher mantissa bit width, that is, has a higher precision; the data farther away from the exponent center has a gradually decreasing mantissa bit width and a decreasing precision. The exponent of the floating-point number can determine the value range of the floating-point number. Therefore, the embodiment of the present application can effectively balance the total bit width, value range, and value precision of the floating-point number, and meet the different requirements for the value range and value precision of the floating-point number in various scenarios without additionally increasing the total bit width, data storage, or data transfer cost, and improve the use effect of the floating-point number.

[0172] Exemplarily, the floating-point representation method provided by the embodiments of the present application can be applied to lossy compression of high-precision data and preliminary approximate solution calculation of a mixed-precision matrix solver. For general computing and high-performance computing, the embodiments of the present application can obtain a higher convergence speed and accuracy of computing tasks under the same total bit width, that is, the same data storage or data transfer overhead. For AI neural network training and inference, the embodiments of the present application can meet the functional and accuracy requirements of neural network training and inference under the same total bit width. When the bit width occupied by the first exponent field in the embodiments of the present application is 0, 3 bits of coding space are provided for the mantissa, and the mantissa accuracy of 3 bits can already meet the needs of AI neural network training and inference. Therefore, the embodiments of the present application set the bit width of the Dot field to 4 bits, and 1 bit of it is used to indicate whether the first mantissa field represents the mantissa or the offset of the exponent. If the first mantissa field represents the offset of the exponent, it means that the floating-point number is DML data. When the floating-point number is DML data, the mantissa can adopt the default value 1, and the mantissa accuracy can meet the needs of AI neural network training and inference. At this time, borrowing the first mantissa field to represent the offset of the exponent can expand the dynamic range of the exponent and the numerical range of the representable floating-point numbers.

[0173] For example, as introduced above, the floating-point representation provided by the embodiments of the present application has a numerical range of [-22, 15], which is much larger than the [-9, 8] or [-16, 15] numerical range of FP8. The numerical range provided by the embodiments of the present application can also meet the needs of AI neural network training and inference. For example, for the end-to-end training of a large language model (LLM), such as Figure 9 shown, when using the traditional floating-point representation FP8 for the end-to-end training of the LLM, due to the relatively small numerical range of the data that can be represented by the floating-point number, when there are a large number of data with very small values, these data cannot be distinguished and are considered the same value, resulting in an inability to obtain the correct training result. It is manifested that when training for more than a hundred times, the loss value suddenly increases, the LLM collapses, and cannot converge. However, when using the floating-point representation HiF8_DML provided by the embodiments of the present application for the end-to-end training of the LLM, as Figure 10 shown, since the numerical range of the data that can be represented by the floating-point number is relatively large and can meet the needs of LLM training, the training accuracy is improved, the LLM can converge normally, and thus the performance and stability of LLM training can be improved.

[0174] In summary, the embodiments of the present application can not only ensure a small impact on the numerical accuracy of AI neural network training and inference, but also free up coding space to express DML data, thus well balancing the numerical accuracy and the numerical range.

[0175] In addition, the embodiments of the present application also define the numerical range that can be encoded by the exponent code in the first exponent code field under different bit widths, as well as the numerical range that can be encoded by the first mantissa field when representing the offset of the exponent code. This effectively avoids the problem of overlap in the numerical values of the exponent code under different bit widths, making the data encoding method in the embodiments of the present application free of information duplication and being a non-redundant encoding. Based on this, the highest bit of the original code amplitude in the first exponent code field can be hidden without storage, which can further reduce the data storage or data transfer cost of floating-point numbers.

[0176] Based on the same design concept as the above method embodiments, the embodiments of the present application also provide a floating-point processing device. This floating-point processing device can be applied to Figure 1 or Figure 2 the computing devices shown, or, applied to Figure 2 the decoder 111 shown. This floating-point processing device can be used to implement the functions of the above method embodiments, and thus can achieve the beneficial effects possessed by the above method embodiments. As Figure 11 shown, the floating-point processing device 1100 can include a floating-point acquisition module 1101 and a decoding module 1102.

[0177] Among them, the floating-point acquisition module 1101 can be used to acquire a first floating-point number; the first floating-point number includes a first sign field, a bit width indication field, a first exponent code field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit width indication field is used to indicate the bit width D occupied by the first exponent code field in the total bit width N of the first floating-point number; the bit width indication field is also used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or, the offset of the exponent code of the first floating-point number;

[0178] The decoding module 1102 can be used to decode the first floating-point number to obtain the sign, exponent code, and mantissa.

[0179] In some embodiments, the decoding module 1102 can specifically be used to: decode the first floating-point number to obtain a second floating-point number; the second floating-point number includes a second sign field, a second exponent code field, and a second mantissa field, and the second sign field is used to indicate the sign, the second exponent code field is used to indicate the exponent code, and the second mantissa field is used to indicate the mantissa.

[0180] In some embodiments, the floating-point processing device 1100 may further include a calculation module, and the calculation module can be used to: perform calculations using the sign, exponent code, and mantissa. In some other embodiments, the calculation module may also be located in a calculation unit outside the decoder, and the floating-point processing device 1100 can transmit the sign, exponent code, and mantissa to the calculation unit, and the calculation unit performs calculations using the sign, exponent code, and mantissa.

[0181] It should be noted that in some embodiments, the floating-point number acquisition module 1101 can be used to execute any step in the floating-point data processing method, and the decoding module 1102 can be used to execute any step in the floating-point data processing method. The steps to be implemented by the floating-point number acquisition module 1101 and the decoding module 1102 can be specified as needed. The floating-point number acquisition module 1101 and the decoding module 1102 respectively implement different steps in the floating-point data processing method to implement all functions of the floating-point processing device.

[0182] In the embodiments of the present application, each functional module can be integrated in one processor, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional units.

[0183] Based on the same design concept as the above method embodiments, the embodiments of the present application also provide a floating-point processing device. This floating-point processing device can be applied to Figure 1 or Figure 2 the computing device shown, or applied to Figure 2 the encoder 112 shown. This floating-point processing device can be used to implement the functions of the above method embodiments, and thus can achieve the beneficial effects of the above method embodiments. As Figure 12 shown, the floating-point processing device 1200 can include a data acquisition module 1201 and an encoding module 1202.

[0184] Among them, the data acquisition module 1201 can be used to acquire a sign, an exponent, and a mantissa; the encoding module 1202 can be used to obtain a first floating-point number based on the sign, the exponent, and the mantissa. The first floating-point number includes a first sign field, a bit-width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit-width indication field is used to indicate the bit-width D occupied by the first exponent field in the total bit-width N of the first floating-point number; the bit-width indication field is also used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number.

[0185] It should be noted that in some embodiments, the data acquisition module 1201 can be used to execute any step in the floating-point data processing method, and the encoding module 1202 can be used to execute any step in the floating-point data processing method. The steps to be implemented by the data acquisition module 1201 and the encoding module 1202 can be specified as needed. The data acquisition module 1201 and the encoding module 1202 respectively implement different steps in the floating-point data processing method to implement all functions of the floating-point processing device.

[0186] In the embodiments of the present application, each functional module can be integrated in a processor, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional unit.

[0187] Based on the same technical concept as the above method embodiments, the embodiments of the present application also provide a computing device. The computing device can be Figure 1 or Figure 2 any one of the computing devices shown in Figure 4 , Figure 6 or Figure 8 shown. The computing device can be used to implement the functions of the method embodiments shown above, and thus can achieve the beneficial effects possessed by the above method embodiments.

[0188] In some embodiments, the structure of the computing device 1300 can be as shown in Figure 13 , including a processor 1301 and a memory 1302 connected to the processor 1301. The processor 1301 and the memory 1302 can be interconnected through a bus. The processor 1301 can be a general-purpose processor, such as a microprocessor, or other conventional processors. The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0189] Among them, the memory 1302 can be used to store software programs and modules. The processor 1301 executes various functional applications and data processing of the terminal device 1300 by running the software programs and modules stored in the memory 1302, such as any floating-point processing method provided by the embodiments of the present application.

[0190] The memory 1302 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs of at least one application, etc.; the data storage area can be used to store user data, etc. In addition, the memory 1302 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0191] The processor 1301 in the computing device 1300 is used to run the computer instructions or programs saved in the memory 1302 and execute the functions in any of the above method embodiments. In some embodiments, the processor 1301 may include one or more processing units. Different processing units may be independent devices or integrated in one or more processors. The processor 1301 may also include a controller. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0192] In some embodiments, the processor 1301 of the computing device 1300 may include a decoder, and the processor 1301 may be used to implement Figure 4 the floating-point processing method shown; in some other embodiments, the processor 1301 of the computing device 1300 may include an encoder, and the processor 1301 may be used to implement Figure 8 the floating-point processing method shown; or, in some other embodiments, the processor 1301 of the computing device 1300 may include a decoder and an encoder, and the processor 1301 may be used to implement Figure 6 the floating-point processing method shown.

[0193] In one embodiment, the computing device 1300 may further include a communication module, and the communication module may be used to communicate with a network device.

[0194] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the terminal device. In some other embodiments of the present application, the terminal device may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0195] The embodiments of the present application further provide a computer program product including computer-executable instructions. In one embodiment, the computer-executable instructions are used to cause a computer to execute the functions in the above method embodiments.

[0196] The computer-executable instructions may be stored in a computer-readable storage medium. The embodiments of the present application further provide a computer-readable storage medium, and the computer-readable storage medium stores executable instructions. In one embodiment, the computer-executable instructions are used to cause a computer to execute the functions in the above method embodiments.

[0197] The computer-readable storage medium provided by the embodiments of the present application may be a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a removable hard disk, a CD-ROM, or any other form of computer-readable storage medium well known in the art.

[0198] Computer-executable instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer program or instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center integrating one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disc (DVD); or it may be a semiconductor medium, such as a solid-state drive.

[0199] In various embodiments of the present application, if there is no special description and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be cross-referenced. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a method, system, product, or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0200] Although the present application has been described in conjunction with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, the present specification and the drawings are merely exemplary illustrations of the solutions defined by the appended claims and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present application.

[0201] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the scope of this application. Thus, if these modifications and variations of the embodiments of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A floating-point number processing method, characterized in that, The method includes: Obtain a first floating-point number, where the first floating-point number includes a first sign field, a bit-width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit-width indication field is used to indicate the bit-width D occupied by the first exponent field in the total bit-width N of the first floating-point number; the bit-width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number. Decode the first floating-point number to obtain the sign, the exponent, and the mantissa.

2. The method according to claim 1, characterized in that, The decoding of the first floating-point number to obtain the sign, the exponent, and the mantissa includes: Decode the first floating-point number to obtain a second floating-point number; the second floating-point number includes a second sign field, a second exponent field, and a second mantissa field, where the second sign field is used to indicate the sign, the second exponent field is used to indicate the exponent, and the second mantissa field is used to indicate the mantissa.

3. The method according to claim 1 or 2, characterized in that, When the bit-width D is 0, the first exponent field does not exist. If the first mantissa field represents the mantissa, the exponent is a first preset exponent.

4. The method according to any one of claims 1 to 3, characterized in that, When the bit-width D is 0, the first exponent field does not exist. If the first mantissa field represents the offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using the offset.

5. The method according to any one of claims 1 to 4, characterized in that When the bit-width D is not 0, the first mantissa field represents the mantissa, and the exponent is the value indicated by the first exponent field.

6. The method according to any one of claims 1 to 5, characterized in that The bit-width DW of the bit-width indication field is negatively correlated with the bit-width D.

7. The method according to any one of claims 1 to 6, characterized in that When the bit-width DW of the bit-width indication field is a preset bit-width and the value indicated by the bit-width indication field is a preset value, the first mantissa field represents the offset.

8. The method according to any one of claims 1 to 7, characterized in that, The obtaining of the first floating-point number includes: Read the first floating-point number from a memory, or obtain the first floating-point number through a communication network.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Perform calculations using the sign, the exponent, and the mantissa.

10. A floating-point number processing method, characterized in that, The method includes: Obtain a sign, an exponent, and a mantissa; Based on the sign, the exponent, and the mantissa, obtain a first floating-point number; the first floating-point number includes a first sign field, a bit-width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit-width indication field is used to indicate the bit-width D occupied by the first exponent field in the total bit-width N of the first floating-point number; the bit-width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number.

11. The method according to claim 10, wherein When the bit-width D is 0, the first exponent field does not exist. If the first mantissa field represents the mantissa, the exponent is a first preset exponent.

12. The method according to claim 10 or 11, characterized in that, When the bit-width D is 0, the first exponent field does not exist. If the first mantissa field represents the offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using the offset.

13. The method according to any one of claims 10 to 12, characterized in that When the bit width D is not 0, the first mantissa field represents the mantissa, and the exponent is the value indicated by the first exponent field.

14. The method according to any one of claims 10 to 13, characterized in that The bit width DW of the bit width indication field is negatively correlated with the bit width D.

15. The method according to any one of claims 10 to 14, characterized in that When the bit width DW of the bit width indication field is a preset bit width and the value indicated by the bit width indication field is a preset value, the first mantissa field represents the offset.

16. A floating-point processing device, characterized in that, The apparatus includes: A floating-point number acquisition module, configured to acquire a first floating-point number, where the first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number. A decoding module, configured to decode the first floating-point number to obtain the sign, the exponent, and the mantissa.

17. The device according to claim 16, wherein Specifically, the decoding module is configured to: Decode the first floating-point number to obtain a second floating-point number; the second floating-point number includes a second sign field, a second exponent field, and a second mantissa field, the second sign field is used to indicate the sign, the second exponent field is used to indicate the exponent, and the second mantissa field is used to indicate the mantissa.

18. A floating-point processing device, characterized in that, The apparatus includes: A data acquisition module, configured to acquire a sign, an exponent, and a mantissa. An encoding module, configured to obtain a first floating-point number based on the sign, the exponent, and the mantissa; the first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number.

19. The device according to claim 18, characterized in that, When the bit width D is 0, the first exponent field does not exist. If the first mantissa field represents the offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using the offset.

20. A computing device, characterized in that, Including a processor and a memory; A computer program is stored on the memory; The processor is configured to read the computer program stored in the memory and execute the following steps: Acquire a first floating-point number, where the first floating-point number includes a first sign field, a bit width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit width indication field is used to indicate the bit width D occupied by the first exponent field in the total bit width N of the first floating-point number; the bit width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number. Decode the first floating-point number to obtain the sign, the exponent, and the mantissa.

21. The computing device according to claim 20, wherein Specifically, the processor is configured to: Decode the first floating-point number to obtain a second floating-point number; the second floating-point number includes a second sign field, a second exponent field, and a second mantissa field, where the second sign field is used to indicate the sign, the second exponent field is used to indicate the exponent, and the second mantissa field is used to indicate the mantissa.

22. A computing device, characterized in that, Comprising a processor and a memory; A computer program is stored on the memory; The processor is configured to read the computer program stored in the memory and execute the following steps: Obtain a sign, an exponent, and a mantissa; Based on the sign, the exponent, and the mantissa, obtain a first floating-point number; the first floating-point number includes a first sign field, a bit-width indication field, a first exponent field, and a first mantissa field; the first sign field is used to indicate the sign of the first floating-point number, and the bit-width indication field is used to indicate the bit-width D occupied by the first exponent field in the total bit-width N of the first floating-point number; the bit-width indication field is further used to indicate that the first mantissa field represents the mantissa of the first floating-point number, or the offset of the exponent of the first floating-point number.

23. The computing device according to claim 22, wherein When the bit-width D is 0, the first exponent field does not exist. If the first mantissa field represents the offset, the mantissa is a preset mantissa, and the exponent is a value obtained by correcting a second preset exponent using the offset.

24. A computer-readable storage medium, characterized in that, Stored with computer-executable instructions for causing a computer to execute the method according to any one of claims 1 to 9; or, the method according to any one of claims 10 to 15.

25. A computer program product, characterized in that, Containing computer-executable instructions for causing a computer to execute the method according to any one of claims 1 to 9; or, the method according to any one of claims 10 to 15.

Citation Information

Cited By

  • Floating point number processing method, apparatus, computing device and storage medium

    WO2025140047A1