Method and apparatus for quantizing floating-point numbers, and electronic device and storage medium

WO2026200578A1PCT designated stage Publication Date: 2026-10-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/083538
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-13
Publication Date
2026-10-01

Smart Images

  • Figure CN2026083538_01102026_PF_FP_ABST
    Figure CN2026083538_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a method and apparatus for quantizing floating-point numbers, and an electronic device and a computer-readable storage medium. During the process of determining a target scale factor required for quantization, the method not only takes all exponent bits of floating-point numbers into consideration, but also takes a preset number of mantissa bits into consideration. In this way, the numerical characteristics of the floating-point numbers can be captured more comprehensively, thereby reducing truncation errors during quantization, and determining a more accurate target scale factor. The precision of floating-point numbers obtained by means of subsequently performing quantization by using the target scale factor will also be higher. That is, the conversion error between floating-point numbers before and after quantization is relatively small, and the data conversion efficiency and precision are also higher. In addition, during the process of determining the target scale factor, the target scale factor is determined on the basis of statistical characteristics of a plurality of floating-point numbers, thereby ensuring that all floating-point numbers can be appropriately represented. This not only avoids the loss of data, but also prevents the overflow of data.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, electronic devices, and storage media for quantizing floating-point numbers

[0001] This application claims priority to Chinese Patent Application No. 202510390457.6, filed on March 28, 2025, entitled “Method, Apparatus, Electronic Device and Storage Medium for Quantizing Floating-Point Numbers”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The embodiments of this application relate to the field of computer technology. The embodiments of this application relate to methods for quantizing floating-point numbers, apparatus for quantizing floating-point numbers, electronic devices, computer-readable storage media, and computer program products. Background Technology

[0003] With the rapid development of Artificial Intelligence (AI), neural networks are becoming increasingly large, leading to a surge in computational demands for training neural network models, such as AI models. This places ever-higher requirements on hardware memory and performance. Currently, quantizing models, such as converting floating-point numbers to low-bit data, is a future trend in AI model training. This is understandable, as training AI models with low-bit data requires less memory and enables robust training and inference accuracy. AI model training typically involves conversions between floating-point data of different precisions. For example, converting high-precision floating-point data to low-precision floating-point data. This process requires rounding the high-precision floating-point data, which introduces conversion errors and affects the training accuracy of the AI ​​model. Summary of the Invention

[0004] This application provides a scheme for quantizing floating-point numbers. In determining the target scaling factor required for quantization, this scheme considers not only all exponent bits of the floating-point number but also the mantissa bits of a preset number of bits. This allows for a more comprehensive capture of the numerical characteristics of the floating-point number, thereby reducing truncation errors during quantization and determining a more accurate target scaling factor. This more accurate target scaling factor not only balances the precision and representation range of the floating-point number, ensuring that the quantized floating-point number maintains high precision without exceeding its representable range, but also results in higher precision for the subsequent floating-point number quantized using the target scaling factor. In other words, the conversion error between the floating-point numbers before and after quantization is smaller, and the data conversion efficiency and precision are higher. Furthermore, the target scaling factor is determined based on the statistical characteristics of multiple floating-point numbers, ensuring that all floating-point numbers are appropriately represented. This avoids data loss and prevents data overflow. According to a first aspect of this application, a method for quantizing floating-point numbers is provided. This method includes obtaining statistical values ​​corresponding to multiple source floating-point numbers, as well as the exponent and mantissa of the statistical values. The method also includes determining a target scaling factor based on the preset number of mantissa bits in the exponent and mantissa. In addition, the method also includes determining multiple target floating-point numbers by quantizing multiple source floating-point numbers based on a target scaling factor.

[0005] In some embodiments of the first aspect, determining the target scaling factor based on a preset number of mantissas in the exponent and mantissa includes: determining a target operation value based on a preset number of mantissas in the exponent and mantissa; and determining the target scaling factor by performing bit-level operations on the target operation value and the statistical value. In this manner, by using the target operation related to the mantissa and exponent of the statistical value to perform bit-level operations on the statistical value, not only can the time complexity of calculating the shared scaling factor be significantly reduced, thereby improving the efficiency of determining the shared scaling factor, but also the accuracy of determining the target scaling factor can be improved.

[0006] In some embodiments of the first aspect, determining the target scaling factor by performing bit-level operations on the target operational value and the statistical value includes: determining the target statistical value corresponding to the statistical value based on the exponent of the statistical value and a preset number of mantissas; and determining the target scaling factor by performing bit-level operations on the target operational value and the target statistical value. In this way, the target statistical value can be used to indicate which bits are important and which can be ignored. Through the operation between the target statistical value and the target operational value, the key bits used to determine the target scaling factor can be extracted. This avoids unnecessary calculations, thereby saving computation time and memory resources and improving the efficiency of determining the target scaling factor.

[0007] In some implementations of the first aspect, determining the target operand value based on a preset number of mantissas in the exponent and mantissa includes: splitting the mantissa portion of the statistical value into high-order mantissas and low-order mantissas, the splitting method being determined according to a preset number; and determining the target operand value based on the exponent and high-order mantissas. Since the high-order mantissas better reflect the influence of the mantissa portion of the statistical value on the target scaling factor, the target operand value determined using the high-order mantissas is more accurate, and the target scaling factor subsequently determined based on this is also more accurate. Furthermore, the high-order mantissas have fewer digits than the total mantissas, thus requiring less computational and storage resources.

[0008] In some embodiments of the first aspect, determining the target scaling factor by performing bit-level operations on the target operational value and the target statistical value includes: determining a first candidate scaling factor based on the sum of the target operational value and the target statistical value; determining a range mapping parameter based on the format of the source floating-point number and the format of the target floating-point number; and determining the target scaling factor based on the range mapping parameter and the first candidate scaling factor.

[0009] In some implementations of the first aspect, determining the target scaling factor based on the range mapping parameters and the first candidate scaling factor includes: determining a second candidate scaling factor based on the range mapping parameters and the first candidate scaling factor; and determining the target scaling factor by shifting the second candidate scaling factor, wherein the shift amount is determined according to the format of the source floating-point number. In this way, the relevant parameters of the range mapping can be determined according to the types of the source and target floating-point numbers, providing a reference basis for subsequent floating-point quantization.

[0010] In some implementations of the first aspect, determining the second candidate scaling factor based on the range mapping parameter and the first candidate scaling factor includes: determining the calculation result by subtracting the range mapping parameter and the first candidate scaling factor; and determining the second candidate scaling factor based on the calculation result. In this way, the second scaling factor can be determined based on simple bitwise operations, providing reliable data support for the subsequent determination of the target scaling factor.

[0011] In some implementations of the first aspect, obtaining the exponent and mantissa of a statistical value includes: determining the corresponding parameters based on the format of the source floating-point number; and obtaining the exponent and mantissa of the statistical value based on the AND value of the statistical value and the parameters. In this manner, by utilizing simple bit operations, the time complexity of obtaining the mantissa and exponent of the statistical value can be significantly reduced.

[0012] In some embodiments of the first aspect, obtaining the statistical values ​​corresponding to multiple source floating-point numbers includes: determining the maximum value among the multiple source floating-point numbers; and using the maximum value as the statistical value corresponding to the multiple source floating-point numbers. Using the maximum value of the multiple floating-point numbers to determine the target scaling factor allows for maximum utilization of the exponential range of the target floating-point numbers.

[0013] In some implementations of the first aspect, the plurality of source floating-point numbers includes a first source floating-point number and a second source floating-point number. Based on a target scaling factor, the plurality of target floating-point numbers are determined by quantizing the plurality of source floating-point numbers, including: determining a first target floating-point number corresponding to the first source floating-point number based on the target scaling factor; and determining a second target floating-point number corresponding to the second source floating-point number based on the target scaling factor. In this manner, each source floating-point number undergoes the same quantization process to obtain its corresponding target floating-point number, thereby ensuring the consistency and predictability of the quantization.

[0014] In some implementations of the first aspect, determining the first target floating-point number corresponding to the first source floating-point number based on the target scaling factor includes: determining the private element corresponding to the first target floating-point number based on the target scaling factor and the first source floating-point number; and determining the first target floating-point number based on the private element and the target scaling factor.

[0015] In some embodiments of the first aspect, bit-level operations include basic arithmetic performed at the bit level, whereby basic arithmetic includes one or more of logical AND, addition, subtraction, and shifting.

[0016] According to a second aspect of this application, an electronic device is provided, comprising: a processing unit and a memory, wherein the processing unit executes instructions in the memory to cause the electronic device to perform a method, the method comprising: acquiring statistical values ​​corresponding to a plurality of source floating-point numbers and the exponent and mantissa of the statistical values; determining a target scaling factor based on a preset number of mantissa bits in the exponent and mantissa; and determining a plurality of target floating-point numbers by quantizing the plurality of source floating-point numbers based on the target scaling factor.

[0017] According to a third aspect of this application, an apparatus for quantizing floating-point numbers is provided, comprising: a statistical value acquisition unit, a target scaling factor determination unit, and a quantization unit. The statistical value acquisition unit is configured to acquire statistical values ​​corresponding to multiple source floating-point numbers, as well as the exponent and mantissa of the statistical values. The target scaling factor determination unit is configured to determine a target scaling factor based on a preset number of mantissas in the exponent and mantissa. The quantization unit is configured to determine multiple target floating-point numbers by quantizing the multiple source floating-point numbers based on the target scaling factor.

[0018] According to a fourth aspect of this application, a computer-readable storage medium is provided that stores one or more computer instructions thereon, wherein one or more computer instructions are executed by a processor to cause the processor to perform the method according to a first aspect of this application.

[0019] According to a fifth aspect of this application, a computer program product is provided, including machine-executable instructions that, when executed by a device, cause the device to perform the method according to a first aspect of this application. Attached Figure Description

[0020] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0021] Figure 1 shows a schematic diagram of the data structure for floating-point numbers provided in some embodiments of this application;

[0022] Figure 2 shows a schematic diagram of the structure of MX format data provided in some embodiments of this application.

[0023] Figure 3 illustrates a schematic diagram of an example environment in which the devices and / or methods of embodiments of this application may be implemented;

[0024] Figure 4 shows a schematic flowchart of some embodiments of this application for quantizing floating-point numbers;

[0025] Figure 5 shows a schematic diagram of a data quantization process provided by some embodiments of this application;

[0026] Figure 6 illustrates a method for determining private elements according to some embodiments of this application;

[0027] Figure 7 shows a schematic diagram of a floating-point quantization process provided by some other embodiments of this application;

[0028] Figure 8 illustrates application scenarios of floating-point quantization according to some embodiments of this application; and

[0029] Figure 9 shows a schematic block diagram of an apparatus for quantizing floating-point numbers according to an embodiment of the present application. Detailed Implementation

[0030] The technical solutions of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments.

[0031] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " indicates "and / or," for example, A / B can mean A or B, or A and B; "and / or" in this text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Furthermore, in the description of the embodiments of this application, "plural" or "multiple" refers to two or more than two.

[0032] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "multiple" means two or more.

[0033] The terminology used in the following embodiments is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to also include expressions such as “one or more,” unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, “at least one” and “one or more” refer to one, two, or more than two. The term “and / or” is used to describe the relationship between related objects, indicating that three relationships may exist; for example, A and / or B can indicate: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character “ / ” generally indicates that the preceding and following related objects are in an “or” relationship.

[0034] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "one embodiment," "some embodiments," "another embodiment," "other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0035] The numbers or values ​​used in this specification are illustrative and are intended only to facilitate understanding of the technology of the embodiments of this application, and are by no means intended to limit the scope of this application.

[0036] Floating-point numbers are a numerical representation used in computers to approximate a real number. A floating-point number consists of an integer part and a fractional part, where the position of the decimal point is not fixed but fluctuates. Floating-point numbers typically consist of three parts: the sign field, the exponent field (also called the exponent field), and the mantissa field. The sign bit indicates whether the floating-point number is positive or negative (e.g., 0 for positive, 1 for negative). The mantissa represents the significant digits of the floating-point number; the number of bits in the mantissa determines the precision of the floating-point number—more bits result in higher precision. The exponent represents the position of the decimal point; the sign of the exponent indicates the direction of the decimal point's movement (positive for right, negative for left). The number of bits in the exponent determines the range of the floating-point number; a longer exponent allows for a wider range of representations.

[0037] Floating-point numbers come in various types, such as single-precision (SP) floating-point numbers, double-precision (DP) floating-point numbers, extended single-precision floating-point numbers, and extended double-precision floating-point numbers. Different types of floating-point numbers have different bit widths. The total bit width of a floating-point number can include the sign bit, the entire bit width of the exponent offset, and the entire bit width of the mantissa. For example, a single-precision floating-point number has a bit width of 32 bits (FP32), a double-precision floating-point number has a bit width of 64 bits (FP64), and a half-precision floating-point number has a bit width of 16 bits (FP16). As an example, Figure 1 shows the data structure of a half-precision floating-point number (FP16), which has a total bit width of 16 bits. Bit 0 is the sign bit, bits 1 to 5 represent the exponent value, and bits 6 to 15 represent the mantissa value.

[0038] In each precision of floating-point data representation, the bit width of each field (e.g., exponent field, mantissa field, etc.) contained in the floating-point number is fixed. Table 1 below shows the data structure of floating-point numbers with different precisions. As shown in Table 1, FP16 may include a 1-bit sign field, a 5-bit exponent field, and a 10-bit mantissa field; FP32 may include a 1-bit sign field, an 8-bit exponent field, and a 23-bit mantissa field; FP64 may include a 1-bit sign field, an 11-bit exponent field, and a 52-bit mantissa field.

[0039] Table 1. Data structure table for floating-point numbers of different precisions

[0040] Floating-point numbers have wide applications in scientific computing, image processing, and neural networks (such as deep learning networks and AI). Taking the application of floating-point numbers in AI training as an example, existing floating-point formats either have insufficient numerical range or insufficient precision, affecting convergence speed and model performance. On the other hand, satisfying both the requirements for numerical range and precision would result in excessive memory usage, leading to excessive overhead for data storage and data transfer.

[0041] Based on this, the industry has introduced a new floating-point format such as Microscaling (MX). MX format floating-point numbers can support lower bit widths for AI training and inference, and require less memory. Data formats conforming to the MX standard can achieve robust model accuracy for AI training and inference using 8 bits or less. Specifically, MX format is a block data format, where several data points can form a block (or a group), with data distributed in blocks. MX format data consists of three parts, which may include, for example, the block size k (how many low-bit data points form a block), a shared scaling factor X, and a private element P. i All k elements (P) i Elements have the same data type and therefore the same bit width. A scaling factor X is applied to all k elements. The data type and scaling factor of each element can be chosen independently. In a sense, MX can be viewed as a mechanism for constructing vector data types from scalar data types. As an example, Figure 2 shows the data structure of floating-point numbers in MX format. As shown in Figure 2, S, E, and M are used to represent the values ​​of the sign, exponent, and mantissa fields of a floating-point number, respectively. The shared scaling factor X is a scaling factor applied to the entire data block, determining the dynamic range of all elements in the block. By introducing the shared scaling factor, MX format data can flexibly represent different ranges of data while maintaining a low bit width. The block size k refers to the number of low-bit data elements that make up a data block (or group). Private element P i This refers to each low-bit data element in a data block. These elements, adjusted by the scaling factor X, collectively represent a high-precision floating-point number or integer. MX data can be of various types, such as MXFP8, MXFP4, MXFP16, MXINT4, etc.

[0042] As mentioned above, while existing MX format floating-point numbers can maintain high model accuracy with low bit width, the quantization process of converting scalar floating-point data to low-bit MX format suffers from low accuracy and excessive latency. For example, in determining the shared scaling factor corresponding to the MX format, an indiscriminate floor function is used, which is mathematically equivalent to indiscriminately truncating the exponent of the floating-point number. For floating-point numbers with a large mantissa, there is a significant truncation error. Thus, quantization using a predetermined shared scaling factor results in large quantization errors and low quantization accuracy.

[0043] For example, in related technologies, the OCP quantization algorithm can be used to determine the shared scaling factor and the quantized MX data. The main process of OCP quantization is as follows: Step 1: Calculate the maximum value in each block size of the quantized scalar floating-point format; Step 2: Calculate using the formula... (Including logarithmic functions) Calculate shared_exp; Step 3: Calculate X using X = 2shared_exp (including exponential functions); Step 4: For each scalar floating-point format data, use P i =quantize_to_element_format(V i / X) calculate the private element P i When the calculation result exceeds the representable range of the MX format, the maximum representable value is taken, and the sign is preserved. Therefore, scalar floating-point numbers can be represented... Convert to MX data blocks

[0044] Where X is the shared scaling factor, emax elem This represents the maximum binary offset of the MX data. During this process, the OCP quantization algorithm uses indiscriminate floor division (equivalent to indiscriminately truncating the exponent bits of the floating-point number) when calculating the shared exponent `shared_exp`. In other words, it only considers the exponent bits of the floating-point number when determining the shared scaling factor `X`, ignoring the influence of the mantissa bits. This results in a significant truncation error for floating-point numbers with large mantissa bits, leading to a substantial quantization error in the MX data.

[0045] Therefore, to address the problem of low quantization accuracy in related technologies, embodiments of this application provide a method for quantizing floating-point numbers. This method includes obtaining statistical values ​​corresponding to multiple source floating-point numbers, as well as the exponent and mantissa of the statistical values. The method further includes determining a target scaling factor based on a preset number of mantissa bits in the exponent and mantissa. Furthermore, the method includes determining multiple target floating-point numbers by quantizing the multiple source floating-point numbers based on the target scaling factor.

[0046] By considering not only all exponent bits of the floating-point number but also the mantissa bits of the preset scalar number in this way when determining the target scaling factor for quantization, the numerical characteristics of the floating-point number can be captured more comprehensively. This reduces truncation errors during quantization and determines a more accurate target scaling factor. A more accurate target scaling factor not only balances the precision and representation range of the floating-point number, ensuring that the quantized floating-point number maintains high precision without exceeding its representable range, but also results in higher precision for subsequent floating-point numbers quantized using the target scaling factor. In other words, the conversion error between the floating-point numbers before and after quantization is smaller, and the data conversion efficiency and accuracy are higher. Furthermore, the target scaling factor is determined based on the statistical characteristics of multiple floating-point numbers, ensuring that all floating-point numbers are appropriately represented. This avoids both data loss and data overflow.

[0047] The method provided in this application embodiment can be applied to the field of neural networks, such as the field of AI. This quantization method can be applied to electronic devices, such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), and other terminal devices. It can also be a server (a single server or a server cluster consisting of multiple servers) or a cloud computing service center. This application embodiment does not impose any restrictions on the specific type of electronic device.

[0048] Figure 3 illustrates a schematic diagram of an example environment 300 in which the device and / or method according to embodiments of this application may be implemented. As shown in Figure 3, the example environment 300 includes electronic device 302 and electronic device 304. Electronic device 302 can communicate with other electronic devices via a network (wired or wireless, such as via wireless local area networks (WLAN) (e.g., wireless fidelity, Wi-Fi)) to send acquired or stored source floating-point numbers to electronic device 302, whereby electronic device 302 quantizes the source floating-point numbers to obtain the quantized target floating-point numbers. Electronic device 302 may include various types, such as a personal computer 302-1, a mobile phone 302-2, and a server 202-3. The method provided in the embodiments of this application can be applied to the training or inference of neural network models, such as the training and inference of AI models. The neural network model can be used to process at least one type of data, including images, language, video, and text; this application does not impose any limitations on this.

[0049] In some embodiments, the source floating-point number can be a standard floating-point number type, such as FP16, FP32, or FP64, or other types of floating-point numbers such as Brain Floating Point (BF) (e.g., BF16, BF32, etc.) or Tensor Float (TF) (e.g., TF32, TF16, etc.). The source floating-point number can be a floating-point number to be quantized, such as "010001001110010". Multiple source floating-point numbers can originate from the same electronic device or from multiple electronic devices, and the types of the multiple floating-point numbers can be the same. The multiple floating-point numbers acquired by electronic device 302 can be input to electronic device 304, where electronic device 304 uses the quantization method provided in this application embodiment to quantize each of the multiple floating-point numbers to determine the quantized floating-point number.

[0050] In some embodiments, the electronic device 302 and electronic device 304 may be a smart wearable device, smartphone, smart home device, tablet computer, laptop computer, desktop computer, in-vehicle computer, or server, etc., with the above-described functions. It may be a single server, a server cluster consisting of multiple servers, or a cloud computing service center, etc., and this application embodiment does not specifically limit this.

[0051] The following is a flowchart of a method 400 for migrating data according to an embodiment of this application, with reference to FIG4. It should be noted that the method 400 according to an embodiment of this application can be implemented, for example, at the electronic device 404 shown in FIG4. It should be understood that the method 400 may also include additional actions not shown, and the scope of this application is not limited in this respect.

[0052] As shown in Figure 4, at step 402, the statistical values, exponents, and mantissas of multiple source floating-point numbers are retrieved. These multiple source floating-point numbers can be multiple floating-point numbers that need to be quantized. The types of these multiple source floating-point numbers can be the same, such as single-precision floating-point numbers, double-precision floating-point numbers, octal-precision floating-point numbers, or even tensor floating-point numbers. Each type of floating-point number has its specific exponent and mantissa bits. In some examples, the multiple source floating-point numbers can be {x1, x2, x3, x4…xn}, where n is a positive integer, and x1, x2, …, xn can be floating-point numbers of the same precision type, for example, x1, x2, …, xn can be BF16. The multiple source floating-point numbers can be stored in an electronic device as an array, list, or other form. The statistical values ​​can be numerical values ​​used to characterize the statistical features of the multiple source floating-point numbers. For example, the statistical values ​​can be the maximum, average, median, mode, minimum, etc., and of course, they can also be other types of statistical values ​​such as standard deviation and mean deviation.

[0053] After determining the statistical value, the exponent and mantissa of the statistical value can be obtained. The method of obtaining them depends on the data type of the statistical value. For example, if the statistical value is also a floating-point number of the same type as multiple source floating-point numbers, the mantissa and exponent of the floating-point number can be obtained directly. If the statistical value is represented in decimal or other types of floating-point representation, the statistical value can be converted to binary floating-point representation in order to extract the exponent and mantissa.

[0054] In 404, a target scaling factor is determined based on a preset number of mantissa bits in the exponent and mantissa. The target scaling factor can be a parameter used to scale the source floating-point number, determining the quantization range of the source floating-point number. For example, the target scaling factor can be 2, 4, 8, or 16. If the quantized target floating-point number is MX data, the target scaling factor can be a shared scaling factor for MX data. In some embodiments, the preset number can be determined by the user according to the actual quantization requirements and the format of the floating-point number; for example, it can be 1 bit, 2 bits, 3 bits, 5 bits, etc. After determining the exponent and the preset number of mantissa bits, they can be used as the basis for calculations to determine the corresponding target scaling factor.

[0055] In 406, multiple target floating-point numbers are determined by quantizing multiple source floating-point numbers based on a target scaling factor. After determining the target scaling factor, the target floating-point number corresponding to each source floating-point number can be determined by quantizing each of the multiple source floating-point numbers individually. The target floating-point number can be of various types, including block size, multiple private elements, and the target scaling factor. In some embodiments, the determination process of the private elements is related to the type of the target floating-point number, and the rounding strategy is also related to the type of the target floating-point number. For example, in some examples, the multiple floating-point numbers are {x1, x2, x3…xk}, and the corresponding private elements after quantization can be {p1, p2, p3…pk}. Here, pi = xi / target scaling factor.

[0056] In this way, the determination of the target scaling factor considers not only the exponent of the floating-point number but also some of its mantissas. This results in a more accurate target scaling factor, leading to more accurate data quantized using the target scaling factor. This approach reduces precision loss during quantization and improves the accuracy of the quantization results.

[0057] In some embodiments, the source floating-point numbers can be floating-point numbers corresponding to the weight parameters of the neural network model, the activation function values ​​of the neural network model, or the gradients of the neural network model. Since the most important factor in quantizing multiple source floating-point numbers is the shared scaling factor, which is the scaling value of the entire data block and determines the dynamic range of all elements in the data block, the basis for determining the shared scaling factor should be determined collectively by the elements in the entire data block, for example, by the statistical characteristics corresponding to multiple source floating-point numbers. In this way, the determined shared scaling factor can balance the precision and representation range of the data, avoiding data loss or overflow. Based on this, a shared scaling factor determined according to the statistical values ​​corresponding to multiple floating-point numbers is more reasonable and accurate.

[0058] As described above, the related technologies do not consider the mantissa of the floating-point number when determining the shared exponent `shared_exp`. Therefore, the determined shared exponent cannot maximize the utilization of the exponent range of the MX format. Thus, the method provided in this application can determine whether the mantissa of the floating-point number is large enough to affect the value of the shared exponent (e.g., whether to add 1 to the shared exponent). This allows for dynamic utilization of the exponent range of the MX format under different circumstances, resulting in a smaller error between the quantized MX data and the corresponding floating-point number. As understood by those skilled in the art, the statistical values ​​corresponding to multiple floating-point numbers better reflect the statistical characteristics of multiple floating-point numbers (e.g., the maximum value of multiple floating-point numbers determines the exponent range of the MX format to a certain extent). For example, in one example, to determine a more accurate shared scaling factor, it can be determined whether to add 1 during the calculation of the shared scaling factor based on whether the maximum value is numerically closer to the next power of 2 (whether the mantissa is greater than a set value). For example, it can be determined whether the value is closer to the next power of 2 based on the exponent and part of the mantissa of the maximum value. If the floating-point number is closer to the next power of 2, then an addition operation is performed.

[0059] Based on this, a shared scaling factor can be determined using the exponent of the maximum value of multiple floating-point numbers and a portion of the mantissa bits. The portion of the mantissa bits can be the higher-order mantissa bits, such as the first n mantissa bits, where n can be determined by the user according to actual quantization requirements (e.g., it can be set to 1, 2, or 3, etc.). For example, after obtaining a preset number, the mantissa portion can be split according to the preset number to determine the required higher-order mantissa bits. It can be understood that the splitting method can be determined based on the preset number, which typically refers to the number of bits retained in the higher-order mantissa bits. For example, if the preset number is 3 bits, the mantissa portion can be split into the first 3 bits as the higher-order mantissa bits, and the remaining part as the lower-order mantissa bits.

[0060] In this way, the determined shared scaling factor is more accurate and more in line with the actual situation, thereby reducing the conversion error during data format conversion, improving the efficiency of data conversion, and improving the training accuracy of the training model.

[0061] In some embodiments, the shared scaling factor can be determined by performing bit-level operations on the target operation value and the statistical value. These bit-level operations can be basic algorithms performed at the bit level, such as bitwise AND operations, addition operations, subtraction operations, left shift operations, and right shift operations. For example, a bitwise AND operation results in a 1 if all corresponding bits are 1, otherwise it results in a 0. That is, a bitwise AND operation can be used to check if certain bits are 1 (e.g., to check if the higher-order mantissa is 1). A left shift operation can shift all bits of a binary number to the left by a specified number of bits, padding the right side with zeros. Similarly, a right shift operation can shift all bits of a binary number to the right by a specified number of bits; for unsigned numbers, zeros are padding the left side; for signed numbers, an arithmetic right shift based on the sign bit may be performed (padding the left side with the value of the sign bit).

[0062] For example, by performing a logical AND operation on the target operand and the statistical value, it can be determined whether the higher-order mantissas of the statistical value affect the value of the sharing exponent. In one example, if all higher-order mantissas are 1 or the number of higher-order mantissas is greater than a threshold, the value of the sharing exponent can be adjusted, and the corresponding sharing scaling factor can be determined based on the adjusted sharing exponent. In this way, by using simple bit operations, the time complexity of calculating the sharing scaling factor can be significantly reduced, thereby improving the efficiency of determining the sharing scaling factor. This avoids the need to use basic elementary functions (exponential, logarithmic) to reduce the computational latency of the sharing scaling factor.

[0063] In some embodiments, after determining the shared scaling factor, the quantized MX data (e.g., determining the block size k and individual private elements of the MX data) can be determined based on the shared scaling factor and an array of multiple floating-point numbers. Here, since floating-point numbers in MX format are essentially in the form of data blocks, similar to a data set composed of multiple data points, the block size k is the number of floating-point elements contained in each MX data block. Based on this, the number of multiple source floating-point numbers is consistent with the k value of the target floating-point number. For example, if the multiple source floating-point numbers are 32 BF16 floating-point numbers, the size k of the target floating-point number, i.e., the floating-point number block in MX format, is 32. After determining the shared scaling factor X, the floating-point number can be divided by the shared scaling factor X to determine the private elements corresponding to that floating-point number. This process ensures that all floating-point numbers are scaled to a suitable range so that they can be efficiently stored in MX format data blocks. In some embodiments, P can be utilized... i =CAST(V i / X) Determine the floating-point number V i The corresponding private element P i The CAST operation refers to the process of converting one data type to another.

[0064] In some embodiments, during quantization, to simplify the algorithm, if the floating-point number is small—for example, if the exponent part of the floating-point number is less than the smallest normalized exponent—the private element corresponding to that floating-point number can be set to 0. Here, a denormalized number (sometimes called a subnormalized number or non-standard number) is a special representation that can be used to represent numbers whose absolute values ​​are so small that they cannot be represented using a normal floating-point format. For example, the exponent of a denormalized number can be all 0, and the mantissa does not implicitly contain leading 1s.

[0065] Exemplary, with reference to FIG5, a floating-point quantization process 500 provided in an embodiment of this application will be described. FIG5 shows a schematic diagram of a data quantization process 500 provided in some embodiments of this application. The example process 500 may be an example implementation of method 400 and is executed, for example, by the electronic device 304 shown in FIG1. ​​It should be understood that the example process 500 may also include additional actions not shown, or some of the actions may be omitted. Furthermore, the order of actions shown in the example process 500 is only an example, and in some other embodiments, the execution order of actions may be changed without departing from the scope of this application. As shown in FIG5, the maximum value 502 (statistical value) of k floating-point numbers can be "0100001001110010" (decimal 60.5). In some embodiments, in order to obtain the mantissa and exponent of the maximum value 502, a "logical AND" operation can be performed by combining the maximum value 502 with a mask. The logical AND operation is a basic logical operation that compares two or more binary numbers. The result is set to 1 only if all bits of the compared numbers are 1; otherwise, the result is set to 0.

[0066] In some embodiments, the mask can be determined based on the type of the floating-point number and the size of a preset number of bits. The mask is used to extract all the exponent bits and the mantissa bits of the maximum value 502. For example, in the case of a floating-point number BF16, it includes a 1-bit sign field, an 8-bit exponent field, and a 7-bit mantissa field. That is, bit 0 is the sign bit, bits 1 to 8 are the exponent bits, and bits 9 to 15 are the mantissa bits. When the preset number is 2 bits, it can be determined that the bits to be extracted are all the exponent bits and the first two mantissa bits. Based on this, the mask can be determined as "0111111111100000", which is converted to hexadecimal as "0X7FE0". That is, by performing an "AND" operation on the maximum value 502 and "0X7FE0", all the exponent bits (10000100) and the first two mantissa bits (11) of the maximum value 502 can be determined (the gray part of the target statistical value 504 in Figure 5).

[0067] In essence, a mask can be used to extract specific bits from floating-point numbers. For example, by setting the bits to be retained to 1 and the bits to be filtered or cleared to 0, the effect of retaining / clearing specific bits can be achieved. It should be noted that the size of the mask varies depending on the type of the source floating-point number (e.g., precision type) and the preset number of bits. For example, when the source floating-point number is FP16 (containing a 1-bit sign field, a 5-bit exponent field, and a 10-bit mantissa field) and the preset number of bits is 2, the mask can be "0111111100000000".

[0068] In some embodiments, subsequent bit-level operations can be directly performed on the maximum value 502 and the target operation value to determine the corresponding shared scaling factor. In other embodiments, to save computing resources and improve the calculation efficiency of the shared scaling factor, after determining all exponent bits and the mantissa bits of a preset number of bits, a corresponding target statistical value 504 can be determined (it can be understood that the exponent bits and the mantissa bits of the target statistical value are the same as the statistical value, such as the maximum value 502, and the remaining positions can be 0). That is, after determining all exponent bits and the mantissa bits of the preset number of bits of the maximum value 502, the target statistical value 504 can be determined by setting the values ​​of other bits to 0 (for example, it can be "0100001001100000"). The corresponding shared scaling factor is determined by performing subsequent bit-level operations on the target statistical value 504 and the target operation value. This target statistical value 504 can be used to filter out key bits in subsequent bit-level operations. Through the operation of the target statistical value 504 and the target operation value, key bits used to determine the shared scaling factor can be extracted. The target operation value can be a reference factor for determining whether to adjust the value of the shared scaling factor. For example, the target operand value can be used to determine whether the high-order mantissa of a floating-point number affects the value of the shared scaling factor. If the high-order mantissa of the floating-point number is large (e.g., all mantissas of the preset quantity bits are 1), the target operand value can be rounded up when determining the shared scaling factor, resulting in higher quantization precision and smaller quantization error. If the high-order mantissa of the floating-point number is small (e.g., not all mantissas of the preset quantity bits are 1), the shared scaling factor value can be determined without adjusting it, and floating-point quantization can proceed normally. Based on this, the target operand value can be determined based on the mantissas of all exponent bits and the preset quantity bits. For example, the target operand value can be determined by setting the last bit of the preset quantity bits to 1 and the remaining bits to 0 (e.g., "0000000000100000"). In other words, the value of the shared scaling factor is influenced by retaining the mantissas of the preset quantity bits and determining whether to carry over to the exponent. For example, in some examples, when the last digit of the preset quantity is all 1 (e.g., 11), the exponent of the maximum value will be incremented by 1; while when the last digit of the preset quantity is not all 1 (e.g., 00, 01, 10), the exponent of the maximum value will not change.

[0069] As shown in Figure 5, the first candidate scaling factor (scale1) 506 ("0100001010000000", corresponding to hexadecimal 0X20) can be determined by performing an "add" operation on the target statistical value 504 and the target operation value (e.g., "0000000000100000"). The addition operation refers to the addition operation; in binary data, binary addition starts from the least significant bit and proceeds bit by bit towards the most significant bit. After determining the first candidate scaling factor 506, the second candidate scaling factor (scale2) 508 ("0100000110010010") can be determined through value range mapping. For example, the second candidate scaling factor 508 can be determined by performing a "subtract" operation on the first candidate scaling factor 506 and the value range mapping parameter. The value range mapping parameter can be determined based on the floating-point format type, MX format type, and a preset number of mantissa bits. For example, with a maximum value of 502 being BF16 and a preset quantity of 2 bits, the value range mapping parameter is determined to be "0X100". In some embodiments, the value range mapping parameter can be the offset of the maximum binary value (emax_elem) of the MX data format.

[0070] In some embodiments, because the MX data format may contain various floating-point types (such as FP8, FP6, FP4, etc.), and each type may have different exponent bit widths and offset settings, the range mapping parameters will also differ. For example, if the exponent of a floating-point number is an n-bit binary number, then the maximum exponent value it can represent is 2^n. n -1 (assuming offset binary representation is used, and the case where the offset is 0 is not considered). In other words, the range mapping parameter can be a technical parameter related to the representation range of floating-point elements in the MX data format. It determines the maximum exponent value that a floating-point element can represent, and thus affects the representation range and precision of the floating-point number.

[0071] As shown in Figure 5, after determining the second candidate scaling factor 508, the shared scaling factor 510 ("0000000010000011") can be determined by right shifting. The number of shifts can be determined based on the floating-point number format and the MX data format, for example, based on the number of bits in the mantissa field corresponding to the maximum value. For instance, the shared scaling factor 510 can be determined by right-shifting the second candidate scaling factor 508 by 7 bits. In some embodiments, the shared scaling factor is converted to decimal 16, then the shared scaling factor X = 16.

[0072] This approach optimizes the calculation of the shared scaling factor, thereby reducing precision loss during quantization and improving the accuracy of the quantization results.

[0073] Figure 6 illustrates a method for determining private elements according to some embodiments of this application. As shown in Figure 6, in the process of converting the floating-point number 602 (0100000001110010) of BF16 to MXFP4, the private element P of the quantized MXFP4 can be determined by dividing the floating-point number of BF16 by the shared scaling factor 16 determined according to the above embodiments. In the process of determining the private element 604 of the quantized MXFP4 data, the corresponding private element 604 of the quantized MXFP4 data can be determined by a CAST operation. In some embodiments, since MX data can only represent data within a certain range. Taking MXFP4 as an example, MXFP4 can only represent 0, 0.5, 1, 1.5, 2, 3, 4, and 6. For example, as shown in Table 2 below, the actual value corresponding to 0111 is 6, the actual value corresponding to 0110 is 4, and the actual value corresponding to 0001 is 0.5.

[0074] Table 2 Data Range Table for MXFP4

[0075] Therefore, when converting floating-point numbers to MX data, a rounding operation is required. For example, 7.526 can be rounded to 6. During this process, the CAST operation can include a predefined rounding strategy. As shown in Figure 6, the private element of the quantized MX data corresponding to the floating-point number 60.5 can be determined to be 4 based on P = CAST(V / X) = CAST(60.5 / 16) = CAST(3.78125) = 4. The shared scaling factor of the quantized MX data is 16, and the restored floating-point number is 16 * 4 = 64. It can be seen that the quantization method provided in this embodiment has a smaller quantization error and higher quantization accuracy.

[0076] In addition, in practical applications, the floating-point array to be quantized is X = [0.,1.,2.,3.,4.,5.,6.,7.,8.,9.,10.,12.,13.,14.,15.,16.,17.,18.,19.,20.,21.,22.,23.,24.,25.,26.,27.,28.,29.,30.,31.]. The quantized array Y = [0.,2.,2.,4.,4.,6.,6.,8.,8.,12.,12.,12.,12.,12.,12.,16.,16.,16.,16.,16.,16.,16.,16.,24.,24.,24.,24.,24.,24.,24.,24.,24.,24.,24.,24.,24. The quantized array Z = [0.,0.,4.,4.,4.,4.,4.,4.,8.,8.,8.,12.,12.,12.,12.,16.,16.,16.,16.,16.,16.,16.,16.,16.,24.,24.,24.,24.,24.,24.,24.,24.] is obtained using the embodiments of this application. By calculating the norm distance between the quantized array and the array to be quantized (||XZ||1 = 48 < 56 = ||XY||1, ||XZ||2 = 10.58 < 14.14 = ||XY||2, etc.), it can be seen that the quantization method provided by some embodiments of this application has smaller quantization error and higher quantization accuracy.

[0077] For example, a floating-point quantization process 700 provided in some embodiments of this application will be described with reference to FIG7. FIG7 shows a schematic diagram of a floating-point quantization process 700 provided in some embodiments of this application. It should be understood that the example process 700 may also include additional actions not shown, or some of the actions may be omitted. As shown in FIG7, the multiple source floating-point numbers to be quantized are of type BF16, and the statistical value corresponding to the multiple source floating-point numbers is 0000001101100100. Based on this, a mask can be determined according to the floating-point number type BF16. All exponent bits and the first two mantissa bits of the mask can be 1, and the remaining bits can be 0. The purpose is to extract all exponent bits and the first two mantissa bits of the statistical value to determine the corresponding target statistical value 0000001101100000.

[0078] As shown in Figure 7, after determining the target statistical value, the target operation value can be determined based on the above method. For example, based on the floating-point number format and preset quantity, the target operation value can be determined as 0000000000100000. That is, all other bits can be set to 0, and the last bit of the first two mantissa bits can be set to 1. By adding the target operation value and the target statistical value, the first candidate scaling factor 0000001110000000 is determined. Similarly, by subtracting the range mapping parameter 0000000110000000 from the first candidate scaling factor, the second candidate scaling factor 0000001000000000 can be determined. By shifting the second candidate scaling factor, the final shared scaling factor 0000001000000100 can be determined.

[0079] In some embodiments, to achieve low-bit training and inference of the neural network model, the floating-point numbers corresponding to the weight parameters or other parameters of the neural network model can be converted into low-bit MX format. Based on this, the source floating-point numbers can be the floating-point numbers corresponding to the weight parameters in the neural network model (e.g., a natural language processing model). In some embodiments, statistical values ​​corresponding to multiple weight floating-point numbers can be determined, such as the maximum value or median. Taking the maximum value as an example, after determining the maximum value, the exponent and mantissa parts of the maximum value can be determined through simple bit operations. For example, determining how many bits the exponent part of the maximum value contains, and the value of each bit, and how many bits the mantissa part contains, and the value of each bit. It can be understood that after determining the mantissa and exponent parts of the maximum value, the total exponent plus a preset number of mantissa bits can be determined. The preset number can be determined by the user according to the actual training requirements of the neural network model. Based on the total exponent plus the preset number of mantissa bits, a shared scaling factor (target scaling factor) can be determined for quantizing the multiple weight floating-point numbers. Based on this shared scaling factor, the weight floating-point numbers can be quantized to obtain the corresponding MX data. Because MX data has a lower bit width and occupies less memory, it can effectively reduce the computational intensity, parameter size, and memory consumption of neural network models.

[0080] Figure 8 illustrates application scenarios of floating-point quantization according to some embodiments of this application. As shown in Figure 8, the neural network model may include multiple attention layers and fully connected layers. The attention layers can focus information on the input data and determine the importance of different input information in the output. The fully connected layers are usually located in the last few layers of the neural network and are used to combine and classify the features extracted by the previous layers. In the fully connected layer, each neuron is connected to all neurons in the previous layer. The fully connected layer includes two parts: forward propagation (calculating the weighted sum of the input data and generating the output through the activation function) and backpropagation (updating the weight values ​​according to the gradient of the loss function during training).

[0081] As shown in Figure 8, the floating-point data quantization method provided in this application embodiment can be applied to the training process of a neural network model, where the neural network model can be an image classification model, a natural language processing model, an autonomous driving model, a medical image recognition model, a speech recognition model, etc. During the training process of the neural network model, the model's parameters and gradients can use the same floating-point number format, such as FP16 data or FP32 data. In other embodiments, the model's parameters and gradients can also use different floating-point number formats, for example, the model's parameters can use FP16 data and the model's gradients can use FP32 data. In some embodiments, the training process of the neural network model may include, but is not limited to, model weight parameter initialization, forward computation, backward computation, weight update process, and multi-machine / multi-card data communication process. For example, it can be applied to the quantization of the Attention layer or the quantization of the linear layer. During the quantization process of the Attention layer, the output of the activation function can be quantized. For example, the quantization method provided in this application embodiment can be used to quantize the FP32 output of the activation function to MXFP4, or it can be quantized to MXFP8. During the quantization process of the linear layer, the weight values ​​can be quantized. In some embodiments, the quantization method provided in this application can be applied not only to the training of neural network models but also to the inference process of neural networks. Of course, the quantization method provided in this application can also be applied to other scenarios, and no limitations are imposed here.

[0082] The following experimental data illustrates the improvement in quantization accuracy achieved by the data quantization methods provided in some embodiments of this application. As shown in Table 3, when the floating-point numbers to be quantized are arrays output by activation functions (e.g., 10,000 sets of BF16 data with a block size of 32), it can be seen that 9,893 sets of data obtained by the quantization methods provided in some embodiments of this application have higher quantization accuracy than the OCP quantization results, with an average accuracy improvement of 27.17%; 107 sets of data obtained by the quantization methods provided in some embodiments of this application have lower quantization accuracy than the OCP quantization results, with an average accuracy degradation of 2.55%. The overall average accuracy improvement across the 10,000 sets of experiments is 26.85%, demonstrating a significant advantage. Furthermore, when the floating-point numbers to be quantized are weight values ​​of a neural network model (weights in BF16 format), it can be seen that the quantization results obtained using the quantization methods provided in some embodiments of this application for each weight matrix have higher accuracy than the OCP quantization results, with an average accuracy improvement of 4.48%.

[0083] Table 3. Improvement in Quantization Accuracy

[0084] Figure 9 shows a schematic block diagram of a device 900 for quantizing floating-point numbers according to an embodiment of this application. The device 900 for quantizing floating-point numbers can be implemented at an electronic device 304. The device 900 for quantizing floating-point numbers includes a statistical value acquisition unit 910, a target scaling factor determination unit 920, and a quantization unit 930.

[0085] The statistical value acquisition unit 910 is configured to acquire statistical values ​​corresponding to multiple source floating-point numbers, as well as the exponent and mantissa of the statistical values. The target scaling factor determination unit 920 is configured to determine the target scaling factor based on a preset number of mantissas in the exponent and mantissa. The quantization unit 930 is configured to determine multiple target floating-point numbers by quantizing multiple source floating-point numbers based on the target scaling factor.

[0086] In some embodiments, the target scaling factor determination unit 920 is further configured to: determine a target operating value based on a preset number of mantissas in the exponent and mantissa; and determine a target scaling factor by performing bit-level operations on the target operating value and the statistical value.

[0087] In some embodiments, the target scaling factor determination unit 920 is further configured to: determine the target statistical value corresponding to the statistical value based on the exponent of the statistical value and the preset number of digits; and determine the target scaling factor by performing bit-level operations on the target operational value and the target statistical value.

[0088] In some embodiments, the apparatus 900 further includes a target operation value determination unit, configured to: split the mantissa portion of a statistical value into a high-order mantissa and a low-order mantissa, wherein the splitting method is determined according to a preset quantity; and determine a target operation value based on the exponent and the high-order mantissa.

[0089] In some embodiments, the target scaling factor determination unit 920 is further configured to: determine a first candidate scaling factor based on the sum of a target operational value and a target statistical value; determine a range mapping parameter based on the format of the source floating-point number and the format of the target floating-point number; and determine a target scaling factor based on the range mapping parameter and the first candidate scaling factor.

[0090] In some embodiments, the target scaling factor determination unit 920 is further configured to: determine a second candidate scaling factor based on the range mapping parameters and the first candidate scaling factor; and determine the target scaling factor by shifting the second candidate scaling factor, wherein the number of shifts is determined according to the format of the source floating-point number.

[0091] In some embodiments, determining the second candidate scaling factor based on the range mapping parameter and the first candidate scaling factor includes: determining the calculation result by subtracting the range mapping parameter and the first candidate scaling factor; and determining the second candidate scaling factor based on the calculation result.

[0092] In some embodiments, the statistical value acquisition unit 910 is further configured to: determine the corresponding parameters based on the format of the source floating-point number; and acquire the exponent and mantissa corresponding to the statistical value based on the AND value of the statistical value and the parameters.

[0093] In some embodiments, the statistical value acquisition unit 910 is further configured to include: determining the maximum value among a plurality of source floating-point numbers; and using the maximum value as the statistical value corresponding to the plurality of source floating-point numbers.

[0094] In some embodiments, the plurality of source floating-point numbers include a first source floating-point number and a second source floating-point number, and the quantization unit 930 is further configured to: determine a first target floating-point number corresponding to the first source floating-point number based on a target scaling factor; and determine a second target floating-point number corresponding to the second source floating-point number based on a target scaling factor.

[0095] In some embodiments, the quantization unit 930 is further configured to: determine a private element corresponding to a first target floating-point number based on a target scaling factor and a first source floating-point number; and determine a first target floating-point number based on the private element and a target scaling factor.

[0096] In some embodiments, bit-level operations include basic arithmetic performed at the bit level, where basic arithmetic includes one or more of logical AND, addition, subtraction, and shifting.

[0097] In some embodiments, the device 900 is applied to a neural network model for processing at least one type of data, including images, speech, video, and text.

[0098] This application may be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this application.

[0099] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0100] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0101] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing the status information of the computer-readable program instructions. These electronic circuits can execute the computer-readable program instructions to implement various aspects of this application.

[0102] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0103] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0104] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0106] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for quantizing floating-point numbers, comprising: Obtain statistical values ​​corresponding to multiple source floating-point numbers, as well as the exponent and mantissa of the statistical values; The target scaling factor is determined based on the index and the preset number of digits in the last digit. as well as Based on the target scaling factor, multiple target floating-point numbers are determined by quantizing the multiple source floating-point numbers.

2. The method according to claim 1, wherein determining the target scaling factor based on a preset number of digits in the exponent and the mantissa comprises: The target operation value is determined based on the index and the preset number of mantissas in the mantissa. as well as The target scaling factor is determined by performing bit-level operations on the target operational value and the statistical value.

3. The method according to claim 2, wherein determining the target scaling factor by performing bit-level operations on the target operational value and the statistical value includes: Based on the index of the statistical value and the preset number of digits, determine the target statistical value corresponding to the statistical value; The target scaling factor is determined by performing bit-level operations on the target operational value and the target statistical value.

4. The method according to claim 2 or 3, wherein determining the target operation value based on a preset number of mantissas in the exponent and the mantissa includes: The last digit of the statistical value is split into a high-order last digit and a low-order last digit, and the splitting method is determined according to the preset quantity; as well as The target operation value is determined based on the exponent and the high-order mantissa.

5. The method according to claim 3, wherein determining the target scaling factor by performing bit-level operations on the target operational value and the target statistical value includes: Based on the sum of the target operational value and the target statistical value, a first candidate scaling factor is determined; Based on the format of the source floating-point number and the format of the target floating-point number, determine the range mapping parameters; as well as The target scaling factor is determined based on the value range mapping parameters and the first candidate scaling factor.

6. The method according to claim 5, wherein determining the target scaling factor based on the range mapping parameter and the first candidate scaling factor comprises: Based on the value range mapping parameters and the first candidate scaling factor, a second candidate scaling factor is determined; as well as The target scaling factor is determined by shifting the second candidate scaling factor, wherein the number of shifts is determined according to the format of the source floating-point number.

7. The method according to claim 6, wherein determining the second candidate scaling factor based on the range mapping parameter and the first candidate scaling factor comprises: The calculation result is determined by subtracting the value range mapping parameter from the first candidate scaling factor. as well as Based on the calculation results, a second candidate scaling factor is determined.

8. The method according to claim 1, wherein obtaining the exponent and mantissa of the statistical value comprises: Based on the format of the source floating-point number, determine the corresponding parameters; as well as Based on the statistical value and the AND value of the parameter, obtain the exponent and tail number corresponding to the statistical value.

9. The method according to claim 1, wherein obtaining the statistical values ​​corresponding to the plurality of source floating-point numbers includes: Determine the maximum value among the plurality of source floating-point numbers; as well as The maximum value is used as the statistical value corresponding to the plurality of source floating-point numbers.

10. The method according to claim 1, wherein the plurality of source floating-point numbers includes a first source floating-point number and a second source floating-point number, and determining the plurality of target floating-point numbers by quantizing the plurality of source floating-point numbers based on the target scaling factor includes: Based on the target scaling factor, determine the first target floating-point number corresponding to the first source floating-point number; as well as Based on the target scaling factor, the second target floating-point number corresponding to the second source floating-point number is determined.

11. The method according to claim 10, wherein determining the first target floating-point number corresponding to the first source floating-point number based on the target scaling factor comprises: Based on the target scaling factor and the first source floating-point number, determine the private element corresponding to the first target floating-point number; as well as The first target floating-point number is determined based on the private element and the target scaling factor.

12. The method of claim 1, wherein bit-level operations include basic arithmetic performed at the bit level, the basic arithmetic including one or more of logical AND, addition, subtraction, and shifting.

13. The method according to claim 1, wherein the method is applied to a neural network model, the neural network model being used to process at least one type of data, including images, language, video, and text.

14. An electronic device comprising: Processing unit and memory, The processing unit executes instructions in the memory, causing the electronic device to perform a method, the method comprising: Obtain statistical values ​​corresponding to multiple source floating-point numbers, as well as the exponent and mantissa of the statistical values; Based on the index and the predetermined number of digits in the last digit, a target scaling factor is determined; and Based on the target scaling factor, multiple target floating-point numbers are determined by quantizing the multiple source floating-point numbers.

15. An apparatus for quantizing floating-point numbers, comprising: The statistical value information acquisition unit is configured to acquire statistical values ​​corresponding to multiple source floating-point numbers, as well as the exponent and mantissa of the statistical values; The target scaling factor determination unit is configured to determine the target scaling factor based on a preset number of mantissas in the exponent and the mantissa; as well as The quantization unit is configured to determine multiple target floating-point numbers by quantizing the multiple source floating-point numbers based on the target scaling factor.

16. A computer-readable storage medium having stored thereon one or more computer instructions, wherein one or more computer instructions, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 13.

17. A computer program product comprising machine-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 13.