Floating-point number processing method and system applied in large model operation and application

By dynamically adjusting the number of effective bits of the mantissa of floating point numbers in large model operations, the problem of floating point numbers adjustment in the existing technology is solved, and the operation speed is improved and time-consuming is reduced, which is suitable for various computing scenarios.

CN119960723APending Publication Date: 2025-05-09SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410633230.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively adjust the accuracy of floating point numbers in large-scale model operations to improve operation speed and reduce time consumption, and there are problems of increased accuracy conversion and calculation amount.

Method used

By adding the on or off status bits of the floating-point format function to the processor status register, inserting the modified NaN number as the indicator bit value before the floating-point operation, dynamically adjusting the number of effective bits of the floating-point mantissa, thereby selecting a simplified processing method.

Benefits of technology

It realizes the improvement of computing speed while maintaining sufficient accuracy, reducing calculation amount, is compatible with existing software and hardware environments, and optimizes special value processing, which is suitable for various computing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004850886290000121
    Figure BDA0004850886290000121
  • Figure HDA0004850886300000011
    Figure HDA0004850886300000011
  • Figure HDA0004850886300000012
    Figure HDA0004850886300000012
Patent Text Reader

Abstract

The invention discloses a floating point number processing method applied to large model operation. The processing method comprises the following steps: step 1, adding a processor status bit in a status register of a processor to represent the opening or closing of a floating point format function; 2, querying a processor status bit before floating point operation; step 3, when the state bit of the processor is open, inserting a modified NaN number as an indication bit value before floating point operation, and indicating the number of effective bits in the mantissa of the floating point number of the current batch; and 4, selecting a processing method according to the number of effective bits in the mantissa of the floating-point number to execute floating-point operation. The invention further discloses a processing system for realizing the floating-point number processing method and application of the processing method or the processing system in large model operation and end-side operation AI model data processing, and the application scene is wide.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology and relates to a floating point number processing method, system and application used in large model calculations. Background Art

[0002] Numbers in computers are basically stored and represented with a limited number of bits, among which floating-point numbers are in the form of 1.2*10^1 (close to the scientific notation in math textbooks. Compared with floating-point numbers, there are other formats such as fixed-point numbers in computers. The numbers in this text, unless otherwise specified, refer to floating-point numbers). Since the 1990s, the de facto standards for floating-point numbers have been IEEE 754's double-precision floating-point numbers (using 64 bits to represent a floating-point number, hereinafter referred to as FP64), single-precision floating-point numbers (using 32 bits to represent a floating-point number, hereinafter referred to as FP32), and half-precision floating-point numbers added in 2008 (using 16 bits to represent a floating-point number, hereinafter referred to as FP16). The range of floating-point numbers is determined by the exponent, and the precision of the representation is determined by the mantissa. The range and precision of floating-point numbers are ranked as follows: FP64 is better than FP32, which is better than FP16. That is, as the number of bits occupied by floating-point numbers decreases, the range and precision of numbers that can be represented by this floating-point format decreases accordingly. However, as the number of floating-point bits decreases, the advantage is that the requirements for computer storage are reduced, and the speed of floating-point operations is improved under the same hardware specifications. For example, if a certain specification of hardware can complete the addition of two FP32 numbers 1 billion times in 1 second, if it is replaced with FP16, it can usually complete the addition of two FP16 numbers 2 billion times in 1 second under similar specifications.

[0003] The existing FP64 / FP32 / FP16 floating point formats are composed of three parts: 1) the sign bit indicating whether the number is positive or negative, which occupies 1 bit; 2) the mantissa indicating the significant digit, which occupies different bits depending on the different floating point formats; 3) the exponent indicating the range, which occupies different bits depending on the different floating point formats. Taking FP32 floating point numbers as an example, Figure 1 shown.

[0004] Figure 1When the three parts are combined to represent a floating point number, it can be approximately considered as ((-1)^sign)*mantissa*(2^exponent). In addition, the mantissa and exponent have different meanings corresponding to floating point numbers in certain specific value combinations: 1) +inf, -inf, positive and negative infinity, indicating that it exceeds the maximum range of the current floating point number, for example, the result of 1 divided by 0 is +inf. The result of -1 divided by 0 is -inf. 2) NaN, the abbreviation of Not a Number, means invalid, such as zero divided by zero. For FP32, when the exponent bit is equal to 127 and the mantissa is not all zero, it represents NaN. The condition that the mantissa is not all zero means that there are many NaNs. In addition, the high and low bits of the bits refer to the order from left to right, and generally the high bits are more important than the low bits.

[0005] In recent years, a single AI model often involves 1 billion or more parameters and their calculations. Many scenarios in AI are insensitive to the precision of floating-point numbers in the calculation process, that is, the precision has little effect on the final output after being reduced to a certain extent. However, the requirements for running time (or running speed) in AI have always existed, and to some extent, the requirements have become higher. For example, a certain business requires that the result be given within 100 milliseconds in order not to affect the user experience. With the iterative development of AI models, more parameters and calculations are often used, but the results still need to be given within 100 milliseconds. To solve this kind of problem, one type of technical means is to propose floating-point formats specifically for AI scenarios, such as BF16 and TF32. The common practice is to select some or all steps in the AI ​​model according to certain criteria, replace the floating-point formats involved in these steps with floating-point formats with lower precision, and then evaluate whether the output of the AI ​​model after replacement can still achieve the expected effect. As mentioned above, after the selected steps are replaced with floating-point numbers with lower precision, the running speed can be improved and the time consumption of the model can be reduced under the same hardware conditions.

[0006] In addition, when the AI ​​model is processing, it usually considers computational efficiency and treats multiple floating-point numbers as a batch. The computing unit processes a batch of data at a time, rather than a single data. Batches can be divided into operations of the same type, or operations of the same type can be divided into multiple batches, or multiple operations of different types can be regarded as one batch. For example, if you want to process 10,000 FP32 numbers for addition, that is, add a fixed value, it can be divided into 10 batches, each batch processing 1,000 FP32 numbers, or only divided into 1 batch, each processing 10,000 FP32 numbers. If the first level operation after the addition operation is multiplication, that is, multiplication by a fixed value, you can also regard addition and multiplication as a compound operation, and then divide the data into batches.

[0007] The sign bit, mantissa, and exponent in the existing floating-point format are fixed for each floating-point format. The existing technology is to replace the floating-point numbers used in certain steps with floating-point numbers with lower precision, which is usually achieved in the following way: through certain criteria, select the steps that can reduce the precision in a series of steps, modify the floating-point numbers involved in these steps to floating-point numbers with lower precision, and use these lower precision floating-point numbers when the subsequent model runs to these steps. This is the most common processing method, which can be understood as "static" to some extent. It is fixed after being determined in advance. Once the AI ​​model starts running, the precision of the floating-point number will not be easily modified. If it is necessary to select operations of different precisions for a certain step according to the precision of the floating-point number in the current running process, it is generally necessary to establish a complex strategy to prepare running paths of different precisions for the AI ​​model in advance, and select them during the running process according to the pre-established strategy. This kind of approach often adds a lot of extra calculations to the model, thereby reducing the speed of the model running, so it is rarely used. If the steps before and after the current step use floating point numbers with different precisions, you will need to convert the original precision floating point numbers to the replaced precision floating point numbers, which will add extra time. This type of conversion is always required unless all steps in a model have completed the replacement of lower precision floating point numbers.

[0008] Another disadvantage is that some floating-point operations will reduce the accuracy of the number during the operation process, or it is called precision loss. For example, the mantissa of FP32 is 23 bits, but only the upper half of the bits are valid, for example (for simplicity, it is expressed in decimal, but it is actually expressed in binary form in the computer): assuming that the mantissa can store 3 decimal digits and the exponent can store 1 decimal digit, then 1.50*10^3 minus 1.49*10^3, according to the rules of floating-point operations, the result should be 1.00*10^1. If expressed according to scientific and technological means, the result should be 1*10^1. Only 1 bit of the mantissa represents the effective value, and the two digits after the decimal point in the floating-point number are added to "make up" the position.

[0009] In non-AI fields, the requirements for the processing quantity and running speed of floating-point numbers are not as high as those for AI. For simplicity, each floating-point number can still be processed according to the number of bits in the floating-point format. Summary of the invention

[0010] In order to solve the deficiencies in the prior art, the purpose of the present invention is to provide a floating point number processing method, system and application for use in large model operations.

[0011] The present invention modifies the floating point of floating point numbers and their processing methods. After use, different steps in the AI ​​model can use the same floating point format, eliminating the conversion between floating point formats. In addition, the precision of floating point numbers can be adjusted according to batches during the operation of the AI ​​model, and the hardware can choose to simplify processing according to different precisions, thereby improving the overall operation speed and reducing the operation time.

[0012] The present invention provides a floating point number processing method applied in large model calculation, the processing method comprising the following steps:

[0013] Step 1: Add a processor status bit in the processor status register to indicate whether the floating point format function is turned on or off;

[0014] Step 2: Query the processor status bit before floating point operation;

[0015] Step 3: When the processor status bit indicates on, insert a modified NaN number as an indicator bit value before the floating-point operation to indicate the number of valid bits in the mantissa of the current batch of floating-point numbers;

[0016] Step 4: Select a processing method to perform floating-point operations according to the number of valid bits in the mantissa of the floating-point number.

[0017] In step one, the processor status bit provides a global indication of whether the floating-point format function is turned on or off; the processor status bit is set externally and / or by the processor; when the processor status bit indicates off, the processing flow in the prior art is used for processing.

[0018] In a specific implementation, one bit is used to indicate whether the floating-point format function is on or off, with a value of 0 indicating off and a value of 1 indicating on.

[0019] In step 2, the interval of querying the processor status bit is greater than the time interval of a batch processing, and subsequent floating-point operations are performed according to the query value of the processor status bit.

[0020] Before step 3, there is also a step of checking whether there is a NaN value in the input data; if there is, screening will be performed to exclude the original NaN value in the input data; if not, no screening is required and the next step is directly entered; and / or,

[0021] Check whether the input data is a special floating-point type including +inf, -inf, NaN, etc.; if it is a special floating-point type, directly process it in a preset specific way; if not, extract the significant digits of the mantissa.

[0022] In step three, one or more bits are set at any preset agreed position of the mantissa of the floating point number to indicate the number of valid bits in the mantissa of the current batch of floating point numbers.

[0023] In a specific implementation, two bits are set in the low bit of the floating point number mantissa to indicate the number of valid bits in the mantissa of the current batch of floating point numbers, and the set bit values ​​include 00, 01, 10, and 11;

[0024] A bit value of 00 indicates that all mantissa bits are valid, a bit value of 01 indicates that the first 3 / 4 mantissa bits starting from the high bit are valid, a bit value of 10 indicates that the first 1 / 2 mantissa bits starting from the high bit are valid, and a bit value of 11 indicates that the first 1 / 4 mantissa bits starting from the high bit are valid.

[0025] In step 4, the floating point operation processing methods include:

[0026] According to the indication bit, only the high-order bits indicated as valid in the mantissa of the floating-point number are calculated; and / or, the Taylor formula calculation of the specified number of terms is performed according to the indication bit; and / or, different error thresholds are set according to the indication bit to perform Newton iteration method calculation; and / or, the input is divided into one or more intervals according to the indication bit, and the activation function is calculated by piecewise interpolation.

[0027] After the floating-point operation is performed in step 4, it is determined whether the result is a special type of floating-point number; if so, the special type of floating-point number is output; if not, the number of valid bits in the output floating-point number is set, and an indicator bit is used at the end of the output data to indicate, that is, a number modified based on NaN is inserted as an indicator bit, indicating the number of valid bits of the mantissa in this batch of data; the indicator bit is obtained according to the current floating-point operation processing method.

[0028] The present invention also provides a processing system for implementing the above processing method, the processing system comprising: a status bit management module, a special NaN value processing module, a floating point operation module, a special floating point processing module, and an output control module;

[0029] The status bit management module is responsible for managing the status bit of the processor and determining whether to enable a new floating point number processing method; the status bit management module may include a status bit register implemented by hardware or software;

[0030] The special NaN value processing module is used to insert a special NaN value into the mantissa of the floating point number before processing the floating point number; and / or, identify and parse the special NaN value in the data stream and extract the effective mantissa digit information;

[0031] The floating-point operation module is used to receive the effective mantissa digit information and perform floating-point operation according to the effective mantissa digit information; the floating-point operation module can not only process standard floating-point operation, but also can perform simplified floating-point operation according to the extracted effective mantissa digit;

[0032] The special floating point number processing module is used to identify and process special types of floating point numbers (such as +inf / -inf / NaN), and directly output the special types of floating point numbers or perform corresponding specific processing as needed;

[0033] The output control module is used to insert a special NaN value at the starting position of each batch of simplified output data to carry the information of the number of significant digits in the mantissa; by controlling the output of the final floating-point number, it ensures that the next-level operation or processing unit can receive the correct data.

[0034] The present invention also provides a hardware system for implementing the above floating point number processing method, the hardware system comprises: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the above floating point number processing method is implemented.

[0035] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above floating-point number processing method is implemented.

[0036] The present invention also provides the above-mentioned floating-point number processing method, or the application of the above-mentioned processing system in large model calculations, data processing of AI models running on the terminal side, etc.

[0037] Specifically, the method of the present invention may have the following two possible application scenarios: 1) terminal AI applications, such as running an AI model (not necessarily a large model) on a mobile phone; 2) always on AI applications, such as running an AI model on a smart sensor.

[0038] The device resources corresponding to the above two categories of application scenarios are often limited. It is hoped that various approximate methods can be used as much as possible while achieving similar results. One of the existing methods of model quantization and model lightweighting is to convert floating points into fixed-point numbers before processing. The present invention can be used to convert high-precision floating points into floating points with lower precision before processing.

[0039] The beneficial effects of the present invention include:

[0040] Dynamic balance between precision and performance: Traditional floating-point number processing methods use fixed precision (such as FP32 or FP64), which will cause a waste of computing resources when processing large amounts of data that require lower precision. The present invention allows for more efficient use of computing resources by dynamically adjusting the number of significant digits of the floating-point mantissa, allowing for increased computing speed while maintaining sufficient precision.

[0041] Reducing the amount of calculation: The method of the present invention reduces the overall amount of calculation of floating-point operations and saves processor time by calculating only the significant bits in the mantissa.

[0042] Compatibility and applicability: The method of the present invention adds a bit in the existing status register of the processor to indicate whether the new floating-point format is enabled, which is compatible with the existing software and hardware environment. The method can be applied without replacing the existing hardware or significantly modifying the software.

[0043] Optimization of special value processing: This method optimizes the batch processing of floating-point numbers by inserting a special NaN value at the beginning of each batch of data to indicate the number of valid bits. This processing method improves the efficiency of special value processing (such as NaN and infinity) and provides necessary information for subsequent data processing.

[0044] Adaptability: This method provides a mechanism to dynamically select accuracy based on different computing requirements, making it applicable to a variety of different computing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0046] Figure 1 It is a schematic diagram of the floating point number format in the prior art.

[0047] Figure 2 It is a flow chart of the floating point number processing method used in large model calculation in the specific implementation mode of the present invention.

[0048] Figure 3 It is a schematic diagram of the actual process of two inputs and one output of the present invention. DETAILED DESCRIPTION

[0049] The present invention is further described in detail with reference to the following specific examples and drawings. The process, conditions, experimental methods, etc. for implementing the present invention, except for the contents specifically mentioned below, are all common knowledge and common common sense in the art and are not particularly limited by the present invention.

[0050] The present invention proposes a new floating point number format, which is obtained by modifying the mantissa in the existing floating point number format. The processor status bit is queried to indicate whether it is turned on or off. When turned on, a modified NaN is inserted at the starting position of each batch of data, and its specific bits indicate the number of valid bits of the mantissa in the current batch of floating point numbers. After that, for the batch of floating point numbers, the operation unit can select a simplified processing method according to the number of valid bits in the mantissa, thereby improving the speed of operation.

[0051] The processor status bit is a bit divided by function in the processor's status register, which records a series of states in the process of the processor executing instructions. It is not reflected in the floating-point format, but one or more bits are added to the processor's existing status register as the status bit. Current processors generally include multiple categories of registers, such as data registers used to store the input, output, and intermediate result data of operations, and another category is the status register, which is equivalent to a global indication, which can be set externally, such as setting the rounding method of floating-point operations, or it can be set by the processor during the calculation process, such as indicating whether the calculation result of the most recent instruction overflowed.

[0052] The present invention provides a floating point number processing method applied in large model calculation, the processing method comprising the following steps:

[0053] Step 1: Add a processor status bit in the processor status register to indicate whether the floating point format function is turned on or off;

[0054] Step 2: Query the processor status bit before floating point operation;

[0055] Step 3: When the processor status bit indicates on, insert a modified NaN number as an indicator bit value before the floating-point operation to indicate the number of valid bits in the mantissa of the current batch of floating-point numbers;

[0056] Step 4: Select a processing method to perform floating-point operations according to the number of valid bits in the mantissa of the floating-point number.

[0057] In step one, the processor status bit provides a global indication of whether the floating-point format function is turned on or off; the processor status bit is set externally and / or by the processor; when the processor status bit indicates off, subsequent processing is performed according to the existing processing flow.

[0058] In a specific implementation, one bit is used to indicate whether the floating-point format function is on or off, with a value of 0 indicating off and a value of 1 indicating on.

[0059] In step 2, the interval of querying the processor status bit is greater than the time interval of a batch processing, and subsequent floating-point operations are performed according to the query value of the processor status bit.

[0060] Before step 3, there is also a step of checking whether there is a NaN value in the input data; if there is, screening will be performed to exclude the original NaN value in the input data; if not, no screening is required and the next step is directly entered; and / or,

[0061] Check whether the input data is a special floating-point type including +inf, -inf, NaN, etc.; if it is a special floating-point type, directly process it in a preset specific way; if not, extract the significant digits of the mantissa.

[0062] In step three, one or more bits are set at any preset agreed position of the mantissa of the floating point number to indicate the number of valid bits in the mantissa of the current batch of floating point numbers.

[0063] In a specific implementation, two bits are set in the low bit of the floating point number mantissa to indicate the number of valid bits in the mantissa of the current batch of floating point numbers, and the set bit values ​​include 00, 01, 10, and 11;

[0064] A bit value of 00 indicates that all mantissa bits are valid, a bit value of 01 indicates that the first 3 / 4 mantissa bits starting from the high bit are valid, a bit value of 10 indicates that the first 1 / 2 mantissa bits starting from the high bit are valid, and a bit value of 11 indicates that the first 1 / 4 mantissa bits starting from the high bit are valid.

[0065] In step 4, the floating point operation processing methods include:

[0066] According to the indication bit, only the high-order bits indicated as valid in the mantissa of the floating-point number are calculated; and / or, the Taylor formula calculation of the specified number of terms is performed according to the indication bit; and / or, different error thresholds are set according to the indication bit to perform Newton iteration method calculation; and / or, the input is divided into one or more intervals according to the indication bit, and the activation function is calculated by piecewise interpolation.

[0067] Step 4: After executing the floating-point operation, determine whether the result is a special type of floating-point number; if so, output the special type of floating-point number; if not, set the number of valid bits in the output floating-point number and indicate it with an indicator bit at the end of the output data.

[0068] Specifically, the present invention provides a floating-point number format with variable precision and a processing method thereof, by adding a processor status bit to indicate whether the floating-point format function of the present invention is turned on. A modified NaN number is inserted before the start of each batch of floating-point data to indicate the number of valid bits in the mantissa of the current batch of floating-point numbers. When processing the current batch of floating-point data, the floating-point operation unit can select a fast processing method according to the number of valid bits indicated in the NaN number, thereby improving the speed of floating-point calculation.

[0069] 1) When the status bit is in the off state, the floating-point number processed is the floating-point number under the existing technology, such as FP64 or FP32 or FP16, which has no impact on the existing format and its processing method, and the existing software and hardware can be reused.

[0070] 2) When the status bit is on, additional format conversion between floating-point numbers is omitted, and different numbers of mantissa valid bits can be set for different steps or different batches of floating-point numbers in the same step according to different scenarios and requirements. The hardware can then selectively process quickly, thereby improving the speed of floating-point calculation.

[0071] For example, when converting from FP32 to FP16, since the number of bits of the mantissa and exponent of FP16 are smaller than that of FP32, and the range of numbers represented by the two is different, a series of operations such as truncation, rounding, and judging whether to overflow to + / -Inf are required according to certain rules. Usually, a dedicated hardware unit is required to perform this conversion step. The "eliminating the additional format conversion between floating-point numbers" mentioned above means that there is no such series of operations, but the calculation unit starts the calculation after taking out the valid bits in the mantissa. Usually, the process of taking out is very simple and does not count as format conversion.

[0072] In a specific implementation process, Figure 2 As shown, the whole process can also be expressed in pseudo code as follows:

[0073] Determine the processor status bit

[0074] First check the processor status bit to determine whether the new floating point format is enabled

[0075] If On:

[0076] Extract modified NaNs before processing each batch of floating point numbers

[0077] Sets the number of significant bits to extract from the mantissa, based on the indicator bit in the NaN.

[0078] Prepare to loop through the floating point numbers in the current batch

[0079] For the input floating point number, determine the type of the number

[0080] If not +inf / -inf / NaN:

[0081] Extract the significant bits from the mantissa

[0082] Else:

[0083] Still treated as +inf / -inf / NaN

[0084] According to the number of valid bits of the mantissa and the characteristics of the current floating-point unit, choose whether to enable the simplified algorithm, and then set the number of valid bits of the mantissa in the output floating-point number.

[0085] Start calculation and get output

[0086] The loop processing ends and a modified NaN is inserted at the end of the output data. The indicator bit is obtained according to the current floating-point operation processing mode.

[0087] In the floating point number processing method of the present invention, specifically,

[0088] The present invention adds a processor status bit, and by querying this status bit, it is determined whether the floating-point format of the present invention is turned on. For example, 1 bit can be used, and a value of 0 indicates off. At this time, the floating-point format is not modified, and subsequent operations can be processed according to the existing floating-point processing method. When the value is 1, it indicates on, that is, the floating-point format proposed by the present invention is used. The interval for querying the status bit is usually greater than the time interval of a batch processing, for example, before the AI ​​model starts running.

[0089] When the processor status bit indicates on, before each batch of floating-point operations, an additional format based on NaN modification will be inserted. In a specific implementation, for example, 2 bits are set in the low order of the floating-point mantissa to indicate the number of valid bits in the mantissa of the current batch of floating-point numbers. For example, a bit value of 00 indicates that all mantissa bits are valid, a bit value of 01 indicates that the first 3 / 4 mantissa bits (counting from the high order) are valid, a bit value of 10 indicates that the first 1 / 2 mantissa bits are valid, and a bit value of 11 indicates that the first 1 / 4 mantissa bits are valid. Since the number of mantissa bits in the existing floating-point format may not be divisible by 4, for each format, the specific relationship between the indicator bit and the number of valid mantissa bits can also be given in a table, as shown in the following table:

[0090] Indicates bit value (binary) Number of significant bits of the mantissa (starting from the high bit) 00 The number of mantissa bits, that is, all are valid 01 ceil(number of mantissa bits / 4*3) 10 ceil(number of mantissa bits / 4*2) 11 ceil(number of mantissa bits / 4)

[0091] Here, ceil represents the rounding up operation, for example, ceil(1.0)=1, ceil(1.1)=2, ceil(1.5)=2, ceil(1.6)=2.

[0092] The above table is for FP32, that is

[0093] Indicates bit value (binary) Number of significant bits of mantissa 00 23 01 18 10 12 11 6

[0094] Since the modified NaN is also a valid NaN for floating-point units not involved in the present invention, if the current step is the first step of the AI ​​model or the previous step may output an unmodified NaN, it is necessary to determine whether to filter the input data to remove the numbers that will interfere with the modified NaN. Since NaN is generally an abnormal data with a lower frequency than +inf and -inf, if it is clear that there is no NaN in the input of the current scenario, there is no need to do this filtering step.

[0095] In other words, when using this new floating-point format, a special NaN value is inserted into the data to indicate the valid number of mantissa bits. However, this modified NaN is still an ordinary NaN for computing units that do not use this new format. Therefore, if unmodified NaNs may be encountered in a certain step of the AI ​​model (especially the first step), it is necessary to determine whether to filter out these NaNs from the input data to prevent them from interfering with the interpretation of the modified NaNs. However, if it can be determined that there is no NaN in the input data in the current scenario, then such filtering is not required. In short, in order to prevent ordinary NaNs from interfering with the interpretation of the new format, these ordinary NaNs may need to be filtered out before processing.

[0096] If the processor status bit is on, it is considered that all NaNs encountered during processing are newly added NaNs. Otherwise, if it is off, it is considered that all NaNs encountered during processing are NaNs in the prior art.

[0097] After obtaining the number of valid bits in the mantissa, the floating-point operation can extract only the valid bits, or extract them according to the mantissa of full precision, and then the floating-point operation unit can choose a simplified processing method for processing.

[0098] The existing floating-point formats also include NaN, inf, and -inf. When encountering these three situations, the hardware does not need to determine the specific bits in the mantissa and directly processes them in a preset specific manner.

[0099] When a floating-point operation is output, the number of valid bits in the output floating-point number must be set according to the current operation type, and then a number modified based on NaN is set as an indicator according to the same rules, and then placed at the end of the output data.

[0100] Figure 3 In the initial batch of floating point numbers: before the floating point batch is processed, a special NaN value is first inserted at the beginning of the floating point batch. This NaN value is modified and its mantissa contains information about the number of significant digits of the mantissa in the next batch of floating point numbers.

[0101] Floating point operation process: The floating point numbers in the batch are operated. The operation unit reads each floating point number, including the starting special NaN value. For floating point numbers that are not special NaN values, the operation unit determines the number of valid bits in the mantissa based on the information encoded in the NaN value and performs the optimized operation process accordingly.

[0102] Optimize operations: By identifying the number of significant digits in the mantissa, the arithmetic unit can use simplified processing to increase the speed of operations. For example, if only part of the mantissa is significant, then when performing operations such as addition or multiplication, the insignificant digits in the mantissa can be ignored, thereby reducing the amount of calculation and improving efficiency.

[0103] Output result: The optimized floating point batches are output, and the special NaN value at the beginning of each batch still exists. In this way, the next step (or the next level of computing unit) can use this information to further process the data.

[0104] Figure 3 The operation unit given in the figure is the most typical one with two inputs and one output. In practice, the input and output of an operation unit may be one or more.

[0105] The dotted box in the operation unit means taking out the valid bits indicated in the mantissa, and then participating in the operation together with the sign bit and the exponent bit.

[0106] The indication of the significant bits of the mantissa can be based on convention. At any position in the mantissa, the figure shows the low-order bits (the low-order bits refer to those counted from right to left).

[0107] The order of input and output data is from left to right.

[0108] The floating-point number processing method of the present invention can calculate only the high-order bits indicated as valid in the mantissa of the floating-point number according to the indication bit; and / or, perform Taylor formula calculation of a specified number of terms according to the indication bit; and / or, set different error thresholds according to the indication bit to perform Newton iteration method calculation; and / or, divide the input into one or more intervals according to the indication bit, and perform activation function calculation by piecewise interpolation.

[0109] Example 1

[0110] In this embodiment, the floating point number processing method is applied to calculate only the high-order bits indicated as valid in the mantissa of the floating point number according to the indication bit position;

[0111] To add floating-point numbers in FP32, first, the mantissa of one of the two floating-point numbers needs to be shifted according to the exponents of the two floating-point numbers so that the exponents of the two numbers remain the same after the shift. Then, the two mantissas are added bit by bit. The existing technology is to add all 23 bits in the mantissa, as shown below: 01000101011110111011011

[0113] +11111001110101011011100

[0114] _______________________ 100111111010100010110111

[0116] If the indicator bit value is 10 (binary), then only 12 bits starting from the high bit need to be calculated, that is, 01000101011110111011011

[0118] +11111001110101011011100

[0119] ____________|___________

[0120] 1001111110100XXXXXXXXXXX

[0121] The lower 11 bits of the mantissa do not need to be calculated and can be filled with an arbitrary value (indicated by X above), which simplifies the calculation.

[0122] Example 2

[0123] In this embodiment, the floating point number processing method is applied to Taylor formula calculation of a specified number of terms according to the indication bit;

[0124] In computers, exponential operations are usually performed using the Taylor formula, i.e., e^x=1+x+x^2 / 2! +x^3 / 3! +...+x^n / n! +...

[0125] The above formula is an infinite accumulation of terms in mathematics, which can get an infinitely accurate result. However, the exponent and mantissa in the floating-point format in the computer are represented by limited bits, which cannot represent infinite precision, so the first few terms of the Taylor formula are usually taken. For the floating-point format involved in the present invention, the number of terms of the Taylor formula when calculating the exponent can be selected according to the number of effective bits of the mantissa, as shown in the following table:

[0126] Indicates bit value (binary) Number of significant bits of mantissa Actual calculation formula for exponential operation 00 23 1+x+x^2 / 2! +x^3 / 3! +x^4 / 4! +x^5 / 5! +x^6 / 6! 01 18 1+x+x^2 / 2! +x^3 / 3! +x^4 / 4! +x^5 / 5! 10 12 1+x+x^2 / 2! +x^3 / 3! +x^4 / 4! 11 6 1+x+x^2 / 2! +x^3 / 3!

[0127] Example 3

[0128] In this embodiment, the floating point number processing method is applied to set different error thresholds according to the indication bit to perform Newton iteration method calculation;

[0129] The square root of a number can be calculated using the Newton-Raphson iteration method. The Newton iteration method usually sets an error threshold or directly sets the number of iterations before calculation, and then calculates multiple times until it is less than the error threshold or reaches the number of iterations. In this way, for the floating-point numbers involved in the present invention, different error thresholds can be set according to different mantissa valid bits, as shown in the following table:

[0130] Indicates bit value (binary) Number of significant bits of mantissa Error threshold 00 23 1e-6 01 18 1e-5 10 12 1e-4 11 6 1e-3

[0131] Example 4

[0132] In this embodiment, the floating point number processing method is applied to divide the input into one or more intervals according to the indication bit, and perform activation function calculation by piecewise interpolation;

[0133] Tanh (hyperbolic tangent), Sigmoid, and Softmax are a type of activation function commonly used in AI. The characteristics of these functions are that the input range is large and the output range is small. In some input intervals (for example, the functions listed above are near zero), the function's rate of change is large. If only the Taylor expansion formula similar to that in Example 2 is used, it is difficult to obtain the required accuracy in the entire input interval. In actual calculations, the piecewise interpolation method is often sampled, that is, the input is divided into multiple intervals, and a set of interpolation methods are used in each interval for calculation. According to the different number of effective bits of the mantissa, it can be divided into different intervals before interpolation. For example, the piecewise interpolation interval of the tanh function is shown in the following table:

[0134]

[0135] It can be known from the description of the above implementation modes that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various implementation modes of the present application or certain parts of the implementation modes.

[0136] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0137] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0138] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.

Claims

1. A floating point number processing method used in large model calculations, characterized in that: The processing method comprises the following steps: Step 1: Add a processor status bit in the processor status register to indicate whether the floating point format function is turned on or off; Step 2: Query the processor status bit before floating point operation; Step 3: When the processor status bit indicates on, insert a modified NaN number as an indicator bit value before the floating-point operation to indicate the number of valid bits in the mantissa of the current batch of floating-point numbers; Step 4: Select a processing method to perform floating-point operations according to the number of valid bits in the mantissa of the floating-point number.

2. The processing method according to claim 1, characterized in that In step 1, the processor status bit provides a global indication of whether the floating-point format function is turned on or off; the processor status bit is set externally and / or by the processor; Use one bit to indicate whether the floating-point format function is on or off. A value of 0 indicates off, and a value of 1 indicates on.

3. The processing method according to claim 1, characterized in that: In step 2, the interval of querying the processor status bit is greater than the time interval of a batch processing, and subsequent floating-point operations are performed according to the query value of the processor status bit.

4. The processing method according to claim 1, characterized in that: Before step 3, there is also a step of checking whether there is a NaN value in the input data; if there is, screening will be performed to exclude the original NaN value in the input data; if not, no screening is required and the next step is directly entered; and / or, Check whether the input data is a special floating point type including +inf, -inf, and NaN; If it is a special floating point number type, it is directly processed in a preset specific way; if not, the significant digits of the mantissa are extracted.

5. The processing method according to claim 1, characterized in that: In step 3, one or more bits are set at any preset agreed position of the mantissa of the floating point number to indicate the number of valid bits in the mantissa of the current batch of floating point numbers; and / or, By setting 2 bits in the low position of the floating-point number mantissa, the number of valid bits in the mantissa of the current batch of floating-point numbers is indicated. The set bit values ​​include 00, 01, 10, and 11. Among them, the bit value 00 indicates that all the mantissa bits are valid, the bit value 01 indicates that the first 3 / 4 of the mantissa bits starting from the high bit are valid, the bit value 10 indicates that the first 1 / 2 of the mantissa bits starting from the high bit are valid, and the bit value 11 indicates that the first 1 / 4 of the mantissa bits starting from the high bit are valid.

6. The processing method according to claim 1, characterized in that: In step 4, the floating point operation processing methods include: According to the indication bit, only the high-order bits indicated as valid in the mantissa of the floating-point number are calculated; and / or, the Taylor formula calculation of the specified number of terms is performed according to the indication bit; and / or, different error thresholds are set according to the indication bit to perform Newton iteration method calculation; and / or, the input is divided into one or more intervals according to the indication bit, and the activation function is calculated by piecewise interpolation; After performing the floating-point operation, determine whether the result is a special type of floating-point number; if so, output the special type of floating-point number; if not, set the number of valid bits in the output floating-point number and indicate it with an indicator bit at the end of the output data.

7. A processing system for implementing the processing method according to any one of claims 1 to 6, characterized in that: The processing system includes: a status bit management module, a special NaN value processing module, a floating point operation module, a special floating point processing module, and an output control module; The status bit management module is responsible for managing the status bit of the processor and determining whether to enable a new floating point number processing method; The special NaN value processing module is used to insert a special NaN value into the mantissa of the floating point number before processing the floating point number; and / or, identify and parse the special NaN value in the data stream and extract the effective mantissa digit information; The floating-point operation module is used to process standard floating-point operations, and / or receive valid mantissa digit information, and perform floating-point operations according to the valid mantissa digit information; The special floating point number processing module is used to identify and process special types of floating point numbers including +inf, -inf, and NaN, and directly output the special types of floating point numbers or perform corresponding processing as needed; The output control module is used to insert a special NaN value at the start position of each batch of simplified output data to carry the information of the number of significant digits in the mantissa.

8. A hardware system for implementing the floating point number processing method according to any one of claims 1 to 6, characterized in that: The hardware system includes: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the floating-point number processing method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the floating-point number processing method according to any one of claims 1 to 6 is implemented.

10. Application of the processing method according to any one of claims 1 to 6, the processing system according to claim 7, the hardware system according to claim 8, or the computer-readable storage medium according to claim 9 in data processing of large model operations and running AI models on the terminal side.