Full FP16 operation softmax activation function optimization method, system and application
By applying the FP16 fast exp algorithm in the large language model to replace the exponential operation in the softmax function, the problem of deploying the softmax activation function on resource-constrained hardware is solved, efficient and low-complexity calculation is achieved, and data overflow is avoided.
Patent Information
- Application Number
- CN202410500052.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-24
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to deploy softmax activation functions in large language models on resource-constrained hardware, especially due to the risk of overflow and high computational complexity caused by exponential operations.
The FP16 fast exp algorithm is used to replace the exponential operation in the softmax function with linear operation, and the calculation is ensured to be carried out within the FP16 accuracy range through vector centralization and truncation processing.
The optimized deployment of softmax functions is implemented, which reduces the complexity and computing overhead of hardware implementation, avoids the risk of data overflow, and saves memory overhead.
Smart Images

Figure BDA0004808394550000031 
Figure BDA0004808394550000071 
Figure BDA0004808394550000072
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology and relates to a full FP16 operation softmax activation function optimization method, system and application. Background Art
[0002] In recent years, fields such as large language models have developed rapidly and achieved excellent results. However, due to the huge number of parameters and computing requirements in large language models, their deployment on resource-constrained edge devices poses huge challenges. Currently, a variety of large language model quantization schemes have been proposed, including INT8, INT4 and other INT type quantization (GPTQ), as well as floating point type quantization such as FP8 (FP8 versus INT8 for efficient deep learning inference published in arXiv preprint in 2023). The softmax activation function is a nonlinear activation function in the transformer structure. The current quantization technology is only applicable to linear operations such as convolution and linear activation functions. However, softmax and LayerNorm nonlinear activation functions are widely used in transformer structures, and the main component of large models is the transformer structure unit. Therefore, in the quantization process of large models, the quantization of nonlinear functions such as softmax has always been a difficult problem. In addition, the softmax function involves exponential operations. Considering that the range of exponential operations can reach infinity and the numerical growth rate is very fast, if it is represented by low-bit data such as FP16, overflow is likely to occur. Therefore, many INT-type quantizations choose to dequantize back to FP32 on the softmax function. Some people have also proposed softmax full quantization inference (I-BERT: Integer-only BERT Quantization published in Proceedings of the 38th International Conference on Machine Learning in 2021). However, in the quantization process of floating-point types such as FP8, due to the existence of exponential operations, the softmax activation function still chooses thirty-two-bit floating-point numbers (FP32) for calculation, just like INT-type quantization. In addition, since the calculation process of the softmax activation function includes exp exponential operations, exponential operations cannot be directly implemented on hardware, so when deploying, approximation methods or table lookup methods are usually used for approximation, but the table lookup method usually requires a large amount of memory to store the entire table, and the polynomial approximation method usually requires a higher order to meet the accuracy. The computational complexity of high-order polynomials is very high, and it will also bring a large computational overhead. In addition, since the softmax function involves exponential summation operations, considering the data range, the softmax function currently uses FP32 for calculations during the inference process, which will also increase memory overhead and reduce inference speed. Summary of the invention
[0003] In order to solve the deficiencies in the prior art, the object of the present invention is to provide a full FP16 operation softmax activation function optimization method, system and application.
[0004] In view of the problem that the softmax nonlinear activation function in large models including Llama is difficult to deploy on hardware, the present invention proposes a full FP16 operation softmax activation function optimization method for the optimized deployment of the softmax activation function. The method of the present invention applies the FP16 fast exp algorithm to the softmax function, so that the optimized softmax function only includes linear operations, and the operation process only needs to use FP16 precision data, making the softmax function a deployment-friendly activation function. Considering the correspondence between the domain and range of the exponential function, the fast exp algorithm is usually only applied to FP32 or Double data types.
[0005] That is, because the exponential function is very sensitive to the input value, a small change in the input can lead to a large change in the output value. Therefore, in order to maintain sufficient calculation accuracy and reduce errors, high-precision data types are usually required to perform these calculations. Due to its lower precision, FP16 is traditionally not suitable for scenarios that require high-precision calculations; therefore, the fast exp algorithm is usually only applied to higher-precision data types, such as FP32 (32-bit single-precision floating point numbers) or Double (double-precision floating point numbers).
[0006] Based on commonly used data sets of large language models such as MMLU, C4, Ptb, etc., tests were conducted on the LLama model and it was found that the accuracy sensitivity of the softmax operator is not high during the reasoning process of the LLama model. Therefore, the present invention applies the FP16 fast exp algorithm to the softmax function to realize the full FP16 softmax function reasoning process.
[0007] Specifically, the present invention provides a method for optimizing a full FP16 operation Softmax activation function, the method comprising the following steps:
[0008] Step 1: Perform a vector centering operation on the input vector and / or matrix to convert all elements in the input vector and / or matrix into non-positive numbers;
[0009] Step 2: Use FP16's exp fast algorithm to convert the exp operation;
[0010] Step 3: truncate the element values in the input vector and / or matrix after the vector is centered according to the data range in the exp fast algorithm;
[0011] Step 4: Replace the exp operation in the softmax calculation formula with the FP16 fast exp calculation converted in step 2 to perform optimized calculation of the softmax activation function.
[0012] In step 1, each element in the input vector and / or matrix is subtracted from the maximum value of the input vector and / or matrix in the input dimension, and the input vector and / or matrix is represented as z=(z1, z2, ..., z k ), the maximum value of the input vector and / or matrix in the input dimension is expressed as M = max(z); the input vector and / or matrix after vector centering can be expressed as z' = (z1-M, z2-M, ..., z k -M), since when the exponent in the exp operation is greater than 0, the operation result will explode with the increase of the exponent, therefore, in the present invention, z is set i -M≤0.
[0013] In step 2, in the FP16 exp fast algorithm, the exponential operation of the natural constant e is converted to the exponential operation of the constant 2. X , that is, e in exp calculation A Convert to That is 2 1.4427*A ; Wherein, X = 1.4427*A;
[0014] The exponential operation with a constant of 2 is transformed into the following steps:
[0015] Step 2.1, first calculate 2 offset (X+bias-C); where offset is the displacement, which is the same as the number of decimal places in the floating-point representation in IEEE754, bias is the offset of the exponent, which is the offset in the floating-point representation in IEEE754, and C is an empirical value used to correct the error in the calculation of the decimal X, which is 0.05798; so for the FP16 floating-point type, offset is 10 and bias is 15 (reference: Kahan W. IEEE standard 754 for binary floating-point arithmetic [J]. Lecture Notes on the Status of IEEE, 1996, 754 (94720-1776): 11., Schraudolph N NA fast, compact approximation of the exponential function [J]. Neural Computation, 1999, 11 (4): 853-862.);
[0016] Step 2.2: Convert the calculation result of step 2.1 into an integer, and adapt the calculation result to the storage format of the integer;
[0017] Step 2.3, read the binary data in the memory of step 2.2 in a floating point data storage mode to obtain the final calculation result;
[0018] In the FP16 exp fast algorithm of step 2, the maximum value that FP16 floating-point data can represent is 65504, and the acceptable range of X satisfies the following relationship:
[0019] X+bias-C≥0
[0020] 2 offset (X+bias-C)≤65504
[0021] Then X∈[-14.92,49.437].
[0022] In step 3 and step 4, according to the value range of X in step 2 and the calculation formula of the softmax activation function, the value of the element in the input vector and / or matrix after vector centering is equivalent to A in step 2, so the following conditions are met:
[0023] -14.92≤1.4427*A=1.4427*(z i -M)≤49.437
[0024] and z in step 1 i -M≤0 can immediately obtain z i -M range is -10.34 ≤ z i -M≤0;z i -All values less than -10.34 in M are replaced by -10.34, completing the truncation of the element values in the vector and / or matrix after vector centering;
[0025] The output of the softmax activation function for the input vector and / or matrix after vector centering is expressed as:
[0026]
[0027] Replace the exp operation in the softmax calculation formula with the FP16 fast exp calculation converted in step 2 to perform optimized calculation of the softmax activation function.
[0028] The present invention also provides an optimization system for implementing the above optimization method, the optimization system comprising: an input preprocessing module, an FP16 fast exp calculation module, a truncation processing module, an optimized softmax calculation module, and a hardware interface module; wherein,
[0029] The input preprocessing module is used to process the input matrix, subtract the input matrix from the maximum value on the input dimension to ensure that all elements in the input matrix become non-positive numbers, prevent data overflow in subsequent exp operations, and optimize the calculation process;
[0030] The FP16 fast exp calculation module is used to implement a fast exp algorithm with FP16 precision, replacing the traditional exponential operation and simplifying the exp calculation process by using linear operations, which can improve the operation speed and reduce the complexity of hardware implementation;
[0031] The truncation processing module is used to further optimize the calculation process and prevent data overflow. The truncation processing module truncates the values in the input matrix according to the input range of the FP16 fast exp algorithm to ensure that all operations are performed within the safe numerical range of FP16;
[0032] The optimized softmax calculation module is responsible for executing the optimized softmax function calculation, replacing the exp operation in the traditional softmax calculation with the operation implemented by the FP16 fast exp calculation module; ensuring that the calculation process of the entire softmax function involves only linear operations, is suitable for FP16 precision data, and is easy to deploy on hardware;
[0033] The hardware interface module is responsible for communicating with the hardware part of the system to ensure that the optimized softmax activation function can be efficiently implemented on various hardware platforms including edge devices; the hardware interface module provides flexibility for the system to adapt to different hardware environments and requirements.
[0034] The present invention also provides the application of the above-mentioned optimization method or optimization system in the optimization deployment of large language models, the optimization deployment of common convolutional network (resenet, mobilenet, etc.) models, and the optimization of other neural network deployments containing softmax activation functions.
[0035] The present invention also provides a hardware system for implementing the above-mentioned optimization method, and the hardware system includes: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned optimization method is implemented.
[0036] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned optimization method is implemented.
[0037] The beneficial effects of the present invention include: the existing softmax function contains exponential operations and exponential sum operations, and the exponential operations cannot be directly deployed on the hardware. In addition, since the exponential operation data grows very quickly, the exponential sum operation is prone to data overflow. The present invention replaces the exponential operation in the softmax operation with a fast exp operation. The process only includes linear operations to facilitate hardware deployment. In view of the risk of data overflow, the present invention makes the input subtract the maximum value of the corresponding dimension according to the dimension parameter before the softmax function performs the exponential operation. At this time, the input is all non-positive numbers, and the value range after the exponential operation is less than or equal to 1, and there will be no data overflow problem. Considering that the data requires more memory and the operation takes a long time during FP32 operation, the present invention truncates the lower bound of the input data so that the data accuracy of the entire softmax operation process is within the FP16 representation range.
[0038] Exponential operation itself is an operation that cannot be implemented by hardware units, so approximate operations must be performed in actual deployment. The commonly used table lookup method requires a large on-chip memory for storage, and a chip with a smaller area may not be able to meet the requirements; polynomial approximation usually requires a higher order to approximate, and the calculation overhead of high-order power operations is very large. The approximation method proposed in the present invention only involves shift operations and linear operations, both of which are extremely easy to implement and run very fast for hardware. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0040] Figure 1 It is the softmax function optimization flow chart of the present invention. DETAILED DESCRIPTION
[0041] The present invention is further described in detail with reference to the following specific examples and drawings. The process, conditions, experimental methods, etc. for implementing the present invention, except for the contents specifically mentioned below, are all common knowledge and common common sense in the art and are not particularly limited by the present invention.
[0042] Some of the nouns involved in the present invention are described as follows:
[0043] FP16, or 16-bit floating point, is a common floating point representation format that uses 16 bits to store values. It consists of 1 sign bit, 5 exponent bits, and 10 mantissa bits. The FP16 format provides a compromise between computational accuracy and storage size, making it very useful in applications that need to reduce model size and computing resource consumption, such as deep learning models on mobile devices.
[0044] FP32, or 32-bit floating point number, is a common floating point representation format that uses 32 bits to store values. It contains 1 sign bit, 8 exponent bits, and 23 mantissa bits, providing higher numerical precision.
[0045] Softmax is an activation function that is often used in the output layer of multi-class classification problems. It can convert a vector into a probability distribution form, where each element of the vector is mapped to a value between 0 and 1, and the sum of all elements is 1. The Softmax function is widely used in deep learning models to predict the probability of each category.
[0046] LayerNorm is a common normalization technique in deep learning that performs normalization on all features of a single sample. LayerNorm is particularly suitable for recurrent neural network (RNN) and Transformer models, helping to speed up the training process and improve the stability of the model.
[0047] LLama refers to a large-scale language model containing billions or even trillions of parameters that can handle complex natural language processing tasks such as text generation, translation, question and answer, etc.
[0048] Transformer is a deep learning architecture with a core mechanism called self-attention, which allows the model to take into account the context of the entire sequence when processing each sequence element. The Transformer architecture has become the basis for many NLP tasks and has also been used in computer vision and other fields.
[0049] The present invention provides a method for optimizing the softmax activation function of full sixteen-bit floating point numbers (FP16), and belongs to the field of artificial intelligence. The optimization method includes: to prevent the overflow of softamx calculation, the input matrix is subtracted from the maximum value on the input dimension, and all elements of the input matrix become non-positive numbers; the exp fast algorithm of FP16 is used to perform exp calculation, and the numerical value of the input range of the FP16 fast algorithm in the input matrix is truncated according to the exp fast algorithm; the exp calculation in the softmax calculation formula is replaced by the FP16 fast exp calculation to obtain the optimized softmax function. The fast exp calculation only involves linear calculation, which solves the problem of the difficulty of deploying nonlinear functions on hardware. This scheme proposes an optimization scheme for the problem of the difficulty of deploying the nonlinear activation function softmax. The optimized softmax function only includes linear calculation, and the data type is FP16, which is a hardware deployment-friendly activation function.
[0050] The present invention provides a method for optimizing a full FP16 operation Softmax activation function, the method comprising the following steps:
[0051] Step 1: Perform a vector centering operation on the input vector and / or matrix to convert all elements in the input vector and / or matrix into non-positive numbers;
[0052] Step 2: Use FP16's exp fast algorithm to convert the exp operation;
[0053] Step 3: truncate the element values in the input vector and / or matrix after the vector is centered according to the data range in the exp fast algorithm;
[0054] Step 4: Replace the exp operation in the softmax calculation formula with the FP16 fast exp calculation converted in step 2 to perform optimized calculation of the softmax activation function.
[0055] In step 1, each element in the input vector and / or matrix is subtracted from the maximum value of the input vector and / or matrix in the input dimension, and the input vector and / or matrix is represented as z=(z1, z2, ..., z k ), the maximum value of the input vector and / or matrix in the input dimension is expressed as M = max(z); the input vector and / or matrix after vector centering can be expressed as z' = (z1-M, z2-M, ..., z k -M), since when the exponent in the exp operation is greater than 0, the operation result will explode with the increase of the exponent, therefore, in the present invention, z is set i -M≤0.
[0056] In step 2, in the FP16 exp fast algorithm, the exponential operation of the natural constant e is converted to the exponential operation of the constant 2. X , that is, e in exp calculation A Convert to That is 2 1.4427*A ; Wherein, X = 1.4427*A;
[0057] The exponential operation with a constant of 2 is transformed into the following steps:
[0058] Step 2.1, first calculate 2 offset (X+bias-C); where offset is the displacement, bias is the exponential offset, and C is an empirical value used to correct the error in the calculation of the decimal X; offset is 10 and bias is 15;
[0059] Step 2.2: Convert the calculation result of step 2.1 into an integer, and adapt the calculation result to the storage format of the integer;
[0060] Step 2.3, read the binary data in the memory of step 2.2 in a floating point data storage mode to obtain the final calculation result;
[0061] In the FP16 exp fast algorithm of step 2, the maximum value that FP16 floating-point data can represent is 65504, and the acceptable range of X satisfies the following relationship:
[0062] X+bias-C≥0
[0063] 2 offset (X+bias-C)≤65504
[0064] Then X∈[-14.92,49.437].
[0065] In step 3 and step 4, according to the value range of X in step 2 and the calculation formula of the softmax activation function, the value of the element in the input vector and / or matrix after vector centering is equivalent to A in step 2, so the following conditions are met:
[0066] -14.92≤1.4427*A=1.4427*(z i -M)≤49.437
[0067] and z in step 1 i -M≤0 can immediately obtain z i -M range is -10.34 ≤ z i -M≤0;z i -All values less than -10.34 in M are replaced by -10.34, completing the truncation of the element values in the vector and / or matrix after vector centering;
[0068] The output of the softmax activation function for the input vector and / or matrix after vector centering is expressed as:
[0069]
[0070] Replace the exp operation in the softmax calculation formula with the FP16 fast exp calculation converted in step 2 to perform optimized calculation of the softmax activation function.
[0071] The optimization method in the present invention is further described as follows:
[0072] The softmax activation function is an activation function that exists in the transformer structure. Given a vector z = (z1, z2, …, z k ), its softmax function expression is:
[0073]
[0074] Considering that when the input of the exponential function is greater than 0, the operation result grows explosively with the increase of the input, the above formula is improved. The vector centering operation is performed on the given vector z. First, the maximum value M = max(z) of the input vector z is obtained, and then the maximum value is subtracted from all elements in the original input vector to obtain the input vector and / or matrix after vector centering, which is expressed as z' = (z1-M, z2-M, ..., z k -M). You can then perform softmax operations:
[0075]
[0076] The fast exp algorithm (A fast, compact approximation of the exponential function published in Neural Computation in 1999) is a fast calculation method based on the IEEE floating point representation method. The floating point number is divided into several parts: a part M for representing the decimal, a part E for representing the exponent, and a sign bit S; the sign bit S determines the positive or negative value, the decimal bit M represents the precision of the value, and the exponent bit E determines the range of the value. The exponent part has a fixed bias. Then any floating point number Y can be represented as:
[0077] Y=(-1) S *(1+M)*2 E-bias
[0078] From the above formula, we can see that the exponential operation of 2 already exists in the representation of floating-point numbers, so if we want to calculate 2 X Just put X+bias in the exponent position, E=X+bias, the exponent position is X+bias-bias=X, and then you can get the corresponding result according to the representation method of the exponent position in floating-point data.
[0079] The above process is completed entirely within the framework of floating-point number representation and operation rules.
[0080] Specifically, an FP32 floating point number occupies 32 bits of memory space in the memory to represent it. Among these 32 bits, 0 to 22 are decimal places represented by M, 23 to 30 are exponent places represented by E, and 31 is the sign place represented by S. For the 32-bit floating point representation, the exponent bit bias is 127; for 32-bit floating point operations, if X is an integer, you only need to shift X+bias left by 23 bits to put it all in the exponent position. However, if X is a decimal, the shift operation will produce an error, which can be corrected by the correction value C. C is an empirical value, which is about 0.05798. At this time, 2 X The exponential operation can be transformed into the following steps:
[0081] Step (1) First calculate 2 offset (X+bias-C); offset is the displacement, bias is the exponent offset (127 in FP32), and C is an empirical value used to correct the error in the calculation of the fraction X;
[0082] Step (2) converts the calculation result of step (1) into an integer and adapts the calculation result to the storage format of the integer;
[0083] Step (3) reads the binary data in the memory of step (2) in a floating point data storage mode to obtain the final calculation result.
[0084] For floating-point data of different precisions, the offset and bias values in step (1) are different. The original paper verifies the FP32 and Double data types. The present invention applies it to FP16 precision floating-point data. In FP16 precision floating-point data, 16 bits of memory space are occupied for representation, 0 to 9 bits are decimal places, represented by M, 10 to 14 bits are exponent bits, represented by E, and 15 bits are sign bits, represented by S; in FP16 data, offset = 10, bias = 15.
[0085] Calculate the acceptable input range of FP16 fast exp calculation. Since negative numbers exist in the form of two's complement in memory, they are no longer applicable to the fast exp algorithm. In addition, the maximum value that FP16 floating-point data can represent is 65504. Therefore, the input needs to meet the following requirements:
[0086] X+bias-C≥0
[0087] 2 offset (X+bias-C)≤65504
[0088] Substituting offset=10,bias=15,C=0.05798,we can obtain X∈[-14.92,49.437].
[0089] For any e A can be converted to:
[0090]
[0091] Log2e is a constant of about 1.4427, so the exponential operation of e in softmax is converted to the exponential operation of 2. Since computers usually have hardware support for exponential functions with a base of 2, the softmax function can be better implemented in computers. At this time, the softmax function is converted to:
[0092]
[0093] Next, the value range of the input data is truncated. The input range X∈[-14.92,49.437] calculated in step 4 is i -M≤0, we can get -10.34≤z i -M≤0, so z i -All values less than -10.34 in M are replaced with -10.34 to make them fit the input range of the FP16 fast exp algorithm.
[0094] The FP16 fast exp algorithm is applied to the exponential operation in the original softmax calculation process to obtain the optimized softmax function. At this time, the softmax only includes linear operations and shift operations, and the entire calculation process only uses FP16 precision data.
[0095] Example
[0096] This embodiment tests the FP16 softmax function on the LLama large language model. The specific experimental method is to replace the softmax function in the transformer module of the LLama large model with the FP16 softmax function obtained by the optimization method of the present invention. If the inference accuracy is set to FP16 in the original model, the input data will be converted to FP32 when performing softmax calculation. The present invention uses the LLama 7B model, and the FP16 weight is tested on the MMLU data set. The test results are as follows:
[0097] LLama FP16 softmax LLama Zero shot 28.5% 28.38% Five shot 35.26% 35.16%
[0098] It can be seen that in the zero shot case, the accuracy drops by 0.12% compared with the original model, and in the five shot case, the accuracy drops by 0.1% compared with the original model.
[0099] First, the precision of the softmax function is reduced from FP32 to FP16, which can save half of the memory. Secondly, the operation at this time does not include exponential operations, and it is a function that can be deployed directly on the hardware.
[0100] Exponential operation itself is an operation that cannot be implemented by hardware units. The original softmax function cannot be directly deployed on hardware, so approximate operations must be performed when it is actually deployed on hardware. The commonly used table lookup method requires a large on-chip memory for storage, and a chip with a smaller area may not be able to meet the requirements; polynomial approximation usually requires a higher order to approximate, and the calculation overhead of high-order power operations is very large. The approximation method proposed in the present invention only involves shift operations and linear operations, which are operations that are extremely easy to implement and run very fast for hardware.
[0101] The present invention verifies the optimized FP16 softmax function in a common CNN classification network. The specific experimental process is to replace the last softmax of the four networks in the following table with FP16 softmax, and convert the input into FP16 data. The test is performed on the Imagenet dataset. The specific experimental results are as follows:
[0102]
[0103] From the results in the table above, we can see that compared with the original network, the maximum classification accuracy loss of the top1 and top5 networks after softmax replacement is only 0.022%. This shows that FP16 softmax can achieve optimization that is easy to deploy, saves memory, and speeds up inference with minimal sacrifice of accuracy.
[0104] The present invention provides an optimization method for a deployment-friendly softmax nonlinear activation function. The optimized softmax function only contains linear operations and the calculation process only requires data with FP16 precision, which is convenient for deployment and can save memory usage and speed up reasoning. It is tested on the LLama large language model and the common CNN network with only a slight loss of precision.
[0105] In addition to implementing the client and server in a purely computer-readable program code, the client and server can also implement the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a client and server can be considered as a hardware component, and the means for implementing various functions included therein can also be considered as a structure within the hardware component. Or even, the means for implementing various functions can be considered as both a software module for implementing the method and a structure within the hardware component.
[0106] It can be seen from the above description of the implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a storage medium, such as ROM / RAM, a disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each implementation method of the present application or some parts of the implementation method.
[0107] Each implementation in this specification is described in a progressive manner, and the same or similar parts between the various implementations can be referred to each other, and each implementation focuses on the differences from other implementations. In particular, for the implementation of the client and the server, both can refer to the introduction of the implementation of the aforementioned method for comparative explanation.
[0108] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0109] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.
Claims
1. A method for optimizing the softmax activation function for full FP16 operation, characterized in that: The optimization method comprises the following steps: Step 1: Perform a vector centering operation on the input vector and / or matrix to convert all elements in the input vector and / or matrix into non-positive numbers; Step 2: Use FP16's exp fast algorithm to convert the exp operation; Step 3: truncate the element values in the input vector and / or matrix after the vector is centered according to the data range in the exp fast algorithm; Step 4: Replace the exp operation in the softmax calculation formula with FP16 fast exp calculation to obtain the optimized softmax function.
2. The optimization method according to claim 1, characterized in that: In step 1, each element in the input vector and / or matrix is subtracted from the maximum value of the input vector and / or matrix in the input dimension, and the input vector and / or matrix is represented as z=(z1, z2, ..., z k ), the maximum value of the input vector and / or matrix in the input dimension is expressed as M = max(z); the input vector and / or matrix after vector centering can be expressed as z' = (z1-M, z2-M, ..., z k -M), z i -M≤0.
3. The optimization method according to claim 1, characterized in that: In step 2, the exponential operation of the natural constant e is converted to the exponential operation of the constant 2 X , that is, e in exp calculation A Convert to That is 2 1.4427*A ; Wherein, X = 1.4427*A; The exponential operation with a constant of 2 is transformed into the following steps: Step 2.1, Calculation 2 offset (X+bias-C); offset is the displacement, bias is the exponential offset, and C is an empirical value used to correct the error in the calculation of the decimal X; Step 2.2: Convert the calculation result of step 2.1 into an integer, and adapt the calculation result to the storage format of the integer; Step 2.3, read the binary data in the memory of step 2.2 in floating point data storage mode to obtain the final calculation result.
4. The optimization method according to claim 3, characterized in that: The maximum representable value of FP16 floating-point data is 65504, and the range of X satisfies the following formula: X+bias-C≥0, 2 offset (X+bias-C)≤65504; Then X∈[-14.92,49.437].
5. The optimization method according to claim 4, characterized in that: In step 3, according to the value range of X and the calculation formula of the softmax activation function, the value of the element in the input vector and / or matrix after vector centering is equivalent to A in step 2, so the following conditions are met: -14.92≤1.4427*A=1.4427*(from i -M)≤49.437 With z i -M≤0 can immediately obtain z i -M range is -10.34 ≤ z i -M≤0;z i -All values less than -10.34 in M are replaced by -10.34, completing the truncation of the element values in the vector and / or matrix after vector centering.
6. The optimization method according to claim 5, characterized in that: The output of the softmax activation function for the input vector and / or matrix after vector centering is expressed as: Replace the exp operation in the softmax calculation formula with the FP16 fast exp calculation converted in step 2, and use the truncated element values in step 3 to optimize the calculation of the softmax activation function.
7. An optimization system for implementing the optimization method according to any one of claims 1 to 6, characterized in that: The optimization system includes: an input preprocessing module, an FP16 fast exp calculation module, a truncation processing module, an optimized softmax calculation module, and a hardware interface module; wherein, The input preprocessing module is used to process the input matrix, and to make a difference between the input matrix and the maximum value on the input dimension to ensure that all elements in the input matrix become non-positive numbers; The FP16 fast exp calculation module is used to implement the fast exp algorithm with FP16 precision; The truncation processing module truncates the values in the input matrix according to the input range of the FP16 fast exp algorithm to ensure that all operations are performed within the safe numerical range of FP16, optimize the calculation process and prevent data overflow; The optimized softmax calculation module is used to perform the optimized softmax function calculation, ensuring that the calculation process of the entire softmax function only involves linear operations, is suitable for FP16 precision data, and is easy to deploy on hardware; The hardware interface module is used to communicate with the hardware part of the system to ensure that the optimized softmax activation function can be implemented on various hardware platforms.
8. Application of the optimization method according to any one of claims 1 to 6, or the optimization system according to claim 7 in the optimization deployment of large language models, the optimization deployment of convolutional network models, and the optimization deployment of neural networks including softmax activation functions.
9. A hardware system for implementing the optimization method according to any one of claims 1 to 6, characterized in that: The hardware system includes: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the optimization method according to any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the optimization method according to any one of claims 1 to 6 is implemented.