Method for increasing chip operation speed based on SIMD instruction

The calculation process of the ELU activation function optimization through the SIMD instruction set solves the problem of slow computing speed of the chip when the cache resources are limited, and efficient computing on the Beijing Junzheng t40 chip is realized, which improves the computing speed.

CN120371398APending Publication Date: 2025-07-25INGENIC SEMICON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311819659.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, chips using ELU activation functions have slow computing speed on platforms with limited cache resources, and cannot effectively use SIMD instructions for high-speed calculations, resulting in low computing efficiency.

Method used

Using the SIMD instruction set, a series of instruction operations are used to realize the rapid solution of ELU activation functions, including data loading, comparison, logical operations and exponential function calculation, which optimizes cache utilization and improves the computing speed.

Benefits of technology

On a platform with limited hardware resources, the computing speed of the ELU activation function has been significantly improved and the overall computing power of the chip has been improved, especially on the Beijing Junzheng t40 chip, which has achieved an 11-fold increase in computing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371398A_ABST
    Figure CN120371398A_ABST
Patent Text Reader

Abstract

The invention provides a method for improving the operation speed of a chip based on an SIMD instruction, and the method comprises the steps: rapidly solving an ELU activation function, and the ELU function formula: # imgabs0 # comprises the steps: S1, loading data; s2, the 16 (flow) data in the Register 1 are compared with 0, and the 16 (flow) data in the Register 1 are compared with 0; s3, carrying out AND on the Register 3 and the Register 1, and carrying out AND on the Register 3 and the Register 1; s4, negating each bit of the number in the Register 3, changing 0 into 1, changing 1 into 0, and performing AND with the number in the Register 1; s5, solving an exp value of the data in the Register 5 by using an implementation method of an exp exponential function based on an SIMD (Single Instruction Multiple Data) instruction, and storing a result in a register Register 6; s6, the data in the Register 6 and the floating point number 1.0 f are subtracted; s7, multiplying the data in the Register 7 by a floating point number alpha; s8, the number larger than 0 in the Register 4 is original data, the number smaller than or equal to 0 is 0, the number larger than 0 in the register Register 8 is 0, the number smaller than or equal to 0 is calculated, and the Register 4 and the Register 8 are subjected to phase OR; and S9, storing a calculation result in a memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural networks, and particularly relates to a method for improving the operation speed of a chip based on SIMD instructions. Background Art

[0002] With the development of technology, especially the advent of the era of artificial intelligence, neural networks are widely used in the field of artificial intelligence. Among them, the operation speed of the chip is also required to be faster and faster. In a neural network, the input of each node is a linear combination of the output values of the nodes in the previous layer, and the output value of this node is the value after the non-linear transformation of this linear combination. The function for performing the non-linear transformation on the linear combination is the activation function. A neuron node receives a value of the linear combination, and then outputs the activated value through the activation function f. The activation function plays a role of non-linear mapping in the neural network, enabling the neural network to extract sufficient features. Without using the activation function, the neural network cannot solve non-linear problems. The activation function must be a non-linear function. The neural network model is a hierarchical model. If the activation function is a linear function, then the multi-layer neural network and the single-layer neural network are equivalent and cannot exert the advantage of the neural network extracting features layer by layer.

[0003] The sigmoid activation function is one of the most common activation functions. This activation function can map any real number to the interval of 0-1. However, the biggest disadvantage of sigmoid is that this activation function will cause the problem of gradient disappearance, making the training of the neural network stop prematurely, so that the characteristics that the deeper the neural network, the more effective the feature extraction cannot be exerted. To solve the problem of gradient disappearance of the sigmoid function, the relu function is introduced. The advantage of this activation function is that it avoids the problem of gradient disappearance, but it will directly eliminate the features less than 0. This will cause the phenomenon of neuron inactivation, that is, the neurons with features less than 0 will not be updated. To solve the phenomenon of neuron inactivation of the relu activation function, a new activation function leakyrelu is generated, which avoids gradient disappearance and the features less than 0 will not be inactivated, but leakyrelu is not saturated on both sides and cannot converge better. The ELU activation function has all the advantages of the leakyrelu activation function and is unilaterally saturated, and can converge better. Among them, the prototype formula of the ELU function is shown in formula (1):

[0004]

[0005] In the prior art, when using the ELU function in a neural network, it is usually through calling the function interface in the deep learning framework, such as the tf.nn.elu function in TensorFlow, or directly using a function defined according to formula (1) in python or C language.

[0006] However, the defects in the prior art are as follows:

[0007] When using the ELU function, it is usually to call the encapsulated function interface or write the ELU function in Python or C language. However, it cannot run on some platforms or chips that do not support the above programming languages. For example, the chips of Beijing Junzheng Integrated Circuit Co., Ltd. (hereinafter referred to as: Beijing Junzheng) only support C language and assembly language. However, since C language can only process one data at a time, when running C language on a platform or chip with limited cache resources (for example, using the t40 chip of Beijing Junzheng) to process a large amount of continuous data, it may cause cache waste and thus lead to slow operation speed. Therefore, for speed considerations, chips such as those of Beijing Junzheng will preferentially use assembly language for calculation, that is, use SIMD instructions.

[0008] In addition, the commonly used technical terms include:

[0009] 1. SIMD, the full name is Single Instruction Multiple Data, single instruction multiple data stream, which can process multiple operands at a time and pack them into a set of instruction sets in a large register.

[0010] 2. The ELU (Exponential Linear Unit) activation function is an activation function in deep learning, and its appearance helps to solve the problem of gradient disappearance. Summary of the Invention

[0011] In order to solve the above problems, the purpose of this application is as follows: The ELU function is widely used in various tasks of deep learning and neural networks, such as image processing, natural language processing, reinforcement learning, audio processing, etc. The function is a non-linear fitting, which can improve the accuracy of the neural network from different directions, and can both improve the accuracy and reduce the value of the function loss. Secondly, ELU has a faster convergence rate and can accelerate the training speed of the model.

[0012] The method of using SIMD instructions and solving the ELU function can solve the problem of slow operation speed of the ELU function on the chip. SIMD instructions can improve cache utilization, load multiple data at one time, make full use of the cache space as much as possible, and reduce the number of data interactions between the hard disk or memory and the cache, thus greatly improving the operation speed, occupying a small amount of resources, and can also operate efficiently under the condition of limited hardware resources. The method of solving the ELU function by this method can implement the calculation process of formula (1) using SIMD, so that it can run on the Beijing Junzheng t40 chip and improve the operation speed of the chip.

[0013] Specifically, the present invention provides a method for improving the operation speed of a chip based on SIMD instructions. The method includes a process for quickly solving the ELU activation function. The prototype formula of the ELU function is shown in formula (1):

[0014]

[0015] The method further includes the following steps:

[0016] S1. Loading data:

[0017] Load in multiples of 32 bits. One register can load up to 512 bits of data. A single-precision floating-point number is 32 bits, and one register can load 16 floating-point numbers. Therefore, one SIMD instruction can calculate 16 floating-point numbers simultaneously;

[0018] Use the SIMD instruction Ingenic_simd512_load for loading data. (float)data are the 16 input 32-bit single-precision floating-point numbers, and Register1 is register 1. Save the input data in Register1, expressed as

[0019] Register1 = Ingenic_simd512_load((float)data);

[0020] S2. Compare the 16 (float)data in Register1 with 0:

[0021] Use the comparison instruction Ingenic_simd512_fcltzw with 0. This instruction means that when the data at the corresponding position in Register1 is greater than 0, the 32 bits at the corresponding position in register Register3 are all set to 1, and when the data at the corresponding position in Register1 is less than or equal to 0, the 32 bits at the corresponding position in register Register3 are all set to 0, expressed as

[0022] Registe3 = Ingenic_simd512_fcltzw(0, Register1);

[0023] S3. AND Register3 and Register1:

[0024] Use the AND instruction Ingenic_simd512_and. The numbers greater than 0 in Register1 AND with 1 remain the original numbers, and the numbers less than or equal to 0 AND with 0 become 0. The result is saved in register Register4, expressed as

[0025] Registe4 = Ingenic_simd512_and(Register3, Register1);

[0026] S4. Take the bitwise complement of each bit in the number in Register3, changing 0 to 1 and 1 to 0, and then perform a bitwise AND with the number in Register1. The Ingenic_simd512_andn instruction first takes the complement of Register3 and then performs a bitwise AND with Register1. For numbers in Register1 that are less than or equal to 0, the bitwise AND with 1 keeps the original number, and for numbers greater than 0, the bitwise AND with 0 changes the number to 0. The result is stored in register Register5, expressed as

[0027] Registe5 = Ingenic_simd512_andn(Register3, Register1);

[0028] S5. Use the implementation method of the exp exponential function based on SIMD instructions for the data in Register5 to find the value of exp. The result is stored in register Register6, expressed as

[0029] Register6 = Ingenic_simd512_exp(Register5);

[0030] Among them, Ingenic_simd512_exp is the method name of the exp exponential function based on SIMD instructions, which means it can cover all steps of solving exp, and these steps are combined and named Ingenic_simd512_exp;

[0031] S6. Subtract the data in Register6 from the floating-point number 1.0f. The Ingenic_simd512_float_sub instruction is the SIMD instruction for floating-point subtraction. Subtract 1 from the data in register Register6, and the result after subtraction is stored in register Register7, expressed as

[0032] Register7 = Ingenic_simd512_float_sub(Register6, 1.0f);

[0033] S7. Multiply the data in Register7 by the floating-point number α. The Ingenic_simd512_float_mul instruction is the SIMD instruction for floating-point multiplication. Multiply the data in register Register7 by α, and the result after multiplication is stored in register Register8, expressed as

[0034] Register8 = Ingenic_simd512_float_mul(Register7, α);

[0035] The numbers greater than 0 in S8.Register4 are the original data, and the numbers less than or equal to 0 are 0. The numbers greater than 0 in register Register8 are 0, and the numbers less than or equal to 0 are the calculated numbers. Performing an OR operation on Register4 and Register8 implements formula (1). The Ingenic_simd512_or instruction is an OR instruction, and the result is stored in register Register9, expressed as

[0036] Registe9 = Ingenic_simd512_or(Register4, Register8);

[0037] S9. Save the calculation result from register Register9 to memory, expressed as

[0038] Result = Ingenic_simd512_store(Register9);

[0039] The Ingenic_simd512_store instruction is a data save instruction, which saves 16 data in register Register9 to the Result memory that can store 16 data.

[0040] The above step S3 implements the following in formula (1):

[0041] When x > 0, ELU(x) = x.

[0042] The above steps S4 to S7 implement the following in formula (1):

[0043] When x ≤ 0, ELU(x) = α(e x -1).

[0044] The Ingenic_simd512_exp in the above step S5 further includes:

[0045] S5.1, Load data, using the SIMD instruction as follows:

[0046] Register11 = Ingenic_simd512_load((float)data),

[0047] where, Ingenic_simd512_load is a data load instruction; (float)data is 16 32-bit single-precision floating-point numbers input; Register11 is register 11, which saves the input data;

[0048] S5.2, limit the range of the input data between -88.7 and 88.7; use the SIMD instructions as follows:

[0049] Register12 = Ingenic_simd512_load_float(-88.7),

[0050] Register13 = Ingenic_simd512_load_float(88.7),

[0051] Register11 = Ingenic_simd512_max(Register11, Register12),

[0052] Register11 = Ingenic_simd512_min(Register11, Register13),

[0053] Ingenic_simd512_load_float is an instruction to load and copy floating-point numbers. After using this instruction,

[0054] there are 16 -88.7s in Register12 and 16 88.7s in register Register13;

[0055] Ingenic_simd512_max is a maximum value fetching instruction. This instruction compares the 16 data in Register11 with the 16 -88.7s in register Register12 respectively to fetch the maximum value, that is, the data greater than -88.7 remains unchanged, and the data less than -88.7 takes -88.7, and the result is saved in Register11;

[0056] Ingenic_simd512_min is a minimum value fetching instruction. This instruction compares the 16 data in the result Register11 from the previous step with the 16 88.7s in Register13 respectively to fetch the minimum value, that is, the data less than 88.7 remains unchanged, and the data greater than 88.7 takes 88.7, and the result is still saved in Register11;

[0057] S5.3, split e x into two parts for multiplication. Where x is the input data after range limitation in step S5.2. Split x into two parts for addition, and split e x into two parts for multiplication. The formula is as follows:

[0058] x_int = int(x / ln2) Formula (5-1)

[0059] x_bas = x_int * ln2 Equation (5-2)

[0060] x_rem = x - x_bas Equation (5-3)

[0061] x = x_bas + x_rem Equation (5-4)

[0062] e x = 2 x_int * e x_rem Equation (5-5)

[0063] Let exp_bas = 2 x_int Equation (5-6)

[0064] Let exp_rem = e x_rem Equation (5-7)

[0065] e x = exp_bas * exp_rem Equation (5-8)

[0066] (1), Use SIMD instructions to calculate Equation (5-1) x_int = int(x / ln2), where x / ln2 is equivalent to x * (1 / ln2), 1 / ln2 = 1.442695; The instructions are as follows:

[0067] Register12 = Ingenic_simd512_load_float(1.442695),

[0068] Register13 = Ingenic_simd512_float_mul(Register11, Register12),

[0069] Register13 = Ingenic_simd512_float_to_int_truncate(Register13),

[0070] First, load the floating-point number 1.442695 into Register12;

[0071] The Ingenic_simd512_float_mul instruction is a floating-point multiplication instruction. Multiply the 16 input data in Register11 and the 16 1.442695 in Register12, and place the results in Register13;

[0072] The Ingenic_simd512_float_to_int_truncate instruction is a floating-point to integer conversion instruction. This instruction converts the floating-point number in Register13 into an integer, directly discarding the data after the decimal point and retaining the integer part to obtain x_int in the formula x_int = int(x / ln2), and the result is stored in Register13;

[0073] (2), Use the SIMD instruction to calculate exp_bas in the formula (5-6) exp_bas = 2 x_int The SIMD instruction is set as:

[0074] Register14 = Ingenic_simd512_load_immediate(127),

[0075] Register14 = Ingenic_simd512_int_add(Register13, Register14),

[0076] Register14 = Ingenic_simd512_shift_letf(Register14, 23),

[0077] The Ingenic_simd512_load_immediate is a load immediate instruction that loads the integer 127 and copies it 16 times and stores them in Register14;

[0078] The Ingenic_simd512_int_add is an integer addition instruction that adds the data in Register13 and Register14, and the result is placed in Register14;

[0079] The Ingenic_simd512_shift_letf is a left shift instruction that shifts the 16 integers in Register14 to the left by 23 bits;

[0080] S5.4, Use the SIMD instruction to calculate the second part exp_rem of e x

[0081] (1), Use the SIMD instruction to calculate the formula (5-3), where ln2 is 6.931472e-1, and the instruction is as follows:

[0082] Register12 = Ingenic_simd512_load_float(6.931472e-1)

[0083] Register13 = Ingenic_simd512_int_to_float(Register13)

[0084] Register12 = Ingenic_simd512_float_mul(Register12, Register13)

[0085] Register11 = Ingenic_simd512_float_sub(Register11, Register12)

[0086] Ingenic_simd512_load_float loads and copies 6.931472e-1 into Register12; Ingenic_simd512_int_to_float is an integer-to-float conversion instruction that converts 16 integers x_int in Register13 into floating-point numbers. Because in the Ingenic_simd512 instruction set, integers cannot be directly operated on with floating-point numbers. The integers need to be converted into floating-point numbers first, and then floating-point instructions are used for operations;

[0087] Ingenic_simd512_float_mul multiplies 16 ln2s in Register12 by 16 data in Register13 to obtain x_bas, and the result is stored in Register12;

[0088] Ingenic_simd512_float_sub is a floating-point subtraction instruction that subtracts the data in Register11 from the data in Register12 to obtain x_rem, and the result is stored in Register11;

[0089] (2) Use the Taylor expansion to find formula (5-7). The Taylor formula of the exponential function with base e is:

[0090]

[0091] When n is 6, it meets the accuracy requirements of this method for the exp function. This method does not calculate the case where n exceeds 6, otherwise it will cause a large amount of calculation and slow down the operation speed; so e x_rem is expanded into formula (5-10):

[0092]

[0093] Formula (5-10);

[0094] In formula (5-10), 1 / 6! is 1.388889e-3, 1 / 5! is 8.333333e-3, 1 / 4! is 4.166667e-2, 1 / 3! is 1.666667e-1, and 1 / 2! is 0.5;

[0095] The steps to implement formula (5-10) using SIMD are as follows. The SIMD instructions are set as:

[0096] Register12 = Ingenic_simd512_load_float(1.388889e-3)

[0097] Register13 = Ingenic_simd512_float_mul(Register12, Register11)

[0098] Register12 = Ingenic_simd512_load_float(8.333333e-3)

[0099] Register13 = Ingenic_simd512_float_add(Register12, Register13)

[0100] Register13 = Ingenic_simd512_float_mul(Register13, Register11)

[0101] Register12 = Ingenic_simd512_load_float(4.166667e-2)

[0102] Register13 = Ingenic_simd512_float_add(Register12, Register13)

[0103] Register13 = Ingenic_simd512_float_mul(Register13, Register11)

[0104] Register12 = Ingenic_simd512_load_float(1.666667e-1)

[0105] Register13 = Ingenic_simd512_float_add(Register12, Register13)

[0106] Register13 = Ingenic_simd512_float_mul(Register13, Register11)

[0107] Register12 = Ingenic_simd512_load_float(0.5)

[0108] Register13 = Ingenic_simd512_float_add(Register12, Register13)

[0109] Register13 = Ingenic_simd512_float_mul(Register13, Register11)

[0110] Register12 = Ingenic_simd512_load_float(1.0)

[0111] Register13 = Ingenic_simd512_float_add(Register12, Register13)

[0112] Register13 = Ingenic_simd512_float_mul(Register13, Register11)

[0113] Register13 = Ingenic_simd512_float_add(Register12, Register13). Ingenic_simd512_load_float loads and copies 1.388889e-3 into Register12, multiplies the 16 x_rem data in Register11 by the 16 1.388889e-3 in Register12 using Ingenic_simd512_float_mul, and saves the result in Register13;

[0114] Load 8.333333e-3 into Register12,

[0115] Ingenic_simd512_float_add is a floating-point addition instruction that adds the 16 data in Register13 to the 16 data in Register12 and saves the result in Register13; then multiplies Register13 by the 16 x_rem in Register11, and the result is still saved in Register13;

[0116] And so on until all the calculations in formula (5-10) are completed. The result of exp_rem calculated through SIMD instructions is finally stored in Register13;

[0117] S5.5, use SIMD instructions to multiply the exp_bas and exp_rem obtained in the above steps S5.3 and S5.4 to obtain the result of formula (5-8), expressed as

[0118] Register13 = Ingenic_simd512_float_mul(Register13, Register14)

[0119] Ingenic_simd512_float_mul is a floating-point multiplication instruction that multiplies the results of Register13 and Register14;

[0120] S5.6: Save the calculation result from the register to the memory, expressed as

[0121] Result = Ingenic_simd512_store(Register13)

[0122] Ingenic_simd512_store is a data storage instruction that stores 16 data in Register13 into the Result memory that can store 16 data.

[0123] In the method, the instruction set based on SIMD instructions can load 512-bit data into a 512-bit register at one time, and can load 1024-bit cache data with two registers at one time; there are a total of 32 registers available for repeated use; during calculation, the same operation can be performed on 512-bit data at one time. The instruction set includes arithmetic operations such as addition, subtraction, multiplication, absolute value, maximum value, and minimum value, logical operations such as AND, OR, NOT, XOR, and shift, and comparison operations such as less than, equal to, and less than or equal to, meeting general operation requirements. The exp exponential function requires the combined use of multiple instructions.

[0124] Therefore, the advantages of this application are as follows:

[0125] 1. The ELU activation function calculated using the SIMD instruction set of Ingenic_simd512 supported by the Beijing Junzheng t40 chip improves cache utilization, consumes less resources, and the calculation speed is 11 times that of the C language method for calculating the ELU activation function, and can run efficiently on platforms with limited hardware resources.

[0126] 2. Since the SIMD instruction set of Ingenic_simd512 supported by the Beijing Junzheng T40 chip does not have instructions for exp and solving the reciprocal of a positive number, the formula (1) cannot be directly used to solve the ELU function. This method uses the exp function method in the implementation method of the exp exponential function based on SIMD instructions to first find the exp value of the input data less than or equal to 0, and then uses subtraction and multiplication to find the function value, obtaining the result of the ELU function.

[0127] 3. At the same time, by using the comparison with 0 in step S2, the AND operation in step S3, the NOT-AND operation in step S4, and the OR operation in step S8, the operations in C language that could not be originally implemented on the relevant chip are realized.

[0128] Finally, the purpose of accelerating the calculation speed and improving the computing power of the chip, especially the calculation speed when encountering the ELU activation function, is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0129] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.

[0130] Figure 1 It is a flowchart of this method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0131] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings.

[0132] As Figure 1 shown, this method proposes a method for improving the computing speed of a chip based on SIMD instructions, especially a method for improving the computing speed of the ELU function of the chip. The method includes a method for using SIMD instructions to implement the solution of the ELU function, and further includes the following steps:

[0133] S1. Loading data is loaded in integer multiples of 32 bits. A register can load up to 512 bits of data at most. A single-precision floating-point number is 32 bits, and a register can load 16 floating-point numbers. Therefore, one SIMD instruction can calculate 16 floating-point numbers simultaneously. Ingenic_simd512_load is the SIMD instruction for loading data; (float)data is the input 16 32-bit single-precision floating-point numbers, and Register1 is register 1. The input data is saved in Register1, which is expressed as

[0134] Register1 = Ingenic_simd512_load((float)data);

[0135] S2. Compare the 16 (float) data in Register1 with 0. The Ingenic_simd512_fcltzw is the instruction for comparing with 0. This instruction means that when the data at the corresponding position in Register1 is greater than 0, all 32 bits at the corresponding position in Register3 are set to 1, and when the data at the corresponding position in Register1 is less than or equal to 0, all 32 bits at the corresponding position in Register3 are set to 0, which is expressed as

[0136] Registe3 = Ingenic_simd512_fcltzw(0, Register1);

[0137] S3. AND Register3 and Register1. The Ingenic_simd512_and is the AND instruction. The numbers greater than 0 in Register1 AND with 1 remain the original numbers, and the numbers less than or equal to 0 AND with 0 become 0. The result is stored in Register4, which is expressed as

[0138] Registe4 = Ingenic_simd512_and(Register3, Register1);

[0139] This step S3 implements that when x > 0 in formula (1), ELU(x) = x;

[0140] S4. Invert each bit of the numbers in Register3, changing 0 to 1 and 1 to 0, and then AND with the numbers in Register1. The Ingenic_simd512_andn is the instruction that first inverts Register3 and then ANDs with Register1. The numbers less than or equal to 0 in Register1 AND with 1 remain the original numbers, and the numbers greater than 0 AND with 0 become 0. The result is stored in Register5, which is expressed as

[0141] Registe5 = Ingenic_simd512_andn(Register3, Register1);

[0142] S5. Use the exp function in the implementation method of the exp exponential function based on SIMD instructions to find the exp value for the data in Register5, and the result is stored in Register6, which is expressed as

[0143] Register6 = Ingenic_simd512_exp(Register5);

[0144] Ingenic_simd512_exp is the method name of the exp exponential function based on SIMD instructions, which means it can cover all steps of solving exp, and these steps are combined and named Ingenic_simd512_exp;

[0145] S6. Subtract the data in Register6 from the floating-point number 1.0f. Ingenic_simd512_float_sub is the SIMD instruction for floating-point subtraction. Subtract 1 from the data in register Register6, and the result after subtraction is stored in register Register7, expressed as

[0146] Register7 = Ingenic_simd512_float_sub(Register6, 1.0f);

[0147] S7. Multiply the data in Register7 by the floating-point number α. Ingenic_simd512_float_mul is the SIMD instruction for floating-point multiplication. Multiply the data in register Register7 by α, and the result after subtraction is stored in register Register8, expressed as

[0148] Register8 = Ingenic_simd512_float_mul(Register7, α);

[0149] Steps S4 to S7 implement the case when x ≤ 0 in formula (1), ELU(x) = α(e x - 1).

[0150] S8. The numbers greater than 0 in Register4 are the original data, and the numbers less than or equal to 0 are 0. The numbers greater than 0 in Register8 are 0, and the numbers less than or equal to 0 are the calculated numbers. Perform an OR operation on Register4 and Register8, then formula (1) is implemented. Ingenic_simd512_or is the OR instruction, and the result is stored in register Register9, expressed as

[0151] Registe9 = Ingenic_simd512_or(Register4, Register8);

[0152] S9. Save the calculation result from register Register9 to memory, expressed as

[0153] Result = Ingenic_simd512_store(Register9);

[0154] Ingenic_simd512_store is a data storage instruction that stores 16 data in register Register9 into Result memory which can store 16 data.

[0155] In the method, the instruction set based on SIMD instructions can load 512-bit data into a 512-bit register at one time, and for 1024-bit cache data, it can be loaded with two registers at one time; there are a total of 32 registers available for repeated use; during calculation, the same operation can be performed on 512-bit data at one time. The instruction set includes arithmetic operations such as addition, subtraction, multiplication, absolute value, maximum value, and minimum value, logical operations such as AND, OR, NOT, XOR, and shift, and comparison operations such as less than, equal to, and less than or equal to, meeting general operation requirements. The exp exponential function requires the combined use of multiple instructions.

[0156] Among them, Ingenic_simd512_exp in step S5 further includes:

[0157] S5.1, load data, using the SIMD instruction as follows:

[0158] Register11 = Ingenic_simd512_load((float)data),

[0159] where Ingenic_simd512_load is a data loading instruction; (float)data is 16 32-bit single-precision floating-point numbers input; Register11 is register 11, which stores the input data;

[0160] S5.2, limit the range of the input data between -88.7 and 88.7; use the SIMD instruction as follows:

[0161] Register12 = Ingenic_simd512_load_float(-88.7),

[0162] Register13 = Ingenic_simd512_load_float(88.7),

[0163] Register11 = Ingenic_simd512_max(Register11, Register12),

[0164] Register11 = Ingenic_simd512_min(Register11, Register13),

[0165] Ingenic_simd512_load_float is an instruction for loading and copying floating-point numbers. After using this instruction,

[0166] There are 16 -88.7s in Register12 and 16 88.7s in Register13;

[0167] Ingenic_simd512_max is an instruction for taking the maximum value. This instruction compares the 16 data in Register11 with the 16 -88.7s in Register12 respectively and takes the maximum value, that is, the data greater than -88.7 remains unchanged, and the data less than -88.7 takes -88.7. The result is saved in Register11;

[0168] Ingenic_simd512_min is an instruction for taking the minimum value. This instruction compares the 16 data in the result Register11 from the previous step with the 16 88.7s in Register13 respectively and takes the minimum value, that is, the data less than 88.7 remains unchanged, and the data greater than 88.7 takes 88.7. The result is still saved in Register11;

[0169] S5.3, split e x into two parts and multiply them. Here, x is the input data after range limitation in step S5.2. Split x into two parts and add them. e x is split into two parts and multiplied. The formula is as follows:

[0170] x_int = int(x / ln2) Formula (5-1)

[0171] x_bas = x_int * ln2 Formula (5-2)

[0172] x_rem = x - x_bas Formula (5-3)

[0173] x = x_bas + x_rem Formula (5-4)

[0174] e x = 2 x_int * e x_rem Formula (5-5)

[0175] Let exp_bas = 2 x_int Formula (5-6)

[0176] Let exp_rem = e x_rem Formula (5-7)

[0177] e x = exp_bas * exp_rem Formula (5-8)

[0178] (1), Use SIMD instructions to calculate the formula (5-1) x_int = int(x / ln2), where x / ln2 is equivalent to x*(1 / ln2), and 1 / ln2 = 1.442695; the instructions are as follows:

[0179] Register12 = Ingenic_simd512_load_float(1.442695),

[0180] Register13 = Ingenic_simd512_float_mul(Register11, Register12),

[0181] Register13 = Ingenic_simd512_float_to_int_truncate(Register13),

[0182] First, load the floating-point number 1.442695 into Register12;

[0183] The Ingenic_simd512_float_mul instruction is a floating-point multiplication instruction. Multiply the 16 input data in Register11 and the 16 1.442695 in Register12 respectively, and place the results in Register13;

[0184] The Ingenic_simd512_float_to_int_truncate instruction is a floating-point to integer conversion instruction. This instruction converts the floating-point number in Register13 into an integer, directly discarding the data after the decimal point and retaining the integer part to obtain x_int in the formula x_int = int(x / ln2), and the result is stored in Register13;

[0185] (2), Use SIMD instructions to calculate exp_bas in the formula (5-6) exp_bas = 2 x_int of exp_bas, and the SIMD instructions are set as:

[0186] Register14 = Ingenic_simd512_load_immediate(127),

[0187] Register14 = Ingenic_simd512_int_add(Register13, Register14),

[0188] Register14 = Ingenic_simd512_shift_left(Register14, 23),

[0189] Ingenic_simd512_load_immediate is a load immediate instruction that loads the integer 127 and copies it 16 times and stores them in Register14;

[0190] Ingenic_simd512_int_add is an integer addition instruction that adds the data in Register13 and Register14, and stores the result in Register14;

[0191] Ingenic_simd512_shift_left is a left shift instruction that shifts 16 integers in Register14 to the left by 23 bits;

[0192] S5.4, use SIMD instructions to calculate e x the second part exp_rem of

[0193] (1), use SIMD instructions to calculate formula (5-3), where ln2 is 6.931472e-1, the instructions are as follows:

[0194] Register12 = Ingenic_simd512_load_float(6.931472e-1)

[0195] Register13 = Ingenic_simd512_int_to_float(Register13)

[0196] Register12 = Ingenic_simd512_float_mul(Register12, Register13)

[0197] Register11 = Ingenic_simd512_float_sub(Register11, Register12)

[0198] Ingenic_simd512_load_float loads and copies 6.931472e-1 into Register12; Ingenic_simd512_int_to_float is an instruction to convert an integer to a floating-point number, which converts 16 integers x_int in Register13 into floating-point numbers. Since in the Ingenic_simd512 instruction set, integers cannot be directly operated with floating-point numbers, the integers need to be converted to floating-point numbers first and then floating-point instructions are used for the operation;

[0199] Ingenic_simd512_float_mul multiplies 16 ln2s in Register12 with 16 data in Register13 to get x_bas, and the result is stored in Register12;

[0200] Ingenic_simd512_float_sub is a floating-point subtraction instruction that subtracts the data in Register11 from the data in Register12 to get x_rem, and the result is stored in Register11;

[0201] (2) Use the Taylor expansion to find formula (5-7). The Taylor formula for the exponential function with base e is:

[0202]

[0203] When n is 6, it meets the accuracy requirements of this method for the exp function. This method does not calculate the case where n exceeds 6, otherwise it will cause a large amount of calculation and slow down the operation speed; so e x_rem is expanded into formula (5-10):

[0204]

[0205] Formula (5-10);

[0206] In formula (5-10), 1 / 6! is 1.388889e-3, 1 / 5! is 8.333333e-3, 1 / 4! is 4.166667e-2, 1 / 3! is 1.666667e-1, and 1 / 2! is 0.5;

[0207] The steps to implement formula (5-10) using SIMD are as follows. The SIMD instructions are set as:

[0208] Register12 = Ingenic_simd512_load_float(1.388889e-3)

[0209] Register13 = Ingenic_simd512_float_mul(Register12, Register11)

[0210] Register12 = Ingenic_simd512_load_float(8.333333e-3)

[0211] Register13 = Ingenic_simd512_float_add(Register12, Register13)

[0212] Register13 = Ingenic_simd512_float_mul(Register13, Register11)

[0213] Register12 = Ingenic_simd512_load_float(4.166667e-2)

[0214] Register13 = Ingenic_simd512_float_add(Register12, Register13)

[0215] Register13 = Ingenic_simd512_float_mul(Register13, Register11)

[0216] Register12 = Ingenic_simd512_load_float(1.666667e-1)

[0217] Register13 = Ingenic_simd512_float_add(Register12, Register13)

[0218] Register13 = Ingenic_simd512_float_mul(Register13, Register11)

[0219] Register12 = Ingenic_simd512_load_float(0.5)

[0220] Register13 = Ingenic_simd512_float_add(Register12, Register13)

[0221] Register13 = Ingenic_simd512_float_mul(Register13, Register11)

[0222] Register12 = Ingenic_simd512_load_float(1.0)

[0223] Register13 = Ingenic_simd512_float_add(Register12, Register13)

[0224] Register13 = Ingenic_simd512_float_mul(Register13, Register11)

[0225] Register13 = Ingenic_simd512_float_add(Register12, Register13) Ingenic_simd512_load_float loads and copies 1.388889e-3 into Register12. Use Ingenic_simd512_float_mul to multiply the 16 x_rem data in Register11 by the 16 1.388889e-3 in Register12, and the result is stored in Register13;

[0226] Load 8.333333e-3 into Register12,

[0227] Ingenic_simd512_float_add is a floating-point addition instruction that adds the 16 data in Register13 to the 16 data in Register12, and the result is stored in Register13; then multiply Register13 by the 16 x_rem in Register11, and the result is still stored in Register13;

[0228] And so on until all the calculations in formula (5-10) are completed. The result of exp_rem obtained through SIMD instructions is finally stored in Register13;

[0229] S5.5, use SIMD instructions to multiply the exp_bas and exp_rem obtained in the above steps S5.3 and S5.4 to obtain the result of formula (5-8), expressed as

[0230] Register13 = Ingenic_simd512_float_mul(Register13, Register14)

[0231] Ingenic_simd512_float_mul is a floating-point multiplication instruction that multiplies the results of Register13 and Register14;

[0232] S5.6: Save the calculation result from the register to memory, expressed as

[0233] Result = Ingenic_simd512_store(Register13)

[0234] Ingenic_simd512_store is a data saving instruction that saves 16 data in register Register13 to the Result memory that can store 16 data.

[0235] Taking the method of solving ELU in C language as a reference, under the premise of the same input data, the precision loss between the calculation result of this method and the calculation result of the C language method is obtained using formula (2). The precision only loses 0.0001%, meeting the precision requirements of practical applications. The calculation speed of this method is 11 times that of the C language.

[0236] Precision_loss = (C_result - SIMD_result) / C_result Formula (2)

[0237] In formula (2), Precision_loss represents the calculation result of precision loss, C_result is the calculation result of the ELU method in C language, and SIMD_result is the calculation result of this method.

[0238] In summary, this method can be deployed on models such as Junzheng T40 to effectively improve the chip operation speed in parallel, especially the speed of calculating the ELU activation function.

[0239] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for improving the operation speed of a chip based on SIMD instructions, characterized in that The method includes a process of quickly solving the ELU activation function, where the prototype formula of the ELU function is shown in formula (1): The method further includes the following steps: S1. Loading data: Load in multiples of 32 bits. A register can load up to 512 bits of data at most; a single-precision floating-point number is 32 bits, and a register can load 16 floating-point numbers. Therefore, a single SIMD instruction can perform calculations on 16 floating-point numbers simultaneously; Use the SIMD instruction Ingenic_simd512_load for loading data; (float)data are the input 16 32-bit single-precision floating-point numbers, and Register1 is register 1. Save the input data in Register1, expressed as Register1 = Ingenic_simd512_load((float)data); S2. Compare the 16 (float)data in Register1 with 0: Use the comparison instruction with 0, Ingenic_simd512_fcltzw instruction. This instruction means that when the data at the corresponding position in Register1 is greater than 0, the 32 bits at the corresponding position in register Register3 are all set to 1, and when the data at the corresponding position in Register1 is less than or equal to 0, the 32 bits at the corresponding position in register Register3 are all set to 0, expressed as Registe3 = Ingenic_simd512_fcltzw(0, Register1); S3. AND Register3 and Register1: Use the AND instruction Ingenic_simd512_and instruction. The numbers greater than 0 in Register1 AND with 1 remain the original numbers, and the numbers less than or equal to 0 AND with 0 become 0. The result is saved in register Register4, expressed as Registe4 = Ingenic_simd512_and(Register3, Register1); S4. Invert each bit of the numbers in Register3, change 0 to 1 and 1 to 0, and then AND with the numbers in Register1: Use the Ingenic_simd512_andn instruction, which is an instruction to first invert Register3 and then AND with Register1. The numbers less than or equal to 0 in Register1 AND with 1 remain the original numbers, and the numbers greater than 0 AND with 0 become 0. The result is saved in register Register5, expressed as Registe5 = Ingenic_simd512_andn(Register3, Register1); S5. Use the implementation method of the exp exponential function based on SIMD instructions to find the exp value of the data in Register5, and save the result in register Register6, expressed as Register6 = Ingenic_simd512_exp(Register5); Among them, Ingenic_simd512_exp is the method name of the exp exponential function based on SIMD instructions, which means it can cover all steps of solving exp, and these steps are combined and named Ingenic_simd512_exp; S6. Subtract the data in Register6 from the floating-point number 1.0f: Use the SIMD instruction Ingenic_simd512_float_sub for floating-point subtraction to subtract the data in register Register6 from 1, and the result after subtraction is saved in register Register7, expressed as Register7 = Ingenic_simd512_float_sub(Register6, 1.0f); S7. Multiply the data in Register7 by the floating-point number α: Use the SIMD instruction Ingenic_simd512_float_mul for floating-point multiplication to multiply the data in register Register7 by α, and the result after multiplication is saved in register Register8, expressed as Register8 = Ingenic_simd512_float_mul(Register7, α); S8. The numbers greater than 0 in Register4 are the original data, and the numbers less than or equal to 0 are 0. The numbers greater than 0 in register Register8 are 0, and the numbers less than or equal to 0 are the calculated numbers. Performing an OR operation on Register4 and Register8 realizes formula (1). The Ingenic_simd512_or instruction is the OR instruction, and the result is saved in register Register9, expressed as Registe9 = Ingenic_simd512_or(Register4, Register8); S9. Save the calculation result from register Register9 to memory, expressed as Result = Ingenic_simd512_store(Register9); The Ingenic_simd512_store instruction is a data save instruction that saves 16 data in register Register9 to the Result memory that can store 16 data.

2. A method for improving the operation speed of a chip based on SIMD instructions according to claim 1, characterized in that The above step S3 realizes in formula (1): When x > 0, ELU(x) = x.

3. A method for improving the operation speed of a chip based on SIMD instructions according to claim 1, characterized in that, The above steps S4 to S7 realize in formula (1): When x ≤ 0, ELU(x) = α(e x - 1).

4. A method for improving the computing speed of a chip based on SIMD instructions according to claim 1, characterized in that, The Ingenic_simd512_exp in the above step S5 further includes: S5.1, load data, use the SIMD instruction as follows: Register11 = Ingenic_simd512_load((float)data), Among them, Ingenic_simd512_load is a data loading instruction; (float)data are 16 32-bit single-precision floating-point numbers as input; Register11 is register 11, which stores the input data; S5.2, limit the range of the input data between -88.7 and 88.7; use the SIMD instructions as follows: Register12 = Ingenic_simd512_load_float(-88.7), Register13 = Ingenic_simd512_load_float(88.7), Register11 = Ingenic_simd512_max(Register11, Register12), Register11 = Ingenic_simd512_min(Register11, Register13), Ingenic_simd512_load_float is an instruction to load and copy floating-point numbers. After using this instruction, there are 16 -88.7s in Register12 and 16 88.7s in register Register13; Ingenic_simd512_max is a maximum value taking instruction. This instruction compares the 16 data in Register11 with the 16 -88.7s in register Register12 respectively to take the maximum value, that is, the data greater than -88.7 remains unchanged, and the data less than -88.7 takes -88.

7. The result is stored in Register11; Ingenic_simd512_min is a minimum value taking instruction. This instruction compares the 16 data in the result Register11 from the previous step with the 16 88.7s in Register13 respectively to take the minimum value, that is, the data less than 88.7 remains unchanged, and the data greater than 88.7 takes 88.

7. The result is still stored in Register11; S5.3, split e x into two parts for multiplication, where x is the input data after range limitation in step S5.

2. Split x into two parts for addition, and split e x into two parts for multiplication. The formula is as follows: x_int = int(x / In2) Formula (5-1) x_bas = x_int * ln2 Formula (5-2) x_rem = x - x_bas Formula (5-3) x = x_bas + x_rem Formula (5-4) e x = 2 x_int * e x_rem Equation (5-5) Let exp_bas = 2 x_int Formula (5-6) Let exp_rem = e x_rem Formula (5-7) e x = exp_bas * exp_rem Equation (5-8) (1), use the SIMD instruction to calculate x_int = int(x / ln2) in Formula (5-1), where x / ln2 is equivalent to x * (1 / ln2), and 1 / ln2 = 1.442695; the instructions are as follows: Register12 = Ingenic_simd512_load_float(1.442695), Register13 = Ingenic_simd512_float_mul(Register11, Register12), Register13 = Ingenic_simd512_float_to_int_truncate(Register13), First, load the floating-point number 1.442695 into Register12; The Ingenic_simd512_float_mul instruction is a floating-point multiplication instruction. It multiplies the 16 input data in Register11 with the 16 1.442695 in Register12, and the result is placed in Register13; The Ingenic_simd512_float_to_int_truncate instruction is a floating-point to integer conversion instruction. This instruction converts the floating-point numbers in Register13 into integers, directly discarding the data after the decimal point and retaining the integer part to obtain x_int in the formula x_int = int(x / ln2), and the result is stored in Register13; (2), Use SIMD instructions to calculate exp_bas in the formula (5 - 6) exp_bas = 2^x_int. The SIMD instructions are set as follows: Register14 = Ingenic_simd512_load_immediate(127), Register14 = Ingenic_simd512_int_add(Register13, Register14), Register14 = Ingenic_simd512_shift_letf(Register14, 23), The Ingenic_simd512_load_immediate is a load immediate instruction that loads the integer 127 and duplicates it 16 times and stores them in Register14; The Ingenic_simd512_int_add is an integer addition instruction that adds the data in Register13 and Register14, and the result is placed in Register14; The Ingenic_simd512_shift_letf is a left shift instruction that shifts the 16 integers in Register14 to the left by 23 bits; S5.4, Using SIMD instructions, find the second part exp_rem of e x ​ (1), Use SIMD instructions to calculate the formula (5 - 3), where ln2 is 6.931472e-1. The instructions are as follows: Register12 = Ingenic_simd512_load_float(6.931472e-1) Register13 = Ingenic_simd512_int_to_float(Register13) Register12 = Ingenic_simd512_float_mul(Register12, Register13) Register11 = Ingenic_simd512_float_sub(Register11, Register12) Ingenic_simd512_load_float loads and copies 6.931472e-1 into Register12; Ingenic_simd512_int_to_float is an instruction for converting integers to floating-point numbers, which converts 16 integers x_int in Register13 to floating-point numbers. Since in the Ingenic_simd512 instruction set, integers cannot be directly operated on with floating-point numbers, the integers need to be converted to floating-point numbers first and then floating-point instructions are used for the operation; Ingenic_simd512_float_mul multiplies the 16 ln2s in Register12 by the 16 data in Register13 to obtain x_bas, and the result is stored in Register12; Ingenic_simd512_float_sub is a floating-point subtraction instruction that subtracts the data in Register11 from the data in Register12 to obtain x_rem, and the result is stored in Register11; (2) Use the Taylor expansion to find formula (5-7). The Taylor formula for the exponential function with base e is: When n is 6, the accuracy requirement of the present method for the exp function is satisfied. The present method does not calculate the case where n exceeds 6, otherwise it will cause a large amount of calculation and slow down the operation speed; so e x_rem is expanded into formula (5-10): In formula (5-10), 1 / 6! is 1.388889e-3, 1 / 5! is 8.333333e-3, 1 / 4! is 4.166667e-2, 1 / 3! is 1.666667e-1, and 1 / 2! is 0.5; The steps to implement formula (5-10) using SIMD are as follows. The SIMD instructions are set as: Register12 = Ingenic_simd512_load_float(1.388889e-3) Register13 = Ingenic_simd512_float_mul(Register12, Register11) Register12 = Ingenic_simd512_load_float(8.333333e-3) Register13 = Ingenic_simd512_float_add(Register12, Register13) Register13 = Ingenic_simd512_float_mul(Register13, Register11) Register12 = Ingenic_simd512_load_float(4.166667e-2) Register13 = Ingenic_simd512_float_add(Register12, Register13) Register13 = Ingenic_simd512_float_mul(Register13, Register11) Register12 = Ingenic_simd512_load_float(1.666667e-1) Register13 = Ingenic_simd512_float_add(Register12, Register13) Register13 = Ingenic_simd512_float_mul(Register13, Register11) Register12 = Ingenic_simd512_load_float(0.5) Register13 = Ingenic_simd512_float_add(Register12, Register13) Register13 = Ingenic_simd512_float_mul(Register13, Register11) Register12 = Ingenic_simd512_load_float(1.0) Register13 = Ingenic_simd512_float_add(Register12, Register13) Register13 = Ingenic_simd512_float_mul(Register13, Register11) Register13 = Ingenic_simd512_float_add(Register12, Register13) Ingenic_simd512_load_float loads and copies 1.388889e-3 into Register12, multiplies the 16 x_rem data in Register11 by the 16 1.388889e-3 in Register12 using Ingenic_simd512_float_mul, and saves the result in Register13; Load 8.333333e-3 into Register12, Ingenic_simd512_float_add is a floating-point addition instruction that adds the 16 data in Register13 to the 16 data in Register12 and saves the result in Register13; then multiplies Register13 by the 16 x_rem in Register11, and the result is still saved in Register13; And so on until all the calculations in formula (5-10) are completed. The result of exp_rem calculated by SIMD instructions is finally saved in Register13; S5.5, use SIMD instructions to multiply exp_bas obtained in the above steps S5.3 and S5.4 by exp_rem to obtain the result of formula (5-8), expressed as Register13 = Ingenic_simd512_float_mul(Register13, Register14) Ingenic_simd512_float_mul is a floating-point multiplication instruction that multiplies the results of Register13 and Register14; S5.6: Save the calculation result from the register to the memory, expressed as Result = Ingenic_simd512_store(Register13) Ingenic_simd512_store is a data storage instruction that stores 16 data in register Register13 into the Result memory that can store 16 data.

5. A method for improving the operation speed of a chip based on SIMD instructions according to claim 1, characterized in that In the method, the instruction set based on SIMD instructions can load 512-bit data into a 512-bit register at one time, and can load 1024-bit cache data with two registers at one time; there are a total of 32 registers available for use and they can be reused; during calculation, the same operation can be performed on 512-bit data at one time. This instruction set includes arithmetic operations such as addition, subtraction, multiplication, absolute value, maximum value, and minimum value, logical operations such as AND, OR, NOT, XOR, and shift, and comparison operations such as less than, equal to, and less than or equal to, meeting general operation requirements. The exp exponential function requires a combination of multiple instructions.