Function calculation method and device, chip, calculation equipment and storage medium
By expanding the function into a polynomial form and multiplication using registers, the problem of complex transcending functions in the prior art is solved, and the effect of simplifying operations and improving calculation efficiency is achieved.
Patent Information
- Application Number
- CN202311454000.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-06
AI Technical Summary
When existing computing devices face complex transcendent functions in potential functions, they directly call mathematical library solutions to lead to complex calculations and high clock cycles, which becomes the main computing bottleneck.
By expanding the function in the form of a polynomial, converting it into a polynomial, and multiplication is performed using registers to simplify the operation and improve the calculation efficiency of the function.
The function calculation process is simplified, the computing power requirement is reduced, the calculation efficiency is improved, and the calculation error is reduced.
Smart Images

Figure CN119939099A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent computing technology, and in particular to a function computing method, apparatus, chip, computing device and storage medium. Background Art
[0002] The N-body problem refers to finding the subsequent motion of multiple objects with known initial positions, velocities, and masses under classical mechanics. The relevant calculations of the N-body problem are involved in the fields of molecular dynamics, industrial processes, astronomical simulations, etc., and often occupy a hot position.
[0003] The N-body problem often involves the calculation of potential functions, and a large number of complex transcendental functions are often used in potential functions. Among them, the relationship between variables in transcendental functions cannot be expressed by a finite number of addition, subtraction, multiplication, division, exponentiation, and square root operations. For example, logarithmic functions, inverse trigonometric functions, exponential functions, and trigonometric functions are all transcendental functions.
[0004] When facing potential functions, current computing devices directly calculate according to the formula. If they encounter complex transcendental functions in potential functions, they directly call the math library to solve them. However, calling the math library to solve the calculation is complex (pure mathematical approximation, precise solution), and the number of clock cycles required to solve transcendental functions such as exp, sin, cos, erfc is high, so this part of the calculation often becomes the main computing bottleneck. Summary of the invention
[0005] The present application provides a function calculation method, apparatus, chip, computing device and storage medium, which can improve function calculation efficiency.
[0006] In a first aspect, the present application provides a function calculation method: first, the coefficients of the polynomial obtained by expanding the first function in polynomial form are loaded into the first register, and different powers of the first target variable value are loaded into the second register, and then, a multiplication operation is performed based on the first register and the second register to obtain a calculation result. The first function is a function to be calculated, which includes a transcendental function term, the first target variable value is within the value range of the independent variable in the first function, and the calculation result includes the function calculation value of the first function when the independent variable in the first function takes the value of the first target variable.
[0007] It should be understood that the first function includes transcendental function terms. If the mathematical library is directly called to solve the first function, the calculation will be more complicated and the number of clock cycles required for the solution will be higher. Therefore, this solution transforms the first function into a polynomial form by expanding it in a polynomial form (i.e., polynomial fitting), and then converts the solution of the first function into the solution of the polynomial, thereby simplifying the operation, improving the efficiency of function calculation, and reducing the computing power requirement. As for the solution of the polynomial, it can be implemented based on registers, storing the coefficients of the polynomial obtained after the conversion of the first function into the first register, and storing the different powers of the first target variable value (i.e., a certain target variable value) of the independent variable in the first function into the second register, and then performing multiplication operations based on the first register and the second register, the function calculation value of the first function when the independent variable in the first function takes the first target variable value can be quickly obtained.
[0008] Based on the first aspect, in a possible implementation scheme, the first function can be Taylor expanded at multiple values to obtain multiple groups of coefficients, wherein the multiple groups of coefficients correspond to the multiple values one by one, the multiple values are all within the value range of the independent variable in the first function, and each group of coefficients includes the coefficients of the polynomial obtained by Taylor expanding the first function at the corresponding value. Then, it is determined which value of the multiple values is closest to the first target variable value (i.e., the gap is the smallest), and then part or all of the group of coefficients corresponding to the value closest to the first target variable value (i.e., the target group of coefficients) is loaded into the first register, which helps to improve the calculation accuracy (reduce the calculation error).
[0009] Based on the first aspect, in a possible implementation manner, one or more of the number of coefficients in the target group of coefficients loaded into the first register, the number of coefficients contained in each group of coefficients in the plurality of groups of coefficients (reflecting the order of the polynomial expansion), and the number of groups of the plurality of groups of coefficients (that is, the number of the plurality of values) may be determined according to at least one of the following:
[0010] (1) The computational accuracy requirement of the first function (for example, it can be measured by accuracy level, precision value, or error level, etc.);
[0011] (2) The bit width of the first register.
[0012] For example, assuming that the number of terms of the polynomial corresponding to the target group coefficient is 6 and the highest order is 5, the target group coefficient contains the coefficients of all terms in the polynomial. If the calculation accuracy requirement of the first function is high, all coefficients in the target group coefficient can be loaded into the first register for subsequent calculation. If the calculation accuracy requirement of the first function is high, some coefficients in the target group coefficient can be loaded into the first register for subsequent calculation.
[0013] Based on the first aspect, in a possible implementation, the first register includes one or more vector registers, and the number and / or bit width of the one or more vector registers may be determined according to at least one of the following:
[0014] (1) The computational accuracy requirement of the first function (for example, it can be measured by accuracy level, accuracy value or error level, etc.);
[0015] (2) The number of coefficients contained in the target group coefficients;
[0016] (3) The number of terms in the polynomial corresponding to the target group coefficients;
[0017] (4) The highest order of the polynomial corresponding to the target group coefficients.
[0018] For example, suppose the polynomial order corresponding to the target group coefficients is 9th order, and the number of coefficients contained in the target group coefficients is 10. If you want to use all 10 coefficients for calculation, the bit width of the requested vector register must be at least able to store these 10 coefficients. If the calculation accuracy requirement of the first function is low (i.e., the allowable error is slightly larger), only the first few coefficients of the 10 coefficients can be loaded into the vector register, such as only the first 6 coefficients. Compared with solving the 6th-order expanded polynomial of the first function, the bit width of the requested vector register can store these 6 coefficients.
[0019] Based on the first aspect, in a possible implementation scheme, the first function may be first identified from the application code, and then the above-mentioned multiple sets of coefficients may be generated and recorded according to the first function. The identification method may be to automatically identify the first function through a function introduction identifier, or to automatically identify the first function directly from the code source, which is not specifically limited here.
[0020] Based on the first aspect, in a possible implementation scheme, when the number of coefficients included in the target group coefficients is less than or equal to a threshold, the multiplication operation is a matrix multiplication operation; when the number of coefficients included in the target group coefficients is greater than the threshold, the multiplication operation is a vector inner product operation.
[0021] It should be understood that if the number of coefficients contained in the target group coefficients is small, then the matrix multiplication operation is performed based on the first register and the second register, and the utilization rate of the first register can be achieved relatively high. If the number of coefficients contained in the target group coefficients is large, then the matrix multiplication operation is performed based on the first register and the second register, and the utilization rate of the first register cannot be achieved relatively high. Therefore, when the number of coefficients contained in the target group coefficients is large, it is more appropriate to select the vector inner product operation based on the first register and the second register, and the fill rate of the first register can be achieved relatively high.
[0022] Based on the first aspect, in a possible implementation, a matrix multiplication operation is performed based on the first register and the second register to obtain an intermediate result, and the intermediate result and the second part of the coefficients in the target group coefficients are loaded into a third vector register, wherein the first register includes the first part of the coefficients in the target group coefficients, and the order corresponding to the first part of the coefficients is higher than the order corresponding to the second part of the coefficients. Then, a vector inner product operation is performed based on the third vector register and the fourth vector register to obtain a calculation result, wherein the fourth vector register is loaded with different powers of the first target variable value.
[0023] Based on the first aspect, in a possible implementation, different powers of the second target variable value can also be loaded into the second register, wherein the second target variable value is within the value range of the independent variable in the first function. Then, a multiplication operation is performed based on the first register and the second register. At this time, the calculation result obtained includes not only the function calculation value of the first function when the independent variable in the first function takes the value of the first target variable, but also the function calculation value of the first function when its independent variable takes the value of the second target variable. In other words, the results of the first function under two different independent variable values can be calculated.
[0024] Based on the first aspect, in a possible implementation scheme, similar to the first function, the second function is also a function containing transcendental function terms, and the coefficients of the polynomial obtained by expanding the second function in polynomial form can also be loaded into the first register, and the third target variable value can be loaded into the second register, wherein the third target variable value is within the value range of the independent variable in the second function. Then, a multiplication operation is performed based on the first register and the second register, and the calculation result obtained at this time includes not only the function calculation value of the first function when the independent variable in the first function takes the value of the first target variable, but also the function calculation value of the second function when the independent variable in the second function takes the value of the third target variable. In other words, the operation results of different functions can be calculated.
[0025] Based on the first aspect, in a possible implementation scheme, the first function is a potential function.
[0026] In a second aspect, the present application also provides a function calculation device, including a control unit and an operation unit. The control unit is used to load the coefficients of the polynomial obtained by expanding the first function in polynomial form into a first register, wherein the first function includes a transcendental function term. The control unit is also used to load different powers of the first target variable value into a second register, wherein the first target variable value is within the value range of the independent variable in the first function. The operation unit is used to perform a multiplication operation based on the first register and the second register to obtain a calculation result, wherein the calculation result includes the function calculation value of the first function when the independent variable in the first function takes the value of the first target variable.
[0027] The function computing device may also include more or fewer units / modules, which are not specifically limited here. The function computing device in the second aspect is specifically used to execute a method of any implementation scheme of the function computing method in the first aspect, which can be referred to in the above description and will not be repeated here.
[0028] In a third aspect, the present application also provides a chip, including an operator, a controller and multiple registers, and the chip is used to execute a method of any implementation scheme of the function calculation method in the first aspect, which can be seen in the previous introduction and will not be repeated here.
[0029] In a fourth aspect, the present application further provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory. The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method of any implementation scheme in the first aspect.
[0030] In a fifth aspect, the present application further provides a computer-readable storage medium comprising computer program instructions. When the above-mentioned computer program instructions are executed by a computing device cluster (including at least one computing device), the computing device cluster executes a method as in any implementation scheme in the first aspect.
[0031] In a sixth aspect, the present application further provides a computer program product comprising instructions. When the instructions are executed by a computing device cluster (including at least one computing device), the computing device cluster executes the method of any embodiment of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments.
[0033] Figure 1 It is a flowchart of a function calculation method provided in an embodiment of the present application;
[0034] Figure 2 It is a schematic diagram of dividing the truncation radius provided in an embodiment of the present application;
[0035] Figure 3 is a schematic diagram of performing a vector inner product operation based on a first register and a second register provided in an embodiment of the present application;
[0036] Figure 4 It is a schematic diagram of filling a variable matrix and a coefficient matrix provided in an embodiment of the present application;
[0037] Figure 5 It is another schematic diagram of filling a variable matrix and a coefficient matrix provided in an embodiment of the present application;
[0038] Figure 6 It is another schematic diagram of filling a variable matrix and a coefficient matrix provided in an embodiment of the present application;
[0039] Figure 7 It is another schematic diagram of filling a variable matrix and a coefficient matrix provided in an embodiment of the present application;
[0040] Figure 8 It is a schematic diagram of calculating multiple sets of polynomials provided in an embodiment of the present application;
[0041] Fig. 9 is a performance comparison chart of SVE and SME provided in an embodiment of the present application;
[0042] Fig.10 It is a schematic diagram of a function calculation process provided by an embodiment of the present application;
[0043] Fig.11 is a structural schematic diagram of a function calculation device provided in an embodiment of the present application;
[0044] Fig.12 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0045] Fig.13 is a schematic diagram of a computing device cluster provided in an embodiment of the present application;
[0046] Fig.14 is a schematic diagram of a scenario in which two computing devices interact via a network provided in an embodiment of the present application;
[0047] Fig.15 It is a schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to facilitate understanding of the technical solutions in the embodiments of the present application, some terms and concepts involved in the embodiments of the present application are briefly introduced below.
[0049] 1. Transcendental functions
[0050] Transcendental functions are functions that do not satisfy any polynomial equation with polynomials as coefficients. The relationship between variables in a transcendental function cannot be expressed by a finite number of addition, subtraction, multiplication, division, exponentiation, or square root operations. For example, logarithmic functions, inverse trigonometric functions, exponential functions, and trigonometric functions are all transcendental functions.
[0051] 2. Potential function
[0052] Potential function is a function that describes the energy and force in a physical system, and its value is related to information such as position, time, velocity, and acceleration in the physical system. In the definition of potential function, there are two different types of potential functions, namely potential energy function and potential field function. The potential energy function is a function that describes the energy of an object at a certain position, while the potential field function is a function that describes the properties of the field in which the object is located.
[0053] Transcendental functions are often used in potential functions. For example, in industrial processes, the calculation formula for the tersoff potential function between atoms is shown in formula (1):
[0054]
[0055] In astronomical simulation, the calculation formula of the gravitational potential caused by universal gravitation is shown in formula (2):
[0056]
[0057] 3. Scalable vector extension (SVE)
[0058] SVE is a vector instruction set in the ARM architecture, which is specially designed for vector computing. It provides strong support for high-performance computing and offers higher parallelism and flexibility.
[0059] The characteristic of the SVE vector register is that its length is scalable and can be 128, 256, 512, 1024 or 2048 bits.
[0060] SVE supports multiple data types and operations, including integers, floating-point numbers, and complex numbers, and can perform basic vector operations such as addition, subtraction, multiplication, and division.
[0061] For example, suppose there are two 128-bit SVE vector registers, each storing four 32-bit integers. Using the SVE addition instruction, these four integers can be added simultaneously.
[0062] 4. Scalable matrix extension (SME)
[0063] SME is an extension of SVE and SVE2. It is closely integrated with SVE and provides higher parallelism and flexibility. SME inherits the scalability of SVE and supports operations of various lengths and bit widths.
[0064] The main purpose of SME is to improve the performance of matrix multiplication, especially in deep learning and other applications that require a large number of matrix multiplications.
[0065] SME is mainly optimized for matrix multiplication operations. Matrix multiplication is a key operation in many applications, such as deep learning, graphics processing, and high performance computing (HPC). SME provides efficient matrix multiplication instructions that can accelerate the execution of these applications.
[0066] For example, suppose there are two matrices A and B, each stored in an SME matrix register. Using the SME matrix multiplication instruction, the product of these two matrices can be efficiently calculated and the result stored in another SME matrix register.
[0067] See also Figure 1 , Figure 1 It is a flowchart of a function calculation method provided in an embodiment of the present application, including steps S101 to S103.
[0068] S101. Load coefficients of a polynomial obtained by expanding a first function in polynomial form into a first register, wherein the first function includes a transcendental function term.
[0069] Transcendental functions are used in the above-mentioned first function, and the first function may include one or more transcendental function terms. The embodiment of the present application does not specifically limit the type of the first function. Optionally, the first function may be a potential function, which involves the use of transcendental functions. For potential functions and transcendental functions, please refer to the above description, which will not be repeated here.
[0070] Specifically, step S101 transforms the first function into a polynomial form by expanding the first function in a polynomial form (i.e., polynomial fitting), and then transforms the solution of the first function into the solution of the polynomial, thereby simplifying the operation, improving the function calculation efficiency, and reducing the computing power requirement. As for the solution of the polynomial, it can be implemented based on registers. Step S101 stores the coefficients of the polynomial obtained after the first function is converted into a first register, and stores different powers of the first target variable value of the independent variable in the first function into a second register (see step S102), and then performs a multiplication operation based on the first register and the second register (see step S103), so as to quickly obtain the function calculation value of the first function when its independent variable takes the value of the first target variable.
[0071] The embodiment of the present application does not specifically limit how to expand the first function into a polynomial form. For example, the first function may be subjected to Taylor expansion (the order and position of the expansion may be selected according to the actual application scenario) to convert the first function into a polynomial form.
[0072] Optionally, the first function can be Taylor expanded at multiple values to obtain multiple sets of coefficients, wherein the multiple sets of coefficients correspond to the multiple values one by one, the multiple values are all within the value range of the independent variable in the first function, and each set of coefficients includes the coefficients of the polynomial obtained by Taylor expanding the first function at the corresponding value. Then, it is determined which value of the multiple values is closest to the first target variable value (i.e., the gap / difference is the smallest), and then part or all of the set of coefficients corresponding to the value closest to the first target variable value (i.e., the target set of coefficients) is loaded into the first register.
[0073] That is to say, the first function can be Taylor expanded at multiple different values within the range of its independent variable value to obtain a set of coefficients corresponding to each value for subsequent calculation selection. Then, according to the actual calculation requirements, that is, when the value of the current independent variable is the value of the first target variable, the set of coefficients corresponding to a value closest to the first target variable value is selected as the target set of coefficients, which helps to improve the calculation accuracy (reduce the calculation error), and then some or all of the coefficients in the target set of coefficients are loaded into the first register for calculating the first function.
[0074] For example, in the N-body problem, end-to-end / peer-to-peer (P2P) computing is a very common computing mode in short-range interactions and often occupies a hot spot. Its calculation formula can be found in formula (3):
[0075]
[0076] Where x is the independent variable in the T function of equation (3), which represents the distance between particles. It can be seen that the T function is relatively complex and contains two transcendental function terms, namely erfc(x) and It involves the complementary error function (erfc) and the exponential function.
[0077] P2P calculations often have a cut-off radius / cut-off distance, denoted by r cut , that is, the range of values of x. cut Divide into N parts, and record the value of the i-th point as x i , where N is a positive integer, which can be selected according to the actual application scenario. Figure 2 As shown, assuming r cut The range is 0 to 3. cut Divide it into 512 equal parts, and record the value of the i-th point as x i , i∈[1,512].
[0078] Take the T function in formula (3) as the first function, and calculate the T function at x iTaylor expansion is performed at , thus converting it into the polynomial form of formula (4):
[0079]
[0080] Among them, x i is the i-th value point in the range of x, ∈=xx i Formula (4) is based on the expansion to the fourth order as an example, and can actually be expanded to a higher order.
[0081] It should be understood that in theory, the Taylor expansion can be expanded to an infinite order, thus infinitely approaching the true value of the T function. The value of ∈ largely determines the speed of convergence. When ∈ is small, the convergence is faster and can converge to the true value within 4-5 orders. When ∈ is large, the convergence order is higher and the cost of polynomial calculation is higher. The larger N is, the more numerical points are divided, the smaller ∈ can be obtained, and the smaller the order required for convergence. The larger N is, the more E needs to be stored. i and x i The more N is, the more expensive it is to obtain data. You can first determine the value of N manually or automatically, and then determine the order p of the polynomial expansion based on the accuracy requirements.
[0082] Optionally, one or more of the number of coefficients in the target group of coefficients loaded into the first register, the number of coefficients contained in each group of the plurality of groups of coefficients (reflecting the order of the polynomial expansion), and the number of groups of the plurality of groups of coefficients (that is, the number of the plurality of numerical values) may be determined according to at least one of the following:
[0083] (1) The computational accuracy requirement of the first function (for example, it can be measured by accuracy level, accuracy value or error level, etc.);
[0084] (2) The bit width of the first register.
[0085] For example, assuming that the number of terms of the polynomial corresponding to the target group coefficient is 8 and the highest order is 7, that is, the target group coefficient contains the coefficients of the polynomial obtained by expanding the first function to the 7th order, and the target group coefficient contains the coefficients of all terms in the polynomial. If the calculation accuracy requirement of the first function is high, all coefficients in the target group coefficient can be loaded into the first register for subsequent calculations. If the calculation accuracy requirement of the first function is low, only some coefficients (corresponding to lower orders) in the target group coefficient can be loaded into the first register for subsequent calculations.
[0086] For another example, assuming that the order of Taylor expansion is 5 (can be selected manually or automatically), check the error between the value calculated by formula (3) (i.e., directly calculating the first function and accurately solving it) and the value calculated by the Taylor expansion polynomial of formula (3): When N = 4, that is, the range of x is divided into 4 parts, there are 4 different x i (Numerical point), the error is at the order of 1.0E-3; when N=8, the error is at the order of 1.0E-4; when N=16, the error is at the order of 1.0E-7; when N=32, the error is at the order of 1.0E-8.
[0087] If the error level of the T function of formula (3) is to be guaranteed to be in the order of 1.0E-7, it can be determined that N is 16, which is a more appropriate choice, that is, the value range of x is divided into 16 parts, with 16 numerical points, and the T function is expanded by 5th order Taylor at each numerical point, and the coefficients of the obtained polynomial are recorded. A total of 16 groups of coefficients need to be recorded, and each group of coefficients contains 6 coefficients (the order of the polynomial is 5 and the number of terms is 6). The storage location of each group of coefficients is not specifically limited in the embodiment of the present application, for example, it can be stored in the form of an array, or stored in a certain memory.
[0088] For another example, assume that the number of terms of the polynomial corresponding to the target group coefficients is 7 and the highest order is 6, that is, the target group coefficients contain the coefficients of the polynomial obtained by expanding the first function to the 6th order. If the bit width of the first register supports the storage of 7 coefficients, all coefficients in the target group coefficients can be loaded into the first register; if the bit width of the first register only supports the storage of 6 coefficients, only the first 6 coefficients (corresponding to 0 to 5 order terms) in the target group coefficients can be loaded into the first register, which is equivalent to solving according to the 5th order expanded polynomial of the first function.
[0089] Optionally, the first register includes one or more vector registers, and the number and / or bit width of the one or more vector registers may be determined according to at least one of the following:
[0090] (1) The computational accuracy requirement of the first function (for example, it can be measured by accuracy level, accuracy value or error level, etc.);
[0091] (2) The number of coefficients contained in the target group coefficients;
[0092] (3) The number of terms in the polynomial corresponding to the target group coefficients;
[0093] (4) The highest order of the polynomial corresponding to the target group coefficients.
[0094] For example, suppose the polynomial order corresponding to the target group coefficients is 9th order, and the number of coefficients contained in the target group coefficients is 10. If you want to use all 10 coefficients for calculation, the bit width of the requested vector register must be at least able to store these 10 coefficients. If the calculation accuracy requirement of the first function is low (i.e., the allowable error is slightly larger), only the first few coefficients of the 10 coefficients (corresponding to the lower order) can be loaded into the vector register, such as only the first 6 coefficients. Compared with solving the 5th-order polynomial expansion of the first function, the bit width of the requested vector register can store these 6 coefficients.
[0095] For another example, assuming that the bit width of each vector register is fixed, a sufficient number of vector registers can be applied for so that the sum of the bit widths of these applied vector registers is sufficient to support the storage of all coefficients included in the target group of coefficients.
[0096] Optionally, before step S101, the first function can be first identified from the application code, and then the above-mentioned multiple sets of coefficients can be generated and recorded according to the first function for use in subsequent calculations. The identification method can be to automatically identify the first function through the function introduction mark, or to automatically identify the first function directly from the code source code, which is not specifically limited in the embodiments of the present application. For example, a function containing transcendental function terms can be identified as the first function, a function containing transcendental function terms whose number is greater than or equal to a threshold can be identified as the first function, a function whose complexity is higher than a preset value can be identified as the first function, and so on. Regarding the storage location of the above-mentioned multiple sets of coefficients, the embodiments of the present application do not make specific limitations.
[0097] S102. Load different powers of the first target variable value into the second register, wherein the first target variable value is within the value range of the independent variable in the first function.
[0098] It should be understood that the first target variable value is a target value of the independent variable in the first function, and currently it is necessary to calculate the result of the first function when the independent variable thereof takes the value of the first target variable.
[0099] Optionally, the second register may be a vector register of a fixed length or a vector register of a non-fixed length, such as an SVE vector register, whose length is scalable. An SVE vector register of a suitable length may be applied for as the second register according to actual needs, so as to be used for loading different powers of the first target variable value, so as to be used for calculation in the subsequent step S103. As to which power values of the first target variable value are loaded into the second register, it is necessary to judge according to the actual calculation situation, which will be introduced in step S103.
[0100] Similarly, the first register in step S101 can apply for an SVE vector register of appropriate length as the first register according to actual needs, so as to load part or all of the coefficients in the target group coefficients, so as to be used for the calculation of the subsequent step S103.
[0101] S103. Perform a multiplication operation based on the first register and the second register to obtain a calculation result, wherein the calculation result includes a function calculation value of the first function when the independent variable in the first function takes the value of the first target variable.
[0102] In a first possible implementation, the multiplication operation is a vector inner product operation, that is, the calculation of the polynomial is implemented by multiplying the vector by the first function.
[0103] Specifically, assuming that the first function is expanded in the form of a p-order polynomial, the polynomial obtained is as shown in formula (5):
[0104] A(x)=a0+a1x+a2x 2 +...+a p x p (5)
[0105] Where x represents the independent variable in A(x), a p x p It is the highest order term in the polynomial A(x), and p is a positive integer, which represents the highest order of A(x).
[0106] Assuming that we want to calculate the function value of the first function when the independent variable x is X1 (within the value range of x, corresponding to the value of the first target variable), we can replace the coefficients a0, a1, a2, ...a in A(x) with p Part or all of X1 is loaded into the first register, and part or all of the 0 to p power values of X1 is loaded into the second register, and then a vector inner product operation is performed based on the first register and the second register to obtain the above function value.
[0107] For example, assuming that the first function is expanded by 5th order at multiple different values within the range of its independent variable, the form of each polynomial obtained is as shown in formula (6):
[0108] A(x)=a0+a1x+a2x 2 +a3x 3 +a4x 4 +a5x 5 (6)
[0109] The coefficients of the polynomial expanded at each numerical point are recorded to obtain multiple groups of coefficients, each group of coefficients corresponding to a numerical value within the range of the independent variable. Assuming that the function value of the first function is to be calculated when the independent variable x takes the values of X1 and X2 respectively, a group of coefficients corresponding to a numerical value closest to X1 among multiple numerical points can be loaded into the first register (which can be composed of multiple SVE vector registers), and a group of coefficients corresponding to a numerical value closest to X2 among multiple numerical points can also be loaded into the first register. At the same time, the 0 to 5 power values of X1 are loaded into multiple SVE vector registers (as the second register), and the 0 to 5 power values of X2 are also loaded into the second register, and then the vector inner product operation is performed based on the first register and the second register, so that the function value of the first function at different independent variable values can be calculated.
[0110] Optionally, the coefficients of the polynomial obtained by expanding the second function in polynomial form can also be loaded into the first register, and the third target variable value can be loaded into the second register, wherein the third target variable value is within the value range of the independent variable in the second function. Then, a vector multiplication (vector inner product) operation is performed based on the first register and the second register. At this time, the calculation result obtained includes not only the function calculation value of the first function when the independent variable in the first function takes the value of the first target variable, but also the function calculation value of the second function when the independent variable in the second function takes the value of the third target variable. In other words, multiple sets of polynomials corresponding to different functions can be calculated simultaneously, and then the operation results of different functions can be calculated.
[0111] For example, similar to the first function described above, the second function is also a function containing a transcendental function term. Assuming that the second function is expanded in the form of a fifth-order polynomial, the polynomial obtained is as shown in formula (7):
[0112] B(y)=b0+b1y+b2y 2 +...+b5y 5 (7)
[0113] Where y represents the independent variable in B(y), b5y 5 is the highest order term in the polynomial B(y), indicating that the highest order of B(y) is 5.
[0114] Suppose we want to calculate the function value of the first function when the independent variable x is X1 (within the value range of x, corresponding to the value of the first target variable), and we also want to calculate the function value of the second function when the independent variable y is Y1 (within the value range of y), such as Figure 3 As shown, the coefficients a1, a2, ...a in A(x) can be pStore it in the first register (here composed of multiple vector registers v1~v5), and store the coefficients b1, b2...b5 in B(y) in the first register. Similarly, load the 1st to 5th power values of X1 into the second register (here composed of multiple vector registers v x1 ~v x5 Then, the vector inner product operation is performed based on the first register and the second register to obtain the intermediate results T1 and T2. T1 represents The calculated value of T2 is Then the above calculated value is loaded into a vector register, and a0 and b0 are loaded into another vector register. By performing the addition operation (T1+a0, T2+b0) on the two vector registers, the function value S1 of the first function when the independent variable x is X1 and the function value S2 of the second function when the independent variable y is Y1 can be obtained. In other words, multiple sets of polynomials can be calculated simultaneously by the above method to obtain corresponding calculation results.
[0115] In a second possible embodiment, the multiplication operation is a matrix multiplication operation, that is, the calculation of the polynomial corresponding to the function is implemented by matrix multiplication.
[0116] Specifically, assuming that the polynomial obtained by expanding the first function in the form of a p-order polynomial is as shown in equation (5), the coefficients a0, a1, a2, ...a in A(x) can be p Part or all of it is loaded into the first register, and part or all of the 0 to p power values of X1 are loaded into the second register, and then a matrix multiplication operation is performed based on the first register and the second register to obtain the function value of the first function when the independent variable x is X1 (within the value range of x, corresponding to the value of the first target variable).
[0117] In a possible implementation, a matrix multiplication operation is performed based on the first register and the second register to obtain an intermediate result, and the intermediate result and the second part of the coefficients in the target group of coefficients are loaded into a third vector register, wherein the first register includes the first part of the coefficients in the target group of coefficients, and the order corresponding to the first part of the coefficients is higher than the order corresponding to the second part of the coefficients. Then, a vector inner product operation is performed based on the third vector register and the fourth vector register to obtain a calculation result, wherein the fourth vector register is loaded with different powers of the first target variable value.
[0118] For example, the polynomial A(x) in formula (6) can be decomposed into the form of formula (8):
[0119] A(x)=(a0+a1x)+x 2 (a2+a3x)+x 4(a4+a5x) (8)
[0120] It can be seen that, in fact, two adjacent terms in equation (6) are merged to obtain equation (8).
[0121] Suppose we want to calculate the function value of the first function when the independent variable x takes the value of X1 (within the value range of x, corresponding to the value of the first target variable), such as Figure 4 As shown, some coefficients a2, a3, a4, and a5 in formula (8) can be loaded into the first register, where the first register can be composed of 8 SVE vector registers (each row represents an SVE vector register) to form an 8*8 coefficient matrix N, and the 0 and 1 powers of X1 (i.e., 1 and X1) are loaded into the second register, where the second register is composed of 8 SVE vector registers (each column represents an SVE vector register) to form an 8*8 variable matrix M, and then a matrix multiplication operation is performed based on the first register and the second register, i.e., M·N is performed, and the result is stored in the matrix T. Wherein, the matrix T is a matrix register, which can be an SME matrix register, and the element T[1,1] of the first row and the first column in the matrix T represents the calculated value of a2+a3X1, and the element T[1,2] of the first row and the second column in the matrix T represents the calculated value of a4+a5X1. Then, the intermediate results T[1, 1] and T[1, 2] are extracted from the matrix T into a vector register, and the vector inner product operation is performed with another vector register loaded with the second power and fourth power of X1 to obtain The calculated value and The above two calculated values are added together and then added with a0 and a1X1 to obtain the function value of the first function when the independent variable x is X1.
[0122] It should be understood that the above-mentioned polynomial splitting method is to appropriately improve the filling rate of the coefficient matrix N, that is, the filling rate of multiple vector registers that constitute the coefficient matrix N. In addition to the above-mentioned splitting method, there may be other methods, which are not specifically limited in the embodiments of the present application.
[0123] For example, when the number of terms in the polynomial is an even number (such as formula 9), the two adjacent terms in the polynomial can be merged; when the number of terms in the polynomial is an odd number, the two adjacent terms in the polynomial except the constant term (i.e., the zero-order term) can be merged, such as the 7-term and 6-order polynomials in formula (9) can be merged into the form of formula (10). Figure 5As shown, the coefficients a1, a2, a3, a4, a3, a6 can be loaded into the coefficient matrix N, and the 0th and 1st powers of X1 can be loaded into the variable matrix M, and then the matrix multiplication operation is performed on M and N, and the result is stored in the matrix T. Among them, the element T[1,1] in the matrix T represents the calculated value of a1+a2X1, the element T[1,2] in the matrix T represents the calculated value of a3+a4X1, and the element [1,3] in the matrix T represents the calculated value of a5+a6X1. Subsequently, the intermediate results T[1,1], T[1,2], and T[1,3] are extracted from the matrix T, and they are multiplied by the first power, the third power, and the fifth power of X1, respectively, to obtain X1(a1+a2X1), The result is obtained by adding the above results and adding a0 to obtain the function value of the first function when the independent variable x takes the value of X1.
[0124] For another example, in addition to merging two or more terms with adjacent orders, more terms can be merged. For example, the 6th-order polynomial in formula (11) is obtained by merging three adjacent terms in formula (9). Then, the corresponding coefficients and multiple powers of the target independent variable value are loaded into M and N for matrix multiplication, and then multiplied with the independent variable power to obtain the final calculation result. We will not introduce it in detail here.
[0125] A(x)=a0+a1x+a2x 2 +a3x 3 +a4x 4 +a5x 5 +a6x 6 (9)
[0126] A(x)=a0+x(a1+a2x)+x 3 (a3+a4x)+x 5 (a5+a6x) (10)
[0127] A(x)=a0+x(a1+a2x+a3x 2 )+x 4 (a4+a5x+a6x 2 ) (11)
[0128] A(x)=a0+(a1x+a2x 2 )+x 2 (a3x+a4x 2 )+x 4 (a5x+a6x 2 ) (12)
[0129] For example, in addition to storing the zeroth power and the first power of the independent variable in the variable matrix M, the first power and the second power of the independent variable can also be stored in the variable matrix by appropriate splitting. For example, formula (9) can be split into the form of formula (12). Figure 6 As shown, the coefficients a1, a2, a3, a4, a3, a6 can be loaded into the coefficient matrix N, and the 1st and 2nd powers of X1 can be loaded into the variable matrix M. Then, the matrix multiplication operation is performed on M and N, and the result is stored in the matrix T. Among them, the element T[1,1] in the matrix T represents The calculated value of, element T[1,2] represents The calculated value of, element T[1,3] represents Then, we extract the intermediate results T[1,2] and T[1,3] from the matrix T and multiply them by the 2nd and 4th powers of X1, respectively, to get The result is added with a0 and T[1,1] to obtain the function value of the first function when the independent variable x takes the value of X1 respectively.
[0130] In a possible implementation, different powers of the first target variable value are loaded into the second register, and different powers of the second target variable value can also be loaded into the second register, wherein the first target variable value and the second target variable value are both within the value range of the independent variable in the first function. Then, a multiplication operation is performed based on the first register and the second register, and the calculation result obtained at this time includes not only the function calculation value of the first function when the independent variable in the first function takes the value of the first target variable, but also the function calculation value of the first function when its independent variable takes the value of the second target variable. In other words, the results of the first function under two different independent variable values can be calculated.
[0131] For example, Figure 7 As shown, assuming that the first function is expanded in polynomial form to obtain equation (10), to calculate the function value of the first function when the independent variable x takes the values of X1 and X2 (both within the value range of x), the coefficients a1, a2...a6 in A(x) can be stored in the first register (composed of multiple vector registers here, each row represents a vector register) to form a coefficient matrix N, and the 0 and 1 power values of X1 are loaded into the second register (composed of multiple vector registers here, each column represents a vector register) to form a vector matrix M. Then, the vector inner product operation is performed based on the above-mentioned first register and second register, and the intermediate result is stored in the matrix T. Among them, the element T[1,1] in the matrix T represents The calculated value of, element T[1,2] represents The calculated value of, element T[1,3] represents The calculated value of, element T[2,1] represents The calculated value of, element T[2,2] represents The calculated value of, element T[2,3] represents Then, we extract the intermediate results T[1,2] and T[1,3] from the matrix T and multiply them by the 2nd and 4th powers of X1, respectively, to get The result of , plus a0 and T[1,1], to get the function value of the first function when the independent variable x is X1. Similarly, extract the intermediate results T[2,2] and T[2,3] from the matrix T, multiply them by the 2nd power and the 4th power of X1 respectively, and get The result of , plus a0 and T[2,1], is obtained to obtain the function value of the first function when the independent variable x is X2.
[0132] It should be noted that Figure 4 to Figure 7 The areas in each matrix (matrix N, M, T) where no specific values are written can be loaded with corresponding data according to actual computing requirements, which does not mean that they cannot be filled.
[0133] In a possible implementation, similar to the first function, the second function is also a function containing transcendental function terms. The coefficients of the polynomial obtained by expanding the second function in polynomial form can also be loaded into the first register, and the third target variable value can be loaded into the second register, wherein the third target variable value is within the value range of the independent variable in the second function. Then, a matrix multiplication operation is performed based on the first register and the second register. The calculation result obtained at this time includes not only the function calculation value of the first function when the independent variable in the first function takes the value of the first target variable, but also the function calculation value of the second function when the independent variable in the second function takes the value of the third target variable. In other words, the operation results of different functions can be calculated.
[0134] For example, Figure 8 As shown, assuming that the first function is expanded in polynomial form as Figure 8 A(x) in the second function is expanded into a polynomial form as Figure 8 B(y) in the third function is expanded in polynomial form as follows: Figure 8 C(z) in the fourth function is expanded in polynomial form as follows: Figure 8D(q) in . If you want to calculate the function value of the first function when the independent variable x takes values of X1 to X8 (all within the value range of x), the function value of the second function when the independent variable y takes values of Y1 to Y8 (within the value range of y), the function value of the third function when the independent variable z takes values of Z1 to Z8 (within the value range of y), and the function value of the fourth function when the independent variable q takes values of Q1 to Q8 (within the value range of q), then the coefficients a2, a3, a4, and a5 in A(x) can be stored in the first register (composed of multiple vector registers, each row represents a vector register), the coefficients b2, b3, b4, and b5 in B(y) can also be stored in the first register, the coefficients c2, c3, c4, and c5 in C(z) can also be stored in the first register, and the coefficients d2, d3, d4, and d5 in D(q) can also be stored in the first register, thereby forming an 8*8 coefficient matrix N. Similarly, the 0 and 1 power values of X1 are loaded into the second register (composed of multiple vector registers, each column represents a vector register), the 0 and 1 power values of Y1 are loaded into the second register, the 0 and 1 power values of Z1 are loaded into the second register, and the 0 and 1 power values of Q1 are loaded into the second register, thereby forming an 8*8 variable matrix M. Then, the matrix multiplication operation is performed based on the above first register and second register, and the calculated value is stored in the matrix register T. Subsequently, the intermediate results T[1,1] and T[1,2] are extracted from the matrix T, which are multiplied by the 2nd power and the 4th power of X1 respectively, to obtain The result of , plus a0 and a1X1, to get the function value of the first function when the independent variable x is X1. Extract the intermediate results T[2,1] and T[2,2] from the matrix T, multiply them by the 2nd power and the 4th power of X2 respectively, and get The result of , plus a0 and a1X2, to get the function value of the first function when the independent variable x is X2. The calculation method of the function value of the first function when the independent variable is X3 to X8 is similar, both of which extract the corresponding intermediate result from the matrix T and multiply it by the power of the independent variable to get the function calculation value.
[0135] Similarly, extract the intermediate results T[1,3] and T[1,4] from the matrix T and multiply them by the 2nd and 4th powers of Y1 respectively to get The result of , plus b0 and b1Y1, to get the function value of the second function when the independent variable y is Y1. Extract the intermediate results T[2,3] and T[2,4] from the matrix T, multiply them by the 2nd power and the 4th power of Y2 respectively, and get The result of , plus b0 and b1Y2, to get the function value of the second function when the independent variable y is Y2. The calculation method of the function value of the second function when the independent variable is Y3 to Y8 is similar, both of which extract the corresponding intermediate result from the matrix T and multiply it by the power of the independent variable to get the function calculation value.
[0136] It can be seen that the above method can calculate multiple groups of polynomials (here there are four groups of polynomials) at the same time, realizing batch processing of multiple polynomials.
[0137] It should be understood that the fill rate of the second register (for loading coefficients) is often lower when using matrix multiplication than when using vector inner product operation, but the number of polynomials that can be calculated simultaneously (at one time) is usually greater when using matrix multiplication operation. For example, Fig. 9 The performance difference between the two hardware units, SVE vector register and SME matrix register, when calculating a sixth-order polynomial is given as an example. The horizontal axis represents the number of polynomials and the vertical axis represents the performance gain. Fig. 9 It can be seen that when the number of polynomials is relatively small (such as 32 or 64), the performance benefit of the SVE method is higher (indicating short calculation time and fast calculation speed), while when the number of polynomials is relatively large (such as 64), the performance benefit of the SME method is higher.
[0138] Therefore, in some possible implementations, when the number of polynomials to be calculated is less than or equal to the first threshold, the first register and the second register perform a vector inner product operation, and when the number of polynomials to be calculated is greater than the first threshold, the first register and the second register perform a matrix multiplication operation. The first threshold can be set and adjusted according to the actual scenario, and the embodiment of the present application does not specifically limit it.
[0139] In some other possible implementation schemes, when the number of coefficients contained in the target group coefficients is less than or equal to the second threshold, a matrix multiplication operation is performed between the first register and the second register; when the number of coefficients contained in the target group coefficients is greater than the second threshold, a vector inner product operation is performed between the first register and the second register. Among them, the second threshold can be set and adjusted according to the actual scenario, and the embodiment of the present application does not make specific restrictions. It can be understood that if the number of coefficients contained in the target group coefficients is small, the matrix multiplication operation is selected based on the first register and the second register, and the utilization rate of the first register can be achieved relatively high. If the number of coefficients contained in the target group coefficients is large, the matrix multiplication operation is selected based on the first register and the second register, and the utilization rate of the first register cannot be achieved high. Therefore, when the number of coefficients contained in the target group coefficients is large, it is more appropriate to select the vector inner product operation based on the first register and the second register, and the fill rate of the first register can be achieved high.
[0140] In summary, in the function calculation method provided in the embodiment of the present application, the first function is expanded in the form of a polynomial (i.e., polynomial fitting) to transform it into the form of a polynomial, and then the solution of the first function is converted into the solution of the polynomial, which does not rely on the solution of the mathematical library, thereby simplifying the operation, improving the efficiency of function calculation, and reducing the computing power requirement. As for the solution of the polynomial, it can be implemented based on registers, the coefficients of the polynomial obtained after the conversion of the first function are stored in the first register, and the different powers of the first target variable value (i.e., a certain target variable value) of the independent variable in the first function are stored in the second register, and then the multiplication operation is performed based on the first register and the second register, so that the function calculation value of the first function when the independent variable in the first function takes the first target variable value can be quickly obtained.
[0141] Next, combine Fig.10 ,right Figure 2 The function calculation method is explained with a specific example.
[0142] like Fig.10 As shown, first, the first function can be automatically identified by the introduction or the first function, and then the first function is converted into a polynomial form. Wherein, the first function includes transcendental function terms, and other contents about the first function can be referred to the previous introduction, which will not be repeated here. Regarding the order, number of terms, number of polynomials, etc. of the polynomial converted by the first function, it can be selected according to the actual application scenario, and the embodiment of the present application is not specifically limited. For example, it can be selected according to the calculation accuracy requirement of the first function (ensuring that the error is within an acceptable range), the bit width of the register, etc., and can be specifically referred to the previous related introduction, which will not be repeated here.
[0143] Then, a selection is made between the two types of hardware, SVE and SME, to give full play to the computing power advantages of SVE / SME, accelerate the calculation of polynomials, and provide acceleration opportunities for applications that require function calculation. The selection method is not specifically limited. For example, the selection can be made based on the order / number of terms of the polynomial transformed from the first function. When the order of the polynomial transformed from the first function is less than or equal to the first threshold (settable), SME is selected for use. When the order of the polynomial transformed from the first function is greater than the first threshold, SVE is selected for use. Alternatively, the selection can also be made based on the number of polynomials that need to be calculated. When the number of polynomials that need to be calculated is less than or equal to the second threshold (settable), SVE is selected for use. When the number of polynomials that need to be calculated is greater than the second threshold, SME is used.
[0144] When SVE is selected, SVE vector registers are applied for according to the order of the polynomial, and then the polynomial coefficients are loaded into some SVE vector registers, and the variable values of the first function are loaded into other SVE vector registers (for details, please refer to the relevant introduction in the previous article). Then, vector inner product operations are performed based on these SVE vector registers (the vector multiplication instructions of SVE are called to implement vectorized operations), and the calculation results are saved in the vector registers to complete the calculation.
[0145] When SME is selected, vector registers and SME matrix registers are applied according to the order of the polynomial, and then the polynomial coefficients are loaded into a part of the vector registers to form the coefficient matrix N, and the variable values of the first function are loaded into another part of the vector registers to form the variable matrix M (for details, please refer to the relevant introduction in the previous text), and then the variable matrix M and the coefficient matrix N perform matrix multiplication operations (calling the matrix multiplication instruction of SME to implement, mathematically it is matrix multiplication, but actually it is implemented through the SVE vector outer product), and the results are saved in the SME matrix register to form the matrix T. Finally, the intermediate result is extracted from the matrix T and multiplied by the power of the independent variable to obtain the calculation result and complete the calculation.
[0146] See also Fig.11 , Fig.11 It is a structural diagram of a function computing device 1100 provided in the present application, including a control unit 1101 and a computing unit 1102 .
[0147] The control unit 1101 is used to load coefficients of a polynomial obtained by expanding a first function in a polynomial form into a first register, wherein the first function includes a transcendental function term.
[0148] The control unit 1101 is further configured to load different powers of the first target variable value into the second register, wherein the first target variable value is within the value range of the independent variable in the first function.
[0149] The operation unit 1102 is used to perform a multiplication operation based on the first register and the second register to obtain a calculation result, wherein the calculation result includes a function calculation value of the first function when the independent variable in the first function takes the value of the first target variable.
[0150] Optionally, the first function mentioned above is a potential function.
[0151] Optionally, the control unit 1101 is specifically used to load part or all of the target group coefficients in the multiple groups of coefficients into the first register. The multiple groups of coefficients correspond to multiple values one by one, and the multiple values are all within the value range of the independent variable in the first function. Each group of coefficients in the multiple groups of coefficients includes the coefficients of the polynomial obtained by Taylor expansion of the first function at the values corresponding to each group of coefficients, and the target group coefficients are a group of coefficients corresponding to a value in the multiple values that is closest to the value of the first target variable.
[0152] Optionally, one or more of the number of coefficients in the target group of coefficients loaded into the first register, the number of coefficients contained in each group of coefficients in the plurality of groups of coefficients (reflecting the order of the polynomial expansion), and the number of groups of the plurality of groups of coefficients (that is, the number of the plurality of numerical values) may be determined according to at least one of the following:
[0153] (1) The computational accuracy requirement of the first function (for example, it can be measured by accuracy level, accuracy value or error level, etc.);
[0154] (2) The bit width of the first register.
[0155] Optionally, the first register includes one or more vector registers, and the number and / or bit width of the one or more vector registers may be determined according to at least one of the following:
[0156] (1) The computational accuracy requirement of the first function (for example, it can be measured by accuracy level, accuracy value or error level, etc.);
[0157] (2) The number of coefficients contained in the target group coefficients;
[0158] (3) The number of terms in the polynomial corresponding to the target group coefficients;
[0159] (4) The highest order of the polynomial corresponding to the target group coefficients.
[0160] Optionally, before the control unit 1101 loads part or all of the target group of coefficients from the multiple groups of coefficients into the first register, the control unit 1101 is further used to identify a first function from the application code, and then generate and record the multiple groups of coefficients according to the first function.
[0161] Optionally, when the number of coefficients included in the target group coefficients is less than or equal to a threshold, the multiplication operation is a vector outer product operation; when the number of coefficients included in the target group coefficients is greater than the threshold, the multiplication operation is a vector inner product operation.
[0162] Optionally, the first register includes a first part of the coefficients in the target group of coefficients, and the operation unit 1102 is specifically used to: perform a vector outer product based on the first register and the second register to obtain an intermediate result; load the intermediate result and the second part of the coefficients in the target group of coefficients into a third vector register, wherein the order corresponding to the first part of the coefficients is higher than the order corresponding to the second part of the coefficients; perform a vector inner product operation based on the third vector register and the fourth vector register to obtain a calculation result, wherein the fourth vector register is loaded with different powers of the first target variable value.
[0163] Optionally, before the operation unit 1102 performs the multiplication operation based on the first register and the second register, the control unit 1101 is also used to load different powers of the second target variable value into the second register, wherein the second target variable value is within the value range of the independent variable in the first function, and the calculation result also includes the function calculation value of the first function when the independent variable takes the value of the second target variable.
[0164] Optionally, before the operation unit 1102 performs the multiplication operation based on the first register and the second register, the control unit 1101 is also used to load the coefficients of the polynomial obtained by Taylor expansion of the second function into the first register, and load the third target variable value into the second register, wherein the third target variable value is within the value range of the independent variable in the second function, and the calculation result also includes the function calculation value of the second function when the independent variable in the second function takes the value of the third target variable.
[0165] It should be noted that the control unit 1101 and the computing unit 1102 may be implemented by software, or may be implemented by hardware, or may be implemented by both software and hardware. Exemplarily, the implementation of the control unit 1101 is described below by taking the control unit 1101 as an example. Similarly, the implementation of the other units / modules may refer to the implementation of the control unit 1101.
[0166] Unit / module As an example of a software functional unit, the control unit 1101 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the control unit 1101 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with close geographical locations. Among them, usually a region can include multiple AZs.
[0167] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0168] As an example of a hardware functional unit, the control unit 1101 may include at least one computing device, such as a server, etc. Alternatively, the control unit 1101 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0169] The multiple computing devices included in the control unit 1101 can be distributed in the same region or in different regions. The multiple computing devices included in the control unit 1101 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the control unit 1101 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0170] It should be noted that, in other embodiments, the control unit 1101 and the computing unit 1102 can both be used to execute Figure 1 The steps that the above modules are responsible for implementing can be specified as needed, and all functions of the function computing device 1100 are realized by implementing different steps in the function computing method respectively.
[0171] It should also be noted that the above-mentioned function calculation device 1100 is used to execute any embodiment of the function calculation method provided in this application. Figure 1 The relevant description is not repeated here. Fig.11 The function computing device 1100 is only used as an example to illustrate the division of the above-mentioned units / functional modules. In actual applications, the above-mentioned functions can be distributed to different units / functional modules as needed, that is, the internal structure of the function computing device 1100 is divided into other different units / functional modules to complete all or part of the functions described above.
[0172] See also Fig.12 The present application also provides a computing device 1200, including a bus 1202, a processor 1204, a memory 1206, and a communication interface 1208. The processor 1204, the memory 1206, and the communication interface 1208 communicate with each other through the bus 1202. The computing device 1200 may be a server, a terminal device, etc., and the present application embodiment does not make a specific limitation, and the present application embodiment does not limit the number of processors and memories in the computing device 1200.
[0173] The bus 1202 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.12The bus 1202 may include a path for transmitting information between various components of the computing device 1200 (eg, the memory 1206, the processor 1204, and the communication interface 1208).
[0174] The processor 1204 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0175] The memory 1206 may include a volatile memory, such as a random access memory (RAM). The processor 1204 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0176] The memory 1206 stores executable program codes. The processor 1204 executes the executable program codes to respectively implement Fig.11 The functions of the control unit 1101 and the operation unit 1102 are implemented to realize any embodiment of the function calculation method provided in this application.
[0177] The communication interface 1208 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1200 and other devices or communication networks.
[0178] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0179] like Fig.13 As shown, the computing device cluster includes at least one computing device 1200. The memory 1206 in one or more computing devices 1200 in the computing device cluster may store the same instructions for executing the function calculation method of the previous embodiment.
[0180] In some possible implementations, the memory 1206 of one or more computing devices 1200 in the computing device cluster may also store partial instructions for executing the function calculation method of the previous embodiment. In other words, the combination of one or more computing devices 1200 can jointly execute the instructions for the function calculation method of the previous embodiment.
[0181] It should be noted that the memory 1206 in different computing devices 1200 in the computing device cluster may store different instructions, respectively used to execute Fig.11 That is, the instructions stored in the memory 1206 in the different computing devices 1200 can implement Fig.11 The functions of one or more units in the control unit 1101 and the computing unit 1102.
[0182] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.14 A possible implementation is shown. Fig.14 As shown, two computing devices 1200A and 1200B are connected via a network. Specifically, they are connected to the network via a communication interface in each computing device. In this type of possible implementation, the memory 1206 in the computing device 1200A stores instructions for executing the functions of the control unit 1101. At the same time, the memory 1206 in the computing device 1200B stores instructions for executing the functions of the operation unit 1102.
[0183] It should be understood that Fig.14 The functions of the computing device 1200A shown in FIG. 1 may also be jointly completed by multiple computing devices 1200. Similarly, the functions of the computing device 1200B may also be jointly completed by multiple computing devices 1200.
[0184] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Fig.14 The connection mode of the computing device cluster is different in that the memory 1206 in one or more computing devices 1200 in the computing device cluster may store the same instructions for executing the function calculation method of the previous embodiment.
[0185] In some possible implementations, the memory 1206 of one or more computing devices 1200 in the computing device cluster may also store partial instructions for executing the function calculation method that can implement the previous embodiment. In other words, the combination of one or more computing devices 1200 can jointly execute instructions for executing the function calculation method that can implement the previous embodiment.
[0186] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct a computing device cluster (including at least one computing device) to execute any embodiment of the function calculation method provided in the present application.
[0187] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes any embodiment of the function calculation method provided in the present application.
[0188] The function calculation method provided in the present application can be used in chips with different computing units, such as chips with SVE or SME. Fig.15 As shown, the embodiment of the present application also provides a chip, including an operator, a controller and multiple registers. The operator, the controller and the multiple registers are interconnected. The multiple registers may include vector registers (such as SVE vector registers) and may also include matrix registers (such as SME registers). The chip can be used to execute Figure 1 The function calculation method of the embodiment, for specific methods, please refer to the above introduction, which will not be repeated here.
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A function calculation method, characterized in that: The method comprises: Loading coefficients of a polynomial obtained by Taylor expansion of a first function into a first register, wherein the first function includes a transcendental function term; Loading different powers of a first target variable value into a second register, wherein the first target variable value is within a value range of an independent variable in the first function; A calculation result is obtained by performing a multiplication operation based on the first register and the second register, wherein the calculation result includes a function calculation value of the first function when the independent variable in the first function takes the value of the first target variable.
2. The method according to claim 1, characterized in that: The step of loading coefficients of a polynomial obtained by performing Taylor expansion on the first function into a first register includes: Load part or all of the target group coefficients in the multiple groups of coefficients into the first register, wherein the multiple groups of coefficients correspond one-to-one to multiple numerical values, and the multiple numerical values are all within the value range of the independent variable in the first function, each group of coefficients in the multiple groups of coefficients includes the coefficients of the polynomial obtained by Taylor expansion of the first function at the numerical values corresponding to each group of coefficients, and the target group coefficients are a group of coefficients corresponding to a numerical value among the multiple numerical values that is closest to the first target variable value.
3. The method according to claim 2, characterized in that One or more of the number of coefficients in the target group of coefficients that are loaded into the first register, the number of coefficients contained in each group of the plurality of groups of coefficients, and the number of groups of the plurality of groups of coefficients is determined according to at least one of the following: The calculation accuracy requirement of the first function; Alternatively, the bit width of the first register.
4. The method according to claim 2, characterized in that: The first register includes one or more vector registers, and the number and / or bit width of the one or more vector registers are determined according to at least one of the following: The calculation accuracy requirement of the first function; Or, the number of coefficients contained in the target group of coefficients; Alternatively, the number of terms of the polynomial corresponding to the target group coefficients; Alternatively, the highest order of the polynomial corresponding to the target group of coefficients.
5. The method according to any one of claims 2 to 4, characterized in that Before loading part or all of the target group coefficients of the plurality of groups of coefficients into the first register, the method further comprises: identifying the first function from application code; The multiple sets of coefficients are generated and recorded according to the first function.
6. The method according to any one of claims 2 to 5, characterized in that When the number of coefficients included in the target group of coefficients is less than or equal to a threshold value, the multiplication operation is a vector outer product operation; When the number of coefficients included in the target group of coefficients is greater than the threshold, the multiplication operation is a vector inner product operation.
7. The method according to any one of claims 2 to 5, characterized in that The first register includes a first part of coefficients in the target group of coefficients, and performing a multiplication operation based on the first register and the second register to obtain a calculation result includes: Perform a vector outer product based on the first register and the second register to obtain an intermediate result; Loading the intermediate result and a second part of coefficients in the target group of coefficients into a third vector register, wherein the order corresponding to the first part of coefficients is higher than the order corresponding to the second part of coefficients; A vector inner product operation is performed based on the third vector register and the fourth vector register to obtain the calculation result, wherein the fourth vector register is loaded with different powers of the first target variable value.
8. The method according to any one of claims 1 to 7, characterized in that Before performing the multiplication operation based on the first register and the second register, the method further includes: Different powers of a second target variable value are loaded into the second register, wherein the second target variable value is within the value range of the independent variable in the first function, and the calculation result also includes the function calculation value of the first function when the independent variable takes the value of the second target variable.
9. The method according to any one of claims 1 to 8, characterized in that Before performing the multiplication operation based on the first register and the second register, the method further includes: The coefficients of the polynomial obtained by Taylor expansion of the second function are loaded into the first register, and the third target variable value is loaded into the second register, wherein the third target variable value is within the value range of the independent variable in the second function, and the calculation result also includes the function calculation value of the second function when the independent variable in the second function takes the value of the third target variable.
10. The method according to any one of claims 1 to 9, characterized in that The first function is a potential function.
11. A function computing device, characterized in that: include: A control unit, configured to load coefficients of a polynomial obtained by Taylor expansion of a first function into a first register, wherein the first function includes a transcendental function term; The control module is further configured to load different powers of the first target variable value into the second register, wherein the first target variable value is within the value range of the independent variable in the first function; An operation module is used to perform a multiplication operation based on the first register and the second register to obtain a calculation result, wherein the calculation result includes a function calculation value of the first function when the independent variable in the first function takes the value of the first target variable.
12. A chip, characterized in that: The chip comprises an operator, a controller and a plurality of registers, and is used to execute the method as claimed in any one of claims 1 to 10.
13. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 10.
14. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 10.
15. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 10.