Vectorization optimization method of cry detection data set based on MXU3 instruction set
By using MXU3 instruction set and SIMD instructions for vectorization optimization in cry sound detection, the problems of slow data loading speed and low computational efficiency are solved, and more efficient calculations and lower accuracy errors are achieved.
Patent Information
- Application Number
- CN202311576185.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the cry detection data is loaded slowly and the calculation efficiency is not high. In particular, the calculation time is too long to calculate the square root of the vector and the accuracy error is large.
The vectorization optimization method based on the MXU3 instruction set is adopted, and data parallel loading and approximation calculation is performed through the SIMD instruction, and the table lookup method is used to correct the interval with large errors to improve the calculation efficiency and accuracy.
It significantly improves the data loading speed and calculation efficiency, and reduces the accuracy error. Especially when calculating the square root of the vector, the theoretical calculation speed is increased by 16 times, with an error less than 0.001.
Smart Images

Figure CN120032669A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of audio processing, and in particular relates to a vectorization optimization method for a crying detection data set based on an MXU3 instruction set. Background Art
[0002] In the existing technology, with the development of science and technology, the leap of artificial intelligence technology such as speech recognition in audio processing, the demand for electronic baby care is getting higher and higher. At the same time, due to the high accuracy requirements of baby care, deep neural network technology is increasingly being used in the field of baby care. Among them, deep neural network is one of the commonly used technologies of artificial intelligence, which is characterized by fitting pre-acquired data through complex and massive parameters, and these parameters must be retained for subsequent prediction of new data.
[0003] Existing smart baby care products include sound detection, image detection and other methods. However, as for crying detection, the data is relatively small and difficult to collect. In addition, the noise in the actual environment will also interfere with the accuracy of crying detection.
[0004] In addition, the prior art still has the following defects:
[0005] 1. Data loading speed is too slow: Computers need to transfer data to the processor before they can process data. Loading data from the memory or external storage takes too long.
[0006] 2. Low data calculation efficiency: When the processor processes a large amount of matrix data, it needs to calculate each number one by one. These numbers have to perform the same operation, but the values of the numbers are different. This calculation method is very time-consuming.
[0007] 3. Calculating the square root of a vector requires a very high amount of calculation. Currently, there are only two solutions. One is to take out the elements in the vector one by one and calculate them. The other is to use the "Newton iteration method" to approximate the calculation. However, the first solution takes too long and is not supported by most SIMD instructions. The second solution has a large precision error. SIMD (single instruction multiple data) refers to single instruction stream multiple data streams, a digital parallel computing method. One operation instruction can execute multiple data streams, which can improve the computing speed of the program.
[0008] In addition, the commonly used terms in the prior art include:
[0009] 1. Floating point numbers ("float" or "FP32" in English) are a method of representing decimals inside computers under the IEEE binary floating point arithmetic standard. A floating point number has a width of 32 bits.
[0010] 2. Feature: When designing algorithms for classification or prediction of data such as pictures and sounds, it can be used to describe data for specific purposes after complex calculations.
[0011] 3. Register: The core component of the computer system, with very high reading and writing speeds. 40 times faster than the memory reading speed.
[0012] 4. Vectorized computing: a method of calculating multiple data that need to perform the same operation at the same time.
[0013] 5. Newton's iteration method: A method proposed by Newton in the 17th century to approximate solutions to equations in the real and complex number fields. Summary of the invention
[0014] In order to solve the above problems, the purpose of this application is to propose to use SIMD instructions to accelerate crying feature data based on MXU3, reduce data reading and calculation time, optimize calculation logic, and improve calculation efficiency. MXU3 is an instruction set designed by Beijing Ingenic Integrated Circuit Co., Ltd. (abbreviated as Ingenic) for application requirements with parallel computing characteristics such as audio and video, graphics, and image signal processing, and supports SIMD instructions to accelerate data operations.
[0015] Specifically, the present invention provides a vectorization optimization method for a crying detection data set based on an MXU3 instruction set, the method comprising the following steps:
[0016] S1, load, use the vectorized read instruction LA(o,VR0,0,addr,0) of the MXU3 instruction set to read the data into the register. Since each register has a 512-bit width, if each data has an m-bit width, when the data is of float type, 512 / m numbers can be loaded at one time. The vectorized read instruction LA(o,VR0,0,addr,0) means reading continuous 32 bytes of data from address addr into the VR0 register;
[0017] S2, calculation, using the calculation instructions of the MXU3 instruction set, can perform vectorized calculation on 512 / m input numbers at one time, and the calculation instructions include but are not limited to floating-point multiplication instructions FMUL and floating-point addition instructions FADDW;
[0018] S3, adjusting the error, using SIMD instructions to perform approximate calculation, wherein the approximate calculation is to calculate the square root using the SIMD instruction approximate method;
[0019] S4, further correction, using the table lookup method to improve the square root calculation. Since there is no instruction to directly calculate the square root in the MXU3 instruction set, this method calculates the square root by combining fitting and table lookup. The interval with large error is corrected by the table lookup method to reduce the error value of the interval with the largest error.
[0020] The method includes FP32 data, each FP32 data has a bit width of 32 bits, that is, m=32, 512 / 32=16.
[0021] In step S2, the improved calculation logic includes:
[0022] Input: the transpose of the crying feature matrix Hermitia t_Hermitia, two rows and N columns, where Hermitia is N rows and two columns; weight matrix filters, 64 rows and 1 column; forward step length bands, which can be set to 1; the accurate square root result of the interval with large error obtained by pre-calculation statistics sqrt_store (need to be calculated and stored in advance), 16 numbers;
[0023] Assume that the maximum error range is 224 to 240;
[0024] Output: The result features of the intermediate feature conversion are stored in a continuous address space;
[0025] start:
[0026] The step length is 16, and 16 numbers are taken for calculation each time, which can be expressed as forn=0,16,32,...,N do; if n==K is expressed as if(n==K), then step S4 is performed, where n represents the number of the group of data, and K is the number of the data with the largest error.
[0027] In the step S4, assuming that after calculation, it is found that n is in the range of 224 to 240, the precision error generated is relatively large, and the judgment statement if (n == 224) can be looked up in the table to obtain the value, and the interval with large precision error from 224 to 240 can be corrected by using the table lookup method. Here, it means that the data error is the largest when K=224 is counted, and the stored result is directly returned here without fitting; including:
[0028] Read sqrt_store to register 13, which is represented by
[0029] reg13←load(reg13,sqrt_store);
[0030] The data of storage register 13 and result 4 are expressed as
[0031] store(register13, result4),
[0032] Among them, result4 is used to store the memory address of the calculation result. The data in the register is volatile, but the calculation speed is fast, so it is read from the memory, calculated in the register, and put back to the memory after the calculation is completed.
[0033] In step S3, using SIMD instructions to approximately calculate the quadratic root includes:
[0034] Read Hermitia[n][0]~Hermitia
[15] [0], a total of 16 inputs, expressed as
[0035] load(register1,t_Hermitia[n]);
[0036] Read Hermitia[n][1] to Hermitia
[15] [1], a total of 16 inputs, expressed as
[0037] load(register2,t_Hermitia[n+N]);
[0038] The 16 inputs of register 1 are squared simultaneously, expressed as
[0039] Register 1 ← pow(register 1);
[0040] The 16 inputs of register 2 are squared simultaneously, expressed as
[0041] Register 2 ← pow(register 2);
[0042] The 16 numbers in register 1 plus the 16 numbers in register 2 are stored in register 3, expressed as register 3←sum(register 1, register 2);
[0043] Read a constant into a register:
[0044] load(register 4, 0.5);
[0045] load(register 5, 1.5);
[0046] load(register 6, 0x5F3759DF);
[0047] The value of register 3 is multiplied by 0.5, expressed as
[0048] Register 7 ← mul(register 3, 0.5);
[0049] Each element of register 3 is logically shifted right by 1 bit, expressed as
[0050] Register 8 ← srliw (register 3, 1);
[0051] The value of register 8 is subtracted from the element of register 5, expressed as
[0052] Register 9 ← sub(register 5, register 8);
[0053] The square of the element in register 9 is expressed as
[0054] Register 10 ← mul(register 9, register 9);
[0055] The element of register 7 minus the element of register 10 is expressed as
[0056] Register 11 ← sub(register 7, register 10);
[0057] The element of register 9 is subtracted from the element of register 11, expressed as
[0058] Register 12 ← mul(register 11, register 9);
[0059] The element of register 3 is multiplied by the element of register 12, expressed as
[0060] Register 13 ← mul(register 3, register 12);
[0061] The 16 square root results stored in register 13 are stored in result4, expressed as;
[0062] store(register13, result4).
[0063] The method further comprises,
[0064] Read weights from the specified memory area, represented as
[0065] filters = filters + n*bands;
[0066] Perform the corresponding calculation operation
[0067] Update the result of the calculation and put the result back.
[0068] All calculations of the method are based on SIMD instructions. Data reading and calculation are no longer performed from the memory, but are all completed in registers. The square root obtained by the approximate method has an accuracy error of less than 0.001 and can be further corrected. If no correction is needed, K can be set to -1.
[0069] In step S1, when using a neural network to detect the collected crying data, it is necessary to load a corresponding floating point model file from the internal memory or external memory.
[0070] In the step S2, the calculation is to perform matrix calculation.
[0071] Therefore, the advantages of this application are:
[0072] 1. Improve loading speed. Take FP32 data as an example. When using the non-vectorized loading method, only one floating point number can be loaded at a time. Use the vectorized read instruction in the MXU3 instruction set to read the data into the register. Since each register has a 512-bit width and each FP32 data has a 32-bit width, 16 (512 / 32) numbers can be loaded at a time.
[0073] 2. Speed up the calculation speed. Using the calculation instructions in the MXU3 instruction set, 16 input numbers can be vectorized at one time, and the theoretical calculation speed is increased by 16 times.
[0074] 3. In terms of reducing precision error: The error between the approximate method in this method and the sqrt function of the C language standard library is less than 0.01. After further correction, the error value of the interval with the largest error can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention.
[0076] Figure 1 It is a schematic diagram of the process of the present application method.
[0077] Figure 2 It is an example of pseudo code of original data calculation logic in the prior art.
[0078] Figure 3 This is a pseudo-code example of the processing step logic optimized by this application. DETAILED DESCRIPTION
[0079] In order to more clearly understand the technical content and advantages of the present invention, the present invention is now further described in detail in conjunction with the accompanying drawings.
[0080] like Figure 1As shown, the present application proposes a vectorized optimization implementation method for crying detection based on the MXU3 instruction set. When using a neural network algorithm to detect the collected crying data, it is necessary to load the corresponding model file (usually a floating point number) from the memory or external memory, and then carry out subsequent matrix calculations. However, the process of the computer loading a floating point number from the memory and then calculating it takes up the most time for the model calculation. Originally, only one number could be loaded and calculated at a time. After the SIMD method, 16 floating point numbers can be loaded and calculated at a time, which can greatly reduce the theoretical calculation time. However, the use of SIMD instructions also has defects, that is, it only supports common operations such as addition, subtraction, and multiplication, and does not support square roots. Therefore, this method proposes an approximate method for calculating square roots, and can manually correct the intervals with large errors according to actual needs. The optimization method includes:
[0081] S1, load, use the vectorized read instruction of the MXU3 instruction set to read the data into the register. Taking FP32 data as an example, since each register has a 512-bit width and each FP32 data has a 32-bit width, 16 (512 / 32) numbers can be loaded at a time;
[0082] S2, calculation, uses the calculation instructions of the MXU3 instruction set to perform vectorized calculations on 16 input numbers at one time;
[0083] S3, adjusting the error, using SIMD instructions to perform approximate calculation, wherein the approximate calculation is to calculate the square root using the SIMD instruction approximate method;
[0084] S4, further correction, use the table lookup method to improve the square root calculation, the interval with larger error is corrected by the table lookup method to reduce the error value of the interval with the largest error.
[0085] For example, prepare N groups (each group contains 16 data) of random input data in advance. When using this method for the first time, first use the SIMD instruction of this step S3 to approximate the calculation of the quadratic root to fit the square root. If the fitting result is not satisfactory (there is an error compared to the actual square root calculation), record the group number of the N groups of data as K, and store the actual error-free result (which can be calculated by a calculator or C language standard library) in the array sqrt_store.
[0086] The core principle of this method is described in detail below:
[0087] The vectorized implementation method of crying detection based on MXU3 uses SIMD instructions to load and approximate data in parallel. Here, it is parallel loading and parallel computing: it means that the data read and involved in the operation at one time is multiple numbers, and the use of SIMD instructions is parallel computing itself. For example, to calculate 0.1+0.2, 0.2+0.3, you need to read 0.1 first and then read 0.2, and then add them. After 0.1+0.2 is calculated, 0.2+0.3 is calculated. After using SIMD, directly read the left side of the plus sign (0.1, 0.2) and the right side of the plus sign (0.2, 0.3), and then add the corresponding positions at the same time: (0.1+0.2, 0.2+0.3), and finally calculate the two results (0.3, 0.5) at the same time.
[0088] In the prior art, the pseudo code example of the original data calculation logic is as follows: Figure 2 shown.
[0089] Input: crying feature matrix Hermitia, N rows and 2 columns; weight matrix filters, 64 rows and 1 column; forward step length bands, can be set to 1;
[0090] Output: The result features of the intermediate feature conversion are stored in a continuous address space;
[0091] start:
[0092] The step size is 1, and 1 number is taken for calculation each time, which is expressed as
[0093] for n=0,1,2,...,N do;
[0094] Read data from memory and square it, expressed as
[0095] result1=Hermitia[n][0]*Hermitai[n][0];
[0096] result2=Hermitia[n][1]*Hermitai[n][1];
[0097] The sum of result1 and result2 is stored in result3, which is expressed as
[0098] result3=result1+result2;
[0099] Calculate the square root of the sum number by number, expressed as
[0100] result4 = sqrt(result3);
[0101] Calculate the weight address corresponding to the feature, expressed as
[0102] filters = filters + n*bands;
[0103] forb=0,1,2,3,...,64do
[0104] Read the weight value from the weight address and multiply it by result4, expressed as
[0105] result5=result4*filters[b];
[0106] Update the calculated result and put the result back, expressed as
[0107] feature[b]=feature[b]+result5.
[0108] The present invention focuses on optimizing data loading and calculation logic, and can read and calculate 16 numbers each time. The square root operation that cannot be completed by SIMD instructions can be obtained by approximation. For intervals with large errors, they can be corrected by table lookup. For example, when it is found after calculation that n is in the range of 224 to 240, the error generated by the above method is large, and the judgment statement if (n == 224) can be used to look up the table value.
[0109] The pseudo code example of the optimized processing steps is as follows: Figure 3 As shown, the implementation method of the vectorized optimization of this method further includes:
[0110] Input: the transpose of the crying feature matrix Hermitia t_Hermitia, two rows and N columns, where Hermitia is N rows and two columns; weight matrix filters, 64 rows and 1 column; forward step length bands, which can be set to 1; the accurate square root result of the interval with large error obtained by statistics sqrt_store (need to be calculated and stored in advance), 16 numbers; assume that the maximum error interval is 224~240;
[0111] Output: The result features of the intermediate feature conversion are stored in a continuous address space;
[0112] start:
[0113] The step size is 16, and 16 numbers are taken for calculation each time, which can be expressed as forn=0,16,32,...,N do;
[0114] If n==K is determined, it is represented as if(n==K), and then the process goes to step S4.
[0115] In the step S4, assuming that after calculation, it is found that n is in the range of 224 to 240, the precision error generated is relatively large, and the judgment statement if (n == 224) can be looked up in the table to obtain the value, and the interval with large precision error from 224 to 240 can be corrected by using the table lookup method. Here, it means that the data error is the largest when K=224 is counted, and the stored result is directly returned here without fitting; including:
[0116] Read sqrt_store to register 13, which is represented by
[0117] reg13←load(reg13,sqrt_store);
[0118] The data of storage register 13 and result 4 are expressed as
[0119] store(register13, result4),
[0120] Among them, sqrt_store and result4 are arrays used to store data. Result4 is used to store the memory address of the calculation result. The data is easy to be lost when it is placed in the register, but the calculation speed is fast, so it is read from the memory, calculated in the register, and put back to the memory after the calculation is completed.
[0121] In step S3, using SIMD instructions to approximately calculate the quadratic root includes:
[0122] Read Hermitia[n][0]~Hermitia
[15] [0], a total of 16 inputs, expressed as
[0123] load(register1,t_Hermitia[n]);
[0124] Read Hermitia[n][1] to Hermitia
[15] [1], a total of 16 inputs, expressed as
[0125] load(register2,t_Hermitia[n+N]);
[0126] The 16 inputs of register 1 are squared simultaneously, expressed as
[0127] Register 1 ← pow(register 1);
[0128] The 16 inputs of register 2 are squared simultaneously, expressed as
[0129] Register 2 ← pow(register 2);
[0130] The 16 numbers in register 1 plus the 16 numbers in register 2 are stored in register 3, expressed as register 3←sum(register 1, register 2);
[0131] Read a constant into a register:
[0132] load(register 4, 0.5);
[0133] load(register 5, 1.5);
[0134] load(register 6, 0x5F3759DF);
[0135] The value of register 3 is multiplied by 0.5, expressed as
[0136] Register 7 ← mul(register 3, 0.5);
[0137] Each element of register 3 is logically shifted right by 1 bit, expressed as
[0138] Register 8 ← srliw (register 3, 1);
[0139] The value of register 8 is subtracted from the element of register 5, expressed as
[0140] Register 9 ← sub(register 5, register 8);
[0141] The square of the element in register 9 is expressed as
[0142] Register 10 ← mul(register 9, register 9);
[0143] The element of register 7 minus the element of register 10 is expressed as
[0144] Register 11 ← sub(register 7, register 10);
[0145] The element of register 9 is subtracted from the element of register 11, expressed as
[0146] Register 12 ← mul(register 11, register 9);
[0147] The element of register 3 is multiplied by the element of register 12, expressed as
[0148] Register 13 ← mul(register 3, register 12);
[0149] The 16 square root results stored in register 13 are stored in result4, expressed as;
[0150] store(register13, result4).
[0151] The method further comprises,
[0152] Reading weights from the specified memory area means updating weights (addresses). Different weight addresses store different weights, expressed as
[0153] filters = filters + n*bands;
[0154] ...perform corresponding calculation operations;
[0155] ...update the result of the calculation and put the result back.
[0156] All calculations after optimization of this method are based on SIMD instructions. Data reading and calculation are no longer carried out from memory, but are all completed in registers. The square root obtained by the approximate method has an accuracy error of less than 0.001 and can be further corrected. If no correction is required, K can be set to -1. The optimized method can use SIMD instructions to accelerate the crying detection method.
[0157] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A vectorized optimization method for crying detection dataset based on MXU3 instruction set, It is characterized in that The method comprises the following steps: S1, load, use the vectorized read instruction LA(o,VR0,0,addr,0) of the MXU3 instruction set to read the data into the register. Since each register has a 512-bit width, if each data has an m-bit width, when the data is of float type, 512 / m numbers can be loaded at one time. The vectorized read instruction LA(o,VR0,0,addr,0) means reading continuous 32 bytes of data from address addr into the VR0 register; S2, calculation, using the calculation instructions of the MXU3 instruction set, can perform vectorized calculation on 512 / m input numbers at one time, the calculation instructions include floating point multiplication instructions FMUL and floating point addition instructions FADDW; S3, adjusting the error, using SIMD instructions to perform approximate calculation, wherein the approximate calculation is to calculate the square root using the SIMD instruction approximate method; S4, further correction, using the table lookup method to improve the square root calculation. Since there is no instruction to directly calculate the square root in the MXU3 instruction set, this method calculates the square root by combining fitting and table lookup. The interval with large error is corrected by the table lookup method to reduce the error value of the interval with the largest error.
2. According to claim 1, a vectorized optimization method for a crying detection data set based on the MXU3 instruction set, It is characterized in that The method includes FP32 data, each FP32 data has a bit width of 32 bits, that is, m=32, 512 / 32=16.
3. According to claim 2, a vectorized optimization method for a crying detection data set based on the MXU3 instruction set, It is characterized in that In the step S2, the improved calculation logic includes: input: the transpose t_Hermitia of the crying feature matrix Hermitia, two rows and N columns, where Hermitia is N rows and two columns; weight matrix filters, 64 rows and 1 column; forward step length bands, which can be set to 1; statistics to obtain the accurate square root result sqrt_store of the interval with large error in the approximate calculation, the sqrt_store needs to be calculated and stored in advance, 16 numbers; Assume that the maximum error range is 224 to 240; Output: The result features of the intermediate feature conversion are stored in a continuous address space; start: The step size is 16, and 16 numbers are taken for calculation each time, which can be expressed as forn=0,16,32,...,N do; If n==K, it is determined as if(n==K), then step S4 is performed, wherein n represents the number of the data group, and K is the number of the data with the maximum error.
4. According to claim 3, a vectorized optimization method for a crying detection data set based on the MXU3 instruction set, It is characterized in that In step S4, assuming that after calculation, it is found that n is in the range of 224 to 240, the precision error generated is relatively large, and the judgment statement if (n == 224) can be used to look up the table value, and the interval between 224 and 240 with a large precision error can be corrected by using the table lookup method, including: Read sqrt_store to register 13, which is represented by reg13←load(reg13,sqrt_store); The data of register 13 and result 4 are stored, expressed as store (register 13, result4). Among them, result4 is used to store the memory address of the calculation result, which is read from the memory, calculated in the register, and put back into the memory after the calculation is completed.
5. According to claim 4, a vectorized optimization method for a crying detection data set based on the MXU3 instruction set, It is characterized in that In step S3, using SIMD instructions to approximately calculate the quadratic root includes: Read 16 inputs from Hermitia[n][0] to Hermitia[15][0], expressed as load(register 1, t_Hermitia[n]); Read 16 inputs from Hermitia[n][1] to Hermitia[15][1], expressed as load(register 2, t_Hermitia[n+N]); The 16 inputs of register 1 are squared simultaneously, expressed as Register 1 ← pow(register 1); The 16 inputs of register 2 are squared simultaneously, expressed as Register 2 ← pow(register 2); The 16 numbers in register 1 plus the 16 numbers in register 2 are stored in register 3, expressed as Register 3 ← sum(register 1, register 2); Read a constant into a register: load(register 4, 0.5); load(register 5, 1.5); load(register 6, 0x5F3759DF); The value of register 3 is multiplied by 0.5, expressed as Register 7 ← mul(register 3, 0.5); Each element of register 3 is logically shifted right by 1 bit, expressed as Register 8 ← srliw (register 3, 1); The value of register 8 is subtracted from the element of register 5, expressed as Register 9 ← sub(register 5, register 8); The square of the element in register 9 is expressed as Register 10 ← mul(register 9, register 9); The element of register 7 minus the element of register 10 is expressed as Register 11 ← sub(register 7, register 10); The element of register 9 is subtracted from the element of register 11, expressed as Register 12 ← mul(register 11, register 9); The element of register 3 is multiplied by the element of register 12, expressed as Register 13 ← mul(register 3, register 12); The 16 square root results stored in register 13 are stored in result4, expressed as; store(register13, result4).
6. According to claim 5, a vectorized optimization method for a crying detection data set based on the MXU3 instruction set, It is characterized in that The method further includes Reading weights from a specified memory area: expressed as filters = filters + n * bands; Performing corresponding calculation operations; Updating the calculation result and putting the result back.
7. A vectorization optimization method for a cry detection data set based on the MXU3 instruction set according to claim 3, It is characterized in that All calculations of the method are based on SIMD instructions. Data reading and calculation are no longer carried out from memory, but are all completed in registers. And the square root obtained by the approximation method has an accuracy error of less than 0.001 and can be further corrected. If no correction is required, K = -1 can be set.
8. A vectorization optimization method for a cry detection data set based on the MXU3 instruction set according to claim 1, It is characterized in that In step S1, when using a neural network to detect the collected cry data, it is necessary to load the corresponding floating-point model file from memory or external memory.
9. A vectorization optimization method for a cry detection data set based on the MXU3 instruction set according to claim 1, It is characterized in that In step S2, the calculation is matrix calculation.