Multiplication and accumulation operation method and device, and readable storage medium

By performing fixed-point operation and splitting operands on the floating point input values ​​in the accelerator, and assigning weight coefficients, the accelerator's problem is solved that it is difficult for the accelerator to meet the accuracy requirements of multiplication and accumulation operations in deep learning tasks, and the calculation speed and accuracy are improved.

CN119937979APending Publication Date: 2025-05-06EDGELESS SEMICON CO LTD OF ZHUHAI +1

Patent Information

Application Number
CN202411807913.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In deep learning tasks, due to the bit width limitation, the accelerator of embedded devices is difficult to meet the accuracy requirements of multiplication and accumulation operations, resulting in large errors in the calculation results.

Method used

By performing fixed-point operation on the floating-point input value in the accelerator, the target input value in the form of an integer is obtained, the error caused by the floating-point accuracy is reduced, and when the byte length of the target input value exceeds the maximum processing length of the accelerator, it is split into multiple operands, and the weight coefficient is assigned to each operand, and the multiplication and accumulation operation is performed.

Benefits of technology

It reduces the cost of accelerator usage, improves the speed and accuracy of multiplication and accumulation operations, and ensures the accuracy of input values ​​during the calculation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937979A_ABST
    Figure CN119937979A_ABST
Patent Text Reader

Abstract

The invention discloses a multiplication and accumulation operation method and device and a readable storage medium, and the method comprises the steps: carrying out the fixed-point operation of three input values in a floating-point number form in an accelerator, and obtaining the target input values corresponding to the three input values; determining the maximum byte length capable of being operated by the accelerator; for each target input value, splitting the target input value into at least two operands under the condition that the first byte length corresponding to the target input value is greater than the maximum byte length; the second byte length of each operand is smaller than or equal to the maximum byte length; according to the corresponding position information of each operand in the corresponding target input value, determining a weight coefficient corresponding to each operand; performing multiplication and accumulation operation based on operands corresponding to the three target input values and weight coefficients of the operands by using an accelerator to obtain target output values; and the speed and the accuracy of multiplication and accumulation operation of the accelerator are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a multiplication and addition operation method, device, and readable storage medium. Background Art

[0002] With the development of deep learning, deep neural networks have been widely used in many industries and fields, including speech recognition, image recognition, data fitting, etc. In these deep learning tasks, there are many application scenarios of embedded devices. In these application scenarios, embedded devices generally have an accelerator, and the bit width of the accelerator is usually 8bit and 16bit. When the accelerator cannot meet the accuracy requirements of the deep learning task, the calculation results obtained using the accelerator will produce large errors. Summary of the invention

[0003] The embodiments of the present application provide a multiplication-accumulation operation method, device, and readable storage medium, which can reduce the use cost of an accelerator and improve the speed and accuracy of the accelerator's multiplication-accumulation operation.

[0004] In a first aspect, an embodiment of the present application provides a multiplication and accumulation operation method, the method comprising:

[0005] Performing fixed-point operations on three floating-point input values ​​in the accelerator to obtain target input values ​​corresponding to the three input values;

[0006] Determining the maximum byte length that the accelerator can operate;

[0007] For each target input value, if the first byte length corresponding to the target input value is greater than the maximum byte length, split the target input value into at least two operands; the second byte length of each operand is less than or equal to the maximum byte length;

[0008] Determine the weight coefficient corresponding to each operand according to the position information corresponding to each operand in the corresponding target input value;

[0009] The accelerator is used to perform a multiplication and accumulation operation based on operands corresponding to the three target input values ​​and weight coefficients of the operands to obtain a target output value.

[0010] Optionally, performing fixed-point operations on the three floating-point input values ​​in the accelerator to obtain target input values ​​corresponding to the three input values ​​respectively includes:

[0011] Determining a target byte length according to the precision of three input values ​​in the accelerator and the precision ranges corresponding to different byte lengths;

[0012] Determine a first parameter according to the target byte length; the first parameter is the maximum value of a signed integer that can be represented by a binary number of the target byte length;

[0013] For each input value, the product of the input value and the first parameter is calculated to obtain a target input value corresponding to the input value.

[0014] Optionally, determining the weight coefficient corresponding to each operand according to the position information corresponding to each operand in the corresponding target input value includes:

[0015] Determine the weight corresponding to each position in the target input value based on the place value principle of the binary number system;

[0016] The weight coefficient corresponding to each operand is determined according to the position information of each operand in the corresponding target input value and the weight corresponding to each position in the target input value.

[0017] Optionally, the using the accelerator to perform a multiplication and accumulation operation based on operands corresponding to the three target input values ​​and weight coefficients of the operands to obtain a target output value includes:

[0018] Multiplying each operand by a weight coefficient corresponding to each operand to obtain a first product corresponding to the operand;

[0019] A multiplication and addition operation is performed based on the first products corresponding to the respective operands to obtain the target output value.

[0020] Optionally, performing a multiplication and addition operation based on the first products corresponding to the operands to obtain the target output value comprises:

[0021] Multiplying the first products corresponding to the operands in the first target input value and the first products corresponding to the operands in the second target input value to obtain a first number of second products; the first number is the product of the number of the operands in the first target input value and the number of the operands in the second target input value;

[0022] Adding the first number of second multiplication results to obtain a first operation result;

[0023] The first operation result and the first product corresponding to each operand in the third target input value are added to obtain a target output value.

[0024] Optionally, performing a multiplication and addition operation based on the first products corresponding to the operands to obtain the target output value includes:

[0025] Using the accelerator to perform a multiplication-addition operation based on the first products corresponding to the operands to obtain a first output value;

[0026] The first output value is shifted right by a second number of bits to obtain the target output value; the second number is equal to the number of bits corresponding to the second byte length.

[0027] In a second aspect, an embodiment of the present application provides a multiplication and accumulation operation device, the device comprising:

[0028] A fixed-point operation module is used to perform fixed-point operation on three floating-point input values ​​in the accelerator to obtain target input values ​​corresponding to the three input values;

[0029] A first determination module, used to determine the maximum byte length that the accelerator can operate;

[0030] a splitting module, configured to split each target input value into at least two operands when a first byte length corresponding to the target input value is greater than the maximum byte length; and a second byte length of each operand is less than or equal to the maximum byte length;

[0031] A second determination module, configured to determine a weight coefficient corresponding to each operand according to position information corresponding to each operand in the corresponding target input value;

[0032] The multiplication-accumulation operation module is used to use the accelerator to perform multiplication-accumulation operations based on the operands corresponding to the three target input values ​​and the weight coefficients of the operands to obtain the target output value.

[0033] Optionally, the fixed-point operation module includes:

[0034] A first determination submodule, configured to determine a target byte length according to the precision of three input values ​​in the accelerator and the precision ranges corresponding to different byte lengths;

[0035] A second determination submodule is used to determine a first parameter according to the target byte length; the first parameter is a maximum value of a signed integer that can be represented by a binary number of the target byte length;

[0036] The calculation module is used to calculate the product of each input value and the first parameter to obtain a target input value corresponding to the input value.

[0037] Optionally, the second determining module includes:

[0038] A weight determination module, used to determine the weight corresponding to each position in the target input value based on the position value principle of the binary number system;

[0039] The third determination submodule is used to determine the weight coefficient corresponding to each operand according to the position information of each operand in the corresponding target input value and the weight corresponding to each position in the target input value.

[0040] Optionally, the product-addition operation module includes:

[0041] A first operation submodule, configured to perform a multiplication operation on each operand and a weight coefficient corresponding to each operand to obtain a first product corresponding to the operand;

[0042] The second operation submodule is used to perform a multiplication and addition operation based on the first products corresponding to the respective operands to obtain the target output value.

[0043] Optionally, the second operator module includes:

[0044] a multiplication operation module, configured to multiply a first product corresponding to each operand in the first target input value and a first product corresponding to each operand in the second target input value to obtain a first number of second products; the first number being the product of the number of each operand in the first target input value and the number of each operand in the second target input value;

[0045] A first addition operation module, used for adding the first number of second multiplication results to obtain a first operation result;

[0046] The second addition operation module is used to add the first operation result and the first product corresponding to each operand in the third target input value to obtain a target output value.

[0047] Optionally, the second operator module includes:

[0048] A third operation submodule, configured to use the accelerator to perform a multiplication and addition operation based on the first products corresponding to the respective operands to obtain a first output value;

[0049] The right shift module is used to right shift the first output value by a second number of bits to obtain the target output value; the second number is equal to the number of bits corresponding to the second byte length.

[0050] In a third aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the multiplication and addition operation method as described in any one of the above items is implemented.

[0051] The embodiments of the present application include the following advantages:

[0052] The three floating-point input values ​​in the accelerator are fixed-pointed to obtain three integer target input values, thereby reducing the error of the operation result caused by the floating-point precision problem, thereby improving the accuracy of the multiplication and accumulation operation, and the fixed-pointing of the floating-point number to an integer can improve the operation speed; when the first byte length corresponding to the target input value is greater than the maximum byte length that the accelerator can operate, the target input value is split into at least two operands, and a corresponding weight coefficient is assigned to each operand according to the position information of each operand in the target input value. The original multiplication and accumulation operation based on the input value in floating-point form is converted to a multiplication and accumulation operation based on at least two operands and the weight coefficients corresponding to the operands, which ensures the accuracy of the input value in the multiplication and accumulation operation process and improves the accuracy of the accelerator in performing the multiplication and accumulation operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0054] Figure 1 is a flowchart of a multiplication and accumulation method embodiment of the present invention;

[0055] Figure 2 It is a schematic diagram of a flow chart of an accelerator of the present invention performing a multiplication and accumulation operation;

[0056] Figure 3 It is a structural block diagram of a multiplication and accumulation operation device of the present invention;

[0057] Figure 4 It is a structural block diagram of an electronic device provided by an example of the present invention. DETAILED DESCRIPTION

[0058] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0059] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0060] Method Embodiment

[0061] The multiplication and accumulation operation method provided in the embodiment of the present application is described in detail below through specific embodiments and application scenarios in conjunction with the accompanying drawings.

[0062] Reference Figure 1 , which shows a flowchart of a multiplication and accumulation method provided by an embodiment of the present application, such as Figure 1 As shown, the method specifically comprises the following steps:

[0063] Step 101: Perform fixed-point operations on three floating-point input values ​​in the accelerator to obtain target input values ​​corresponding to the three input values.

[0064] The multiplication-accumulation operation method provided by the present application can be applied to the field of digital signal processing. For example, in the process of audio processing, various audio processing algorithms can be efficiently implemented by performing fast Fourier transform and inverse transform on the audio signal in combination with multiplication-accumulation operation. The multiplication-accumulation operation method can also be applied to the fields of machine learning and deep learning. For example, in deep learning, the forward propagation and back propagation processes of neural networks use a lot of multiplication-accumulation operations. In particular, in convolutional neural networks and fully connected layers, multiplication-accumulation operations are used to calculate the dot product between input data and weights, thereby obtaining output feature maps or classification results.

[0065] Multiply Accumulate (MAC) is a special operation that adds the product of the multiplication to the value in the accumulator to obtain the operation result, and then stores the operation result in the accumulator. Multiply-accumulate operation can significantly improve the operation efficiency. Multiply-accumulate operation can complete both multiplication and addition operations in one instruction cycle, while the traditional programming method requires two instruction cycles to complete both multiplication and addition operations. In hardware implementation, multiply-accumulate operation usually relies on a specific multiply-accumulate circuit.

[0066] An accelerator refers to a hardware or software component used to improve computing efficiency. Specifically, an accelerator in deep learning refers to a hardware or software component specifically designed to accelerate the training and reasoning process of deep learning models. An accelerator can include multiple multiplication and accumulation units, and multiple units can operate in parallel to improve the operating efficiency of deep learning models. Common accelerators include: 8-bit accelerators, 16-bit accelerators, 32-bit accelerators, etc.

[0067] It should be noted that the multiplication and accumulation formula is mac(A, B, C)=A×B+C, where A, B, and C represent three floating-point input values ​​respectively, where when C=0, A and B are multiplied; when A=0, B and C are added; when B=0, A and C are added, and when A, B, and C are all non-zero, A and B are multiplied first, and the multiplication result is added to C to obtain the multiplication and accumulation result based on A, B, and C.

[0068] Among them, the value range of the input value in floating-point form is between -1 and 1; the fixed-point operation refers to converting the input value in floating-point form into a target input value in integer form, specifically, according to the mapping relationship between the input value and the target input value, the input value in floating-point form is converted into the target input value in integer form.

[0069] Optionally, before performing fixed-point operations on the three floating-point input values ​​in the accelerator to obtain target input values ​​corresponding to the three input values, the method further includes:

[0070] Normalize the three initial input values ​​and convert them into three floating-point input values.

[0071] Among them, normalization refers to scaling the value to a specified range according to a specific ratio. In the embodiment of the present application, the three initial input values ​​are scaled to the interval [-1, 1] according to a specific ratio. Common normalization methods include: deviation normalization, standard score normalization and other methods. Deviation normalization is a linear normalization method that maps the data value to between (-1, 1) by performing a linear transformation on the initial input value. The specific formula is as follows:

[0072]

[0073] Among them, A is the original data, min(A) is the minimum value in the data set, max(A) is the maximum value in the data set, B is the normalized data, and the value range of B is between (-1, 1).

[0074] Step 102: Determine the maximum byte length that the accelerator can operate.

[0075] Among them, byte length refers to the number of bytes. For example, an 8-bit accelerator refers to a hardware or software component that can support 8-bit integer precision calculations, and the maximum byte length that an 8-bit accelerator can operate is 1 byte; a 16-bit accelerator is a hardware or software component that can support 16-bit integer precision calculations, and the maximum byte length that a 16-bit accelerator can operate is 2 bytes. A 16-bit accelerator can not only operate 2-byte operands, but also 1-byte operands. In other words, an accelerator that supports high-precision operations not only supports operations with operands with the same precision as the accelerator, but also supports operations with operands with lower precision than the accelerator.

[0076] Step 103: For each target input value, if the first byte length corresponding to the target input value is greater than the maximum byte length, split the target input value into at least two operands; the second byte length of each operand is less than or equal to the maximum byte length.

[0077] When the first byte length corresponding to the target input value is less than or equal to the maximum byte length, the accelerator is used to directly perform multiplication and accumulation operations based on the three target input values ​​to obtain the target output value. For example, when the 8-bit accelerator is used to perform multiplication and accumulation operations on three floating-point input values ​​of 0.5, 0.5, and 0, the target byte length required for fixed-point conversion of the three input values ​​is first determined based on the maximum precision of the three input values ​​and the precision ranges corresponding to different byte lengths. The precision of 0.5 is 1 decimal place, and the precision range corresponding to 1 byte length includes 2 decimal places. Therefore, the target byte length for fixed-point operation of the three input values ​​of "0.5, 0.5, 0" is 1 byte. The maximum value of a signed integer that can be represented by a 1-byte binary number is 127, 0.5×127=63.5, and 63.5 is rounded to 64. The binary form of 64 is "1000000". The byte length of the target input value is equal to the maximum byte length that the accelerator can operate. Therefore, there is no need to split the target input value, and the multiplication and accumulation operation is directly performed based on the three target input values, that is, the multiplication operation is performed based on 64 and 64.

[0078] It should be noted that, when the first byte length of the target input value is greater than the maximum byte length, there is at least one splitting strategy for splitting the target input value, and the byte length corresponding to each operand is the same.

[0079] For example, for an 8-bit accelerator, when the target input value is 2 bytes, the target input value needs to be split into two operands, and the second byte length of each operand is 1 byte, that is, the second byte length of the operand is equal to the maximum byte length that the accelerator can operate. For a 16-bit accelerator, it can support operations on 2-byte operands and 1-byte operands. When the target input value is 4 bytes, the target input value can be split into two operands, and the byte length corresponding to each operand is 2 bytes; the target input value can also be split into 4 operands, and the byte length corresponding to each operand is 1 byte.

[0080] In an embodiment of the present application, splitting the target input value into at least two operands means storing the positive and negative values ​​of the target input value in an extended byte, splitting the absolute value of the target input value into at least two operands, the byte length of each operand is less than or equal to the maximum byte length, and the length of the extended byte is equal to the byte length corresponding to the operand.

[0081] Step 104: Determine the weight coefficient corresponding to each operand according to the position information corresponding to each operand in the corresponding target input value.

[0082] Among them, after splitting the target input value into at least two operands, the role of assigning a corresponding weight coefficient to each operand is to add the product results after each operand is multiplied by the weight coefficient, and the addition result is equal to the target input value. The position information can indicate the position of the operand in the target input value. For example, when an 8-bit accelerator is used for calculation, and the byte length corresponding to the target input value is 2 bytes, the high 8 bits of the target input value are split into the first operand, and the low 8 bits are split into the second operand. The weight coefficient corresponding to the first operand is 0x100, and the weight coefficient corresponding to the second operand is 1.

[0083] Step 105 : Using the accelerator, perform a multiplication and accumulation operation based on the operands corresponding to the three target input values ​​and the weight coefficients of the operands to obtain a target output value.

[0084] In an embodiment of the present application, the multiplication and accumulation operation of three floating-point input values ​​using an accelerator is converted into the multiplication and accumulation operation based on operands and weight coefficients of the operands using an accelerator. Among them, the first input value and the second input value are multiplied, and the multiplication result is added to the third input value to obtain the result of the multiplication and accumulation operation of the three floating-point input values. The input values ​​in floating-point form are fixed-pointed to obtain three target input values. When the byte length corresponding to the target input value is greater than the maximum byte length that the accelerator can operate, each target input value is split into at least two operands and weight coefficients corresponding to the operands; and the multiplication and accumulation operation is performed based on the operands corresponding to each input value and the weight coefficients corresponding to the operands.

[0085] In the embodiment of the present application, three floating-point input values ​​in the accelerator are fixed-pointed to obtain three integer target input values, thereby reducing the error of the operation result caused by the floating-point precision problem, thereby improving the accuracy of the multiplication and accumulation operation, and the fixed-pointing of the floating-point number to an integer can improve the operation speed; when the first byte length corresponding to the target input value is greater than the maximum byte length that the accelerator can operate, the target input value is split into at least two operands, and a corresponding weight coefficient is assigned to each operand according to the position information of each operand in the target input value. The original multiplication and accumulation operation based on the input value in floating-point form is converted into a multiplication and accumulation operation based on at least two operands and the weight coefficients corresponding to the operands, which ensures the accuracy of the input value during the multiplication and accumulation operation and improves the accuracy of the multiplication and accumulation operation.

[0086] Optionally, in step 101, fixed-point operations are performed on three floating-point input values ​​in the accelerator to obtain target input values ​​corresponding to the three input values, including:

[0087] Step 11: Determine the target byte length according to the precision of the three input values ​​in the accelerator and the precision ranges corresponding to different byte lengths;

[0088] Step 12: Determine a first parameter according to the target byte length; the first parameter is the maximum value of a signed integer that can be represented by a binary number of the target byte length;

[0089] Step 13: For each input value, calculate the product of the input value and the first parameter to obtain a target input value corresponding to the input value.

[0090] The precision of the input value in the form of floating-point numbers refers to the numerical range and accuracy that the floating-point number can accurately express. A floating-point number consists of two parts: integer digits and decimal digits. The number of decimal digits is positively correlated with the precision of the floating-point number. The more decimal digits there are, the higher the precision of the floating-point number.

[0091] In the embodiment of the present application, the way in which a byte represents a value is to use the highest bit as a sign bit and the remaining bits as value bits. The sign bit is used to indicate the positive or negative value, 0 for a positive number and 1 for a negative number. The precision range corresponding to different byte lengths is determined according to the maximum value of the integer that can be represented by the byte length. A byte includes 8 bits, the highest bit is the sign bit, and the remaining 7 bits are used as value bits. The value range that a byte can represent is (-2 7 , 2 7 -1), that is, the value range that 1 byte can represent is (-128, 127), 1 / 127=0.0078740, it can be considered that the precision range corresponding to 1 byte is less than or equal to 0.01, that is, the maximum precision corresponding to 1 byte is two decimal places; the value range that 2 bytes can represent is (-32768, 32768), 1 / 32767=0.0000305, it can be considered that the precision range corresponding to 2 bytes is less than or equal to 0.0001, that is, the maximum precision corresponding to 2 bytes is 4 decimal places; the value range that 3 bytes can represent is (-8388608, 8388607), 1 / 8388607=0.0000001195, it can be considered that the precision range corresponding to 3 bytes is less than or equal to 0.000001, that is, the maximum precision corresponding to 2 bytes is 6 decimal places.

[0092] Compare the precision of the input value in floating point form with the precision ranges corresponding to different byte lengths, and select the minimum byte length whose precision range is greater than the precision of the input value as the target byte length. For example, the first input value is 0.5, the precision of the first input value is less than 0.01, and the target byte length is 1 byte.

[0093] It should be noted that the precisions corresponding to the three input values ​​in floating-point form can be the same or different. When the precisions corresponding to the three input values ​​in floating-point form are different, the maximum precision among the three input values ​​is selected for comparison with the precision ranges corresponding to different byte lengths, that is, the precision corresponding to the input value with the smallest absolute value among the three floating-point forms is selected for comparison with the precision ranges corresponding to different byte lengths to determine the target byte length. For example, the three input values ​​in floating-point form are 0.5, 0.006, and 0.01, respectively. The precision of 0.006 is selected for comparison with the precision ranges corresponding to different byte lengths. The precision of 0.006 is 3 decimal places. The precision of 0.006 is greater than the precision range corresponding to 1 byte, but less than the precision range corresponding to 2 bytes. Therefore, the target byte length corresponding to the three floating-point input values ​​0.5, 0.006, and 0.01 is 2 bytes.

[0094] After the target byte length is determined, the first parameter is the maximum value of the signed integer that can be represented by the binary number of the target byte length. The target byte length corresponds to N bits, and the maximum value of the signed integer is equal to (2 N-1 -1). For example, the maximum value of a signed integer that can be represented by a 1-byte binary number is 127, the maximum value of a signed integer that can be represented by a 2-byte binary number is 32767, and the maximum value of a signed integer that can be represented by a 3-byte binary number is 8388607.

[0095] As an example, the three floating-point input values ​​are 0.5, 0.006, and 0.01, and the target byte length corresponding to the three floating-point input values ​​is 2 bytes. The maximum value of a signed integer that can be represented by a 2-byte binary number is 32767. The three floating-point input values ​​are multiplied by 32767, and the three values ​​corresponding to the three floating-point input values ​​are 16383.5, 196.602, and 327.27. After rounding these three values, the final target input values ​​are 16384, 197, and 327, respectively.

[0096] In the embodiment of the present application, the target byte length is determined according to the precision of the three input values ​​in the accelerator and the precision range corresponding to different byte lengths, and the first parameter is determined according to the target byte length, wherein the first parameter is the maximum value of a signed integer that can be represented by a binary number of the target byte length; for each input value, the product of the input value and the first parameter is calculated to obtain the target input value corresponding to the input value. By mapping the input value in floating-point form to an integer value through fixed-point operation, the error of the operation result caused by the precision problem of the floating-point number is reduced, and the fixed-point conversion of the floating-point number to an integer can increase the operation speed.

[0097] Optionally, in step 104, determining the weight coefficient corresponding to each operand according to the position information corresponding to each operand in the corresponding target input value includes:

[0098] Step 21: Determine the weight corresponding to each position in the target input value based on the position value principle of the binary number system;

[0099] Step 22: Determine the weight coefficient corresponding to each operand according to the position information of each operand in the corresponding target input value and the weight corresponding to each position in the target input value.

[0100] In the binary number system, the weight of each bit is a power of 2. Specifically, the weight of the nth bit from right to left (i.e., from low to high) is 2. n-1 , n starts from 0. For example, the binary number "1101", the weights of each bit are shown in Table 1:

[0101] Table 1 The weight of each bit of the binary number "1101"

[0102] Sequence Binary Bit Weight 0 1 <![CDATA[2 0 =1]]> 1 0 <![CDATA[2 1 =2]]> 2 1 <![CDATA[2 2 =4]]> 3 1 <![CDATA[2 3 =8]]>

[0103] In an embodiment of the present application, the weight coefficient corresponding to each operand is determined according to the order of the lowest bit value in the operand in the corresponding target input value and the weight corresponding to each position in the target input value. Exemplarily, an 8-bit accelerator is used for multiplication and accumulation operations, the target input value is 11111101000, the hexadecimal number of the target input value is 0x7E8, the byte length of the target input value is greater than the maximum byte length that the accelerator can operate, and the target input value is split into two operands. The high 8 bits 00000111 in the target input value are used as the first operand, and the low 8 bits 11101000 of the target input value are used as the second operand. The lowest bit in the first operand, that is, the rightmost bit 1, has a position order of 8 in the target input value, and the weight corresponding to the position with a position order of 8 is 2 8 =256, the binary form of the weight is 100000000, then the weight coefficient corresponding to the first operand is 100000000; the lowest bit in the second operand is 0 in the target input value, and the weight corresponding to the position with the bit sequence 0 is 2 0 =1, the weight coefficient corresponding to the second operand is 1.

[0104] In the embodiment of the present application, based on the position value principle of the binary number system, the weight corresponding to each position in the target input value is determined; according to the position information of each operand in the corresponding target input value and the weight corresponding to each position in the target input value, the weight coefficient corresponding to each operand is determined. The target input value with a byte length greater than the maximum byte length that the accelerator can operate is split into operands with a byte length less than or equal to the maximum byte length, which retains the accuracy of the numerical value during the operation, thereby improving the accuracy of the multiplication and accumulation operation based on the operands and the weight coefficients corresponding to the operands.

[0105] Optionally, in step 105, using the accelerator to perform a multiplication and accumulation operation based on operands corresponding to the three target input values ​​and weight coefficients of the operands to obtain a target output value includes:

[0106] Step 31, multiplying each operand by a weight coefficient corresponding to each operand to obtain a first product corresponding to the operand;

[0107] Step 32: Perform a multiplication and addition operation based on the first products corresponding to the operands to obtain the target output value.

[0108] As an example, an 8-bit accelerator is used to perform multiplication and accumulation operations on the input values ​​in floating-point form. After the floating-point input values ​​are fixed-pointed, the target input values ​​obtained are A and B, B = -2024, and the hexadecimal form of "-2024" is [0x7E8, 0xFF], where 0x7E8 is 2024 and 0xFF represents a negative value. The byte length of 0x7E8 is greater than the maximum byte length that the accelerator can operate. 0x7E8 is split into two operands 0x7E8. , 0xE8, the weight coefficient corresponding to 0x7 is 0x100, and the weight coefficient corresponding to 0xE8 is 1. 0x7 and 0x100 are multiplied, and 0xE8 and 1 are multiplied to obtain two first products. Similarly, the target input value A is split into two operands, and the operands and the weight coefficients corresponding to the operands are multiplied to obtain two first products. Based on the four first products and the positive and negative signs corresponding to the target input values ​​​​A and B, a multiplication and addition operation is performed to obtain the target output value.

[0109] In an embodiment of the present application, by splitting a target input value whose byte length is greater than the maximum byte length that the accelerator can operate into at least two operands and weight coefficients corresponding to the operands, and using an accelerator to calculate the product between the operands and the weight coefficients corresponding to the operands, the accuracy of the numerical values ​​during the operation process can be retained and the accuracy of the multiplication and addition operations can be improved.

[0110] Optionally, in step 32, performing a multiplication and addition operation based on the first products corresponding to the operands to obtain the target output value comprises:

[0111] Step 41: multiply the first products corresponding to the operands in the first target input value and the first products corresponding to the operands in the second target input value to obtain a first number of second products; the first number is the product of the number of the operands in the first target input value and the number of the operands in the second target input value;

[0112] Step 42, adding the first number of second multiplication results to obtain a first operation result;

[0113] Step 43: Add the first operation result and the first product corresponding to each operand in the third target input value to obtain a target output value.

[0114] As an example, an 8-bit accelerator is used to perform multiplication and accumulation operations. When the three target input values ​​are A, -2024, and 0, A and -2024 are first multiplied, and the result of the multiplication and 0 are added to obtain the final target output value. The hexadecimal form corresponding to -2024 is 0x7E8 and 0xFF, where the byte length of 0x7E8 is greater than the maximum byte length that the 8-bit accelerator can operate. 0x7E8 is split to obtain two operands 0x7 and 0xE8. The weight coefficient corresponding to 0x7 is 0x100, and the weight coefficient corresponding to 0xE8 is 1. 0x7 and 0x100 are multiplied, and 0xE8 and 1 are multiplied; then the multiplication operation of A and -2024 means that the first products corresponding to each operand in the target input value A and the target input value B are multiplied to obtain the second product result, and the second product results are added to obtain the first operation result. The calculation formula is as follows:

[0115] A×-2024=(A×0x7×0x100+A×0xE8)×0xFF (2)

[0116] like Figure 2 As shown, Figure 2 A schematic diagram of a process of performing a multiplication-accumulation operation by an accelerator is shown. In a first instruction cycle, mac(A, 0x7, 0) and mac(A, 0xE8, 0) are executed in parallel; in a second instruction cycle, mac(mac(A, 0x7, 0), 0x100, 0) is executed based on the operation result of the first instruction cycle; in a third instruction cycle, mac(mac(mac(A, 0x7, 0), 0x100, 0), 1, mac(A, 0xE8, 0)) is executed based on the operation results of the first two instruction cycles; in a fourth instruction cycle, mac(mac(mac(A, 0x7, 0), 0x100, 0), 1, mac(A, 0xE8, 0)), 0xFF, 0) is executed based on the operation results of the first three instruction cycles.

[0117] Similarly, when using an 8-bit accelerator to perform multiplication and accumulation operations, when the three target input values ​​are "-2277", "-2024", and "1732", "-2277" and "-2024" are multiplied first, and the result of the multiplication is added to 0 to obtain the final target output value. The hexadecimal form corresponding to "-2277" is 0x8E5 and 0xFF, where the byte length of 0x8E5 is greater than the maximum byte length that the 8-bit accelerator can operate. 0x8E5 is split to obtain two operands 0x8 and 0xE5. The weight coefficient corresponding to 0x8 is 0x100, and the weight coefficient corresponding to 0xE5 is 1. 0x8 and 0x100 are multiplied, and 0xE5 and 1 are multiplied to obtain the first product corresponding to each operand in the first target input value.

[0118] The hexadecimal form corresponding to "-2024" is 0x7E8 and 0xFF, where the byte length of 0x7E8 is greater than the maximum byte length that the accelerator can operate. 0x7E8 is split to obtain two operands 0x7 and 0xE8. The weight coefficient corresponding to 0x7 is 0x100, and the weight coefficient corresponding to 0xE8 is 1. 0x7 and 0x100 are multiplied, and 0xE8 and 1 are multiplied to obtain the first product corresponding to each operand in the second target input value.

[0119] The multiplication operation of "-2277" and "-2024" means that the first products corresponding to "0x8" and "0xE5" in the hexadecimal system of "-2277" and the first products corresponding to "0x7" and "0xE8" in the hexadecimal system of "-2024" are multiplied to obtain four second products, and the second product results are added to obtain the first operation result. The calculation formula is as follows:

[0120] (0x8×0x100×0x7×0x100+0x8×0x100×0xE8+0xE5×0x7×0x100+0xE5×0xE8)×0xFF×0xFF(3)

[0121] The hexadecimal form corresponding to 1732 is 0x6C4. The byte length of 0x6C4 is greater than the maximum byte length that the 8-bit accelerator can operate. 0x6C4 is split to obtain two operands 0x6 and 0xC4. The weight coefficient corresponding to 0x6 is 0x100, and the weight coefficient corresponding to 0xC4 is 1. 0x6 and 0x100 are multiplied, and 0xC4 and 1 are multiplied to obtain the first product corresponding to each operand in the third target input value.

[0122] The first operation result and the first product corresponding to each operand in the third target input value are added to obtain the target output value. The calculation formula is as follows:

[0123] (0x8×0x100×0x7×0x100+0x8×0x100×0xE8+0xE5×0x7×0x100+0xE5×0xE8)×0xFF×0xFF+0x6×0x100+0xC4(4)

[0124] In an embodiment of the present application, the first product corresponding to each operand in the first target input value and the first product corresponding to each operand in the second target input value are multiplied to obtain a first number of second products; the first number of second product results are added to obtain a first operation result; the first operation result and the first product corresponding to each operand in the third target input value are added to obtain a target output value. By splitting the target input value into operands whose byte length is less than or equal to the maximum byte length, a multiplication and accumulation operation is performed based on the operands and the weight coefficients corresponding to the operands, the precision of the input value during the operation is retained, and the accuracy of the multiplication and accumulation operation is improved.

[0125] Optionally, the step 32 of performing a multiplication and addition operation based on the first products corresponding to the operands to obtain the target output value includes:

[0126] Step 51: using the accelerator to perform a multiplication and addition operation based on the first products corresponding to the operands to obtain a first output value;

[0127] Step 52: Shift the first output value right by a second number of bits to obtain the target output value; the second number is equal to the number of bits corresponding to the second byte length.

[0128] The first product is equal to the product of the operand and the weight coefficient corresponding to the operand.

[0129] For example, using an 8-bit accelerator to perform multiplication and accumulation operations on three floating-point input values ​​of 0.5, 0.5, and 0 respectively refers to performing a multiplication operation on 0.5 and 0.5. According to the precision of the three input values ​​and the precision ranges corresponding to different byte lengths, the target byte length required for fixing the three input values ​​is determined. The precision of 0.5 is 1 decimal place, and the precision range corresponding to 1 byte length includes 2 decimal places. Therefore, the target byte length for fixed-point operations on 0.5, 0.5, and 0 is 1 byte. The maximum value of a signed integer that can be represented by a 1-byte binary number is 127, 0.5×127=63.5, and 63.5 is rounded to 64. The binary form of 64 is 1000000. The byte length of the target input value is equal to the maximum byte length that the accelerator can operate. There is no need to split the target input value. The multiplication and accumulation operations are directly performed based on the three target input values. 64 and 64 are multiplied to obtain the first output value 4096. The first output value 4096 is shifted right by 8 bits to obtain 16, and 16 is used as the target output value.

[0130] In an embodiment of the present application, an accelerator is used to perform a multiplication and addition operation based on the first product corresponding to each operand to obtain a first output value; the first output value is right shifted a second number of bits to obtain a target output value; the second number is equal to the number of bits corresponding to the second byte length.

[0131] In summary, the multiplication and accumulation method provided in the embodiment of the present application can perform fixed-point operations on the three floating-point input values ​​in the accelerator to obtain three target input values ​​in the form of integers, reduce the error of the operation result caused by the precision problem of floating-point numbers, and thus improve the accuracy of the multiplication and accumulation operation, and the fixed-point conversion of floating-point numbers to integers can improve the operation speed; when the first byte length corresponding to the target input value is greater than the maximum byte length that the accelerator can operate, the target input value is split into at least two operands, and a corresponding weight coefficient is assigned to each operand according to the position information of each operand in the target input value. The multiplication and accumulation operation originally performed based on the input value in the form of floating-point numbers is converted into a multiplication and accumulation operation based on at least two operands and the weight coefficients corresponding to the operands, which ensures the accuracy of the input value during the multiplication and accumulation operation and improves the accuracy of the multiplication and accumulation operation.

[0132] Device Embodiment

[0133] like Figure 3 As shown, Figure 3 A logic block diagram of a multiplication and accumulation operation device provided in an embodiment of the present application is shown.

[0134] The present application provides a multiplication and accumulation operation device, the device comprising:

[0135] A fixed-point operation module 310 is used to perform fixed-point operation on three floating-point input values ​​in the accelerator to obtain target input values ​​corresponding to the three input values;

[0136] A first determination module 320, configured to determine a maximum byte length that the accelerator can operate;

[0137] a splitting module 330, configured to split each target input value into at least two operands when a first byte length corresponding to the target input value is greater than the maximum byte length; and a second byte length of each operand is less than or equal to the maximum byte length;

[0138] A second determination module 340, configured to determine a weight coefficient corresponding to each operand according to position information corresponding to each operand in the corresponding target input value;

[0139] The multiplication-accumulation operation module 350 is used to use the accelerator to perform a multiplication-accumulation operation based on the operands corresponding to the three target input values ​​and the weight coefficients of the operands to obtain a target output value.

[0140] Optionally, the fixed-point operation module includes:

[0141] A first determination submodule, configured to determine a target byte length according to the precision of three input values ​​in the accelerator and the precision ranges corresponding to different byte lengths;

[0142] A second determination submodule is used to determine a first parameter according to the target byte length; the first parameter is a maximum value of a signed integer that can be represented by a binary number of the target byte length;

[0143] The calculation module is used to calculate the product of each input value and the first parameter to obtain a target input value corresponding to the input value.

[0144] Optionally, the second determining module includes:

[0145] A weight determination module, used to determine the weight corresponding to each position in the target input value based on the position value principle of the binary number system;

[0146] The third determination submodule is used to determine the weight coefficient corresponding to each operand according to the position information of each operand in the corresponding target input value and the weight corresponding to each position in the target input value.

[0147] Optionally, the product-addition operation module includes:

[0148] A first operation submodule, configured to perform a multiplication operation on each operand and a weight coefficient corresponding to each operand to obtain a first product corresponding to the operand;

[0149] The second operation submodule is used to perform a multiplication and addition operation based on the first products corresponding to the respective operands to obtain the target output value.

[0150] Optionally, the second operator module includes:

[0151] a multiplication operation module, configured to multiply a first product corresponding to each operand in the first target input value and a first product corresponding to each operand in the second target input value to obtain a first number of second products; the first number being the product of the number of each operand in the first target input value and the number of each operand in the second target input value;

[0152] A first addition operation module, used for adding the first number of second multiplication results to obtain a first operation result;

[0153] The second addition operation module is used to add the first operation result and the first product corresponding to each operand in the third target input value to obtain a target output value.

[0154] Optionally, the second operator module includes:

[0155] A third operation submodule, configured to use the accelerator to perform a multiplication and addition operation based on the first products corresponding to the respective operands to obtain a first output value;

[0156] The right shift module is used to right shift the first output value by a second number of bits to obtain the target output value; the second number is equal to the number of bits corresponding to the second byte length.

[0157] In summary, the multiplication and accumulation operation device provided in the embodiment of the present application can perform fixed-point operation on three floating-point input values ​​in the accelerator to obtain three target input values ​​in the form of integers, reduce the operation result error caused by the floating-point precision problem, and thus improve the accuracy of the multiplication and accumulation operation, and the fixed-point conversion of floating-point numbers to integers can improve the operation speed; when the first byte length corresponding to the target input value is greater than the maximum byte length that the accelerator can operate, the target input value is split into at least two operands, and a corresponding weight coefficient is assigned to each operand according to the position information of each operand in the target input value. The multiplication and accumulation operation originally performed based on the input value in the form of floating-point numbers is converted into a multiplication and accumulation operation based on at least two operands and the weight coefficients corresponding to the operands, which ensures the accuracy of the input value during the multiplication and accumulation operation and improves the accuracy of the multiplication and accumulation operation.

[0158] The multiplication and accumulation operation device in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or it can be other devices other than the terminal. Exemplarily, the electronic device can be a GPU BOX, a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0159] The multiplication and accumulation operation device provided in the embodiment of the present application can realize Figure 1 To avoid repetition, the various processes implemented by the method embodiment are not described here.

[0160] Alternatively, if Figure 4 As shown, an embodiment of the present application also provides an electronic device, including a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the various steps of the above-mentioned multiplication and addition operation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0161] In an embodiment of the present application, the memory may be used to store software programs and various data. The memory may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory may include a volatile memory or a non-volatile memory, or the memory may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0162] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor.

[0163] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned multiplication and addition operation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0164] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0165] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned multiplication and addition method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0166] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises one..." does not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0167] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0168] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A method for multiplying and adding, characterized in that: The method comprises: Performing fixed-point operations on three floating-point input values ​​in the accelerator to obtain target input values ​​corresponding to the three input values; Determining the maximum byte length that the accelerator can operate; For each target input value, if the first byte length corresponding to the target input value is greater than the maximum byte length, split the target input value into at least two operands; the second byte length of each operand is less than or equal to the maximum byte length; Determine the weight coefficient corresponding to each operand according to the position information corresponding to each operand in the corresponding target input value; The accelerator is used to perform a multiplication and accumulation operation based on operands corresponding to the three target input values ​​and weight coefficients of the operands to obtain a target output value.

2. The method according to claim 1, characterized in that The fixed-point operation is performed on the three floating-point input values ​​in the accelerator to obtain target input values ​​corresponding to the three input values, including: Determining a target byte length according to the precision of three input values ​​in the accelerator and the precision ranges corresponding to different byte lengths; Determine a first parameter according to the target byte length; the first parameter is the maximum value of a signed integer that can be represented by a binary number of the target byte length; For each input value, the product of the input value and the first parameter is calculated to obtain a target input value corresponding to the input value.

3. The method according to claim 1, characterized in that The step of determining the weight coefficient corresponding to each operand according to the position information corresponding to each operand in the corresponding target input value includes: Determine the weight corresponding to each position in the target input value based on the place value principle of the binary number system; The weight coefficient corresponding to each operand is determined according to the position information of each operand in the corresponding target input value and the weight corresponding to each position in the target input value.

4. The method according to claim 1, characterized in that: The step of using the accelerator to perform a multiplication and accumulation operation based on operands corresponding to the three target input values ​​and weight coefficients of the operands to obtain a target output value includes: Multiplying each operand by a weight coefficient corresponding to each operand to obtain a first product corresponding to the operand; A multiplication and addition operation is performed based on the first products corresponding to the respective operands to obtain the target output value.

5. The method according to claim 4, characterized in that The step of performing a multiplication and addition operation based on the first products corresponding to the operands to obtain the target output value comprises: Multiplying the first products corresponding to the operands in the first target input value and the first products corresponding to the operands in the second target input value to obtain a first number of second products; the first number is the product of the number of the operands in the first target input value and the number of the operands in the second target input value; Adding the first number of second multiplication results to obtain a first operation result; The first operation result and the first product corresponding to each operand in the third target input value are added to obtain a target output value.

6. The method according to claim 4, characterized in that The performing a multiplication and addition operation based on the first products corresponding to the operands to obtain the target output value includes: Using the accelerator to perform a multiplication-addition operation based on the first products corresponding to the operands to obtain a first output value; The first output value is shifted right by a second number of bits to obtain the target output value; the second number is equal to the number of bits corresponding to the second byte length.

7. A multiplication and accumulation operation device, characterized in that: The device comprises: A fixed-point operation module is used to perform fixed-point operation on three floating-point input values ​​in the accelerator to obtain target input values ​​corresponding to the three input values; A first determination module, used to determine the maximum byte length that the accelerator can operate; a splitting module, configured to split each target input value into at least two operands when a first byte length corresponding to the target input value is greater than the maximum byte length; and a second byte length of each operand is less than or equal to the maximum byte length; A second determination module, configured to determine a weight coefficient corresponding to each operand according to position information corresponding to each operand in the corresponding target input value; The multiplication-accumulation operation module is used to use the accelerator to perform multiplication-accumulation operations based on the operands corresponding to the three target input values ​​and the weight coefficients of the operands to obtain the target output value.

8. The device according to claim 7, characterized in that The fixed-point operation module comprises: A first determination submodule, configured to determine a target byte length according to the precision of three input values ​​in the accelerator and the precision ranges corresponding to different byte lengths; A second determination submodule is used to determine a first parameter according to the target byte length; the first parameter is a maximum value of a signed integer that can be represented by a binary number of the target byte length; The calculation module is used to calculate the product of each input value and the first parameter to obtain a target input value corresponding to the input value.

9. The device according to claim 7, characterized in that The second determining module comprises: A weight determination module, used to determine the weight corresponding to each position in the target input value based on the position value principle of the binary number system; The third determination submodule is used to determine the weight coefficient corresponding to each operand according to the position information of each operand in the corresponding target input value and the weight corresponding to each position in the target input value.

10. The device according to claim 7, characterized in that The multiplication and accumulation operation module comprises: A first operation submodule, configured to perform a multiplication operation on each operand and a weight coefficient corresponding to each operand to obtain a first product corresponding to the operand; The second operation submodule is used to perform a multiplication and addition operation based on the first products corresponding to the respective operands to obtain the target output value.

11. The device according to claim 10, characterized in that The second operator module comprises: a multiplication operation module, configured to multiply a first product corresponding to each operand in the first target input value and a first product corresponding to each operand in the second target input value to obtain a first number of second products; the first number being the product of the number of each operand in the first target input value and the number of each operand in the second target input value; A first addition operation module, used for adding the first number of second multiplication results to obtain a first operation result; The second addition operation module is used to add the first operation result and the first product corresponding to each operand in the third target input value to obtain a target output value.

12. The device according to claim 10, characterized in that The second operator module comprises: A third operation submodule, configured to use the accelerator to perform a multiplication and addition operation based on the first products corresponding to the respective operands to obtain a first output value; The right shift module is used to right shift the first output value by a second number of bits to obtain the target output value; the second number is equal to the number of bits corresponding to the second byte length.

13. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the multiplication and accumulation method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN115756385A

  • Arithmetic unit, processor and electronic equipment

    CN117270813A

  • Large integer multiplication method, device and equipment

    CN117608522A

Cited By

  • Error upper bound device suitable for dynamic precision floating point multiply-accumulate operation, judgment method, medium, terminal and program product

    CN121523639A