Computing method, apparatus, chips, electronic devices and storage medium

KR103024922B1Active Publication Date: 2026-09-29KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
KR1020230063492
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-20
Filing Date
2023-05-17
Publication Date
2026-09-29
Estimated Expiration
2043-05-17

Smart Images

  • Figure 112023054523446-PAT00020_ABST
    Figure 112023054523446-PAT00020_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer technology, and more particularly to the field of chip technology and artificial intelligence, provides a method of computation, an apparatus, a chip, an electronic device, and a storage medium. An implementation method comprises obtaining a plurality of first fixed points and a plurality of first exponents corresponding to a plurality of first floating points and a plurality of second floating points corresponding to a plurality of second floating points, based on a plurality of first floating points of a first vector input to a computation device; obtaining a fixed-point multiplication value and a corresponding fixed-point multiplication exponent of each first fixed point among the plurality of first fixed points and a second fixed point corresponding thereto; obtaining a fixed-point dot product calculation result of the first vector and the second vector based on a fixed-point multiplication exponent corresponding to each fixed-point multiplication value among the plurality of fixed-point multiplication values; and obtaining a floating-point dot product calculation result in a floating-point data format corresponding to the fixed-point dot product calculation result based on the fixed-point dot product calculation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to the field of computer technology, particularly to the field of chip technology and artificial intelligence, and specifically to a method of computation, a device, a chip, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the advancement of artificial intelligence technology, an increasing number of applications based on AI technology have achieved results that far surpass existing algorithms; deep learning is currently a core technology of artificial intelligence. Deep learning is a data-intensive and computationally intensive algorithm, as well as an algorithm that evolves rapidly through iteration.

[0003] Conventional general-purpose processing devices such as CPUs, GPUs, and DSPs are designed for general-purpose computational tasks, but they have disadvantages such as low computational performance and low efficiency when processing deep learning applications, making it difficult to effectively support the large-scale deployment of deep learning algorithms in environments such as data centers. Dedicated deep learning accelerators based on ASICs / FPGAs can achieve higher computational performance and efficiency compared to existing devices such as CPUs, GPUs, and DSPs by deeply customizing the hardware structure to match the computational characteristics of deep learning.

[0004] The methods described in this section are not necessarily previously conceived or used. Unless otherwise specified, no method described in this section shall be considered prior art merely because it is included in this section. Likewise, unless otherwise specified, the problems raised in this section shall not be considered known in any prior art.

[0005] The present invention provides a computation method, a device, a chip, an electronic device, a computer-readable storage medium, and a computer program product. one all.

[0006] According to one aspect of the present invention, the present invention provides a method of operation performed by an arithmetic device, wherein the method comprises the steps of: obtaining a plurality of first fixed points and a plurality of first exponents of a binary representation corresponding to a plurality of first floating points and a plurality of second floating points of a plurality of second floating points, based on a plurality of first floating points of a first vector input to the arithmetic device and a plurality of second floating points of a second vector, wherein the plurality of first floating points correspond one-to-one with a plurality of second floating points, and each fixed point among the plurality of first fixed points and the plurality of second fixed points each comprises a sign bit and a first preset quantity of fixed point number bits; and obtaining a fixed point multiplication value and a corresponding fixed point multiplication exponent of each first fixed point among the plurality of first fixed points and a second fixed point corresponding to the first fixed point. The method includes the step of obtaining a fixed-point dot product calculation result of a first vector and a second vector based on a fixed-point product exponent corresponding to each fixed-point multiplication value among a plurality of fixed-point multiplication values ​​corresponding to a plurality of first fixed-point points; and the step of obtaining a floating-point dot product calculation result of a floating-point data format corresponding to the fixed-point dot product calculation result based on the fixed-point dot product calculation result.

[0007] According to another aspect of the present invention, the present invention provides a method of operation performed by an arithmetic device, wherein the method comprises: a step of obtaining a first matrix and a second matrix, wherein the first matrix includes a row vector of a first quantity and the second matrix includes a column vector of a second quantity, and the vector lengths of the row vector and the column vector are the same; and a step of obtaining the result of the inner product of each row vector in the first matrix and each column vector in the second matrix, respectively, according to the method of operation performed by the arithmetic device for calculating the inner product of the vectors described above, and obtaining the result of the inner product of the first matrix and the second matrix.

[0008] According to another aspect of the present invention, the present invention provides an arithmetic device, wherein the arithmetic device is configured to obtain, based on a plurality of first floating-point numbers of a first vector input to the arithmetic device and a plurality of first fixed-point numbers and a plurality of first exponents of a binary representation corresponding to a plurality of first floating-point numbers, and a plurality of second fixed-point numbers and a plurality of second exponents of a binary representation corresponding to a plurality of second floating-point numbers, wherein the plurality of first floating-point numbers correspond one-to-one with a plurality of second floating-point numbers, and each fixed-point number among the plurality of first fixed-point numbers and the plurality of second fixed-point numbers comprises a first acquisition unit including a sign bit and a first preset quantity of fixed-point number bits; and a multiplier configured to obtain a fixed-point multiplication value and a corresponding fixed-point multiplication exponent of each first fixed-point number among the plurality of first fixed-point numbers and a second fixed-point number corresponding to the first fixed-point number. It includes a second acquisition unit configured to acquire a fixed-point dot product calculation result of a first vector and a second vector based on a fixed-point product exponent corresponding to each of a plurality of fixed-point product values ​​corresponding to each of a plurality of first fixed-point products; and a third acquisition unit configured to acquire a floating-point dot product calculation result of a floating-point data format corresponding to the fixed-point dot product calculation result based on the fixed-point dot product calculation result.

[0009] According to another aspect of the present invention, the present invention provides an arithmetic device, wherein the arithmetic device is configured to acquire a first matrix and a second matrix, wherein the first matrix includes a row vector of a first quantity and the second matrix includes a column vector of a second quantity, and the vector lengths of the row vector and the column vector are the same; and a fifth acquisition unit configured to acquire the results of the inner product of each row vector in the first matrix and each column vector in the second matrix, respectively, according to an arithmetic method performed by the arithmetic device for calculating the inner product of the vectors described above, and thereby acquire the inner product result matrix of the first matrix and the second matrix.

[0010] According to another aspect of the present invention, a chip is provided comprising at least one of an arithmetic device for calculating the inner product of the vector and an arithmetic device for calculating the inner product of the matrix.

[0011] According to another aspect of the present invention, an electronic device comprising the chip is provided.

[0012] According to another aspect of the present invention, an electronic device is provided comprising at least one processor; and a memory connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform an operation method for calculating the inner product of a vector or an operation method for calculating the inner product of a matrix.

[0013] According to another aspect of the present invention, the present invention provides a non-transient computer-readable storage medium in which computer instructions are stored, wherein the computer instructions enable the computer to perform an operation method for calculating the inner product of the vector or an operation method for calculating the inner product of the matrix.

[0014] According to another aspect of the present invention, the present invention provides a computer program stored in a computer-readable storage medium, said computer program comprising instructions, wherein the instructions implement an operation method for calculating the inner product of said vectors or an operation method for calculating the inner product of said matrices when executed by at least one processor.

[0015] According to one or more embodiments of the present invention, floating-point data input to the computing device is converted into fixed-point data through operations within the computing device, and related operations can be completed based on the fixed-point data, thereby implementing a senseless operation of converting floating-point data into fixed-point data and saving labor costs and computing resources.

[0016] It should be understood that the description in this section does not indicate the core or important features of the embodiments of the present invention, nor does it limit the scope of the present invention. Other features of the present invention will be easily understood through the specification below. Brief explanation of the drawing

[0017] The attached drawings illustrate embodiments exemplarily and constitute part of the specification, and together with the description in the specification, describe exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Identical reference numerals in all drawings are similar but not necessarily identical elements. FIG. 1 illustrates a flowchart of a calculation method performed by a calculation device according to an embodiment of the present invention. FIG. 2 illustrates a structural block diagram of a vector multiplication device according to an embodiment of the present invention. FIG. 3 illustrates a flowchart for obtaining a plurality of first fixed points and a plurality of first exponents and a plurality of second fixed points and a plurality of second exponents according to an embodiment of the present invention. FIG. 4 illustrates a flowchart for obtaining a fixed-point multiplication value corresponding to a plurality of first fixed points and a corresponding fixed-point multiplication exponent according to an embodiment of the present invention. FIG. 5 illustrates a structural block diagram of a vector multiplication device according to another embodiment of the present invention. FIG. 6 illustrates a flowchart of an operation method performed by an operation device that calculates matrix multiplication according to an embodiment of the present invention. FIG. 7 illustrates a structural block diagram of a matrix multiplication device of an embodiment of the present invention. FIG. 8 illustrates a structural block diagram of an arithmetic device for calculating the inner product of vectors according to an embodiment of the present invention. FIG. 9 illustrates a structural block diagram of an arithmetic device for calculating matrix multiplication according to an embodiment of the present invention. FIG. 10 illustrates a structural block diagram of an exemplary electronic device that can be used to implement an embodiment of the present invention. Specific details for implementing the invention

[0018] Exemplary embodiments of the present invention are described in conjunction with the drawings below. While various details of the embodiments are included herein to aid understanding, they should be understood as merely illustrative. Accordingly, those skilled in the art should understand that various modifications and variations may be made to the embodiments described herein without departing from the scope of the present invention. Likewise, for clarity and brevity, descriptions of known functions and structures are omitted from the following description.

[0019] In the present invention, unless otherwise noted, terms such as "first," "second," etc. are used to describe various elements and are not intended to limit the positional, chronological, or significant relationships of these elements; rather, these terms are used merely to distinguish one element from another. In some examples, the first element and the second element may refer to the same embodiment of the element, whereas in some cases, they may refer to different embodiments based on contextual description.

[0020] The terms used in the description of the various examples in this invention are intended merely to describe specific examples and are not to be limiting. Unless otherwise clearly known from the context, the elements may be one or many, unless the quantity of the elements is specifically limited. Additionally, the term "and / or" as used in this invention includes any one of the listed items and all possible combinations.

[0021] Below, embodiments of the present invention will be described in detail with reference to the drawings.

[0022] The core operation of deep learning algorithms is matrix multiplication, and mainstream language models (e.g., BERT, ERNIE, etc.) include a large amount of matrix multiplication, and mainstream machine vision models (e.g., networks such as RESNET, MASK-RCNN, YOLO, SSD, etc.) include a large amount of convolution, and convolution operations are generally converted into matrix operations and implemented.

[0023] Deep learning networks all perform matrix multiplication operations based on data in a single-precision floating-point format (i.e., float format). With the advancement of deep learning, related engineers have discovered that using fixed-point precision, such as fixed-point data with 16-bit sign, for matrix multiplication and convolution operations can significantly reduce the area and power consumption of the hardware, achieve better performance-to-power ratio and performance-to-area ratio, and at the same time guarantee no significant loss of precision.

[0024] In the relevant technology, matrix multiplication operations within a deep learning network can be completed through a matrix multiplication unit within a deep learning chip. Although the matrix multiplication unit supports fixed-point operations, a technician must first convert floating-point precision data into fixed-point data using an independent unit, and then input the fixed-point data into the matrix multiplication unit to perform the operation.

[0025] An embodiment of the present invention provides a method of operation performed by an arithmetic device, and by converting floating-point data input to the arithmetic device into fixed-point data through an operation within the arithmetic device, and then performing vector multiplication or matrix multiplication operations based on the fixed-point data, it implements a senseless operation of converting floating-point data into fixed-point data and saves labor costs and computational resources.

[0026] According to an embodiment of the present invention, as illustrated in FIG. 1, a method of operation performed by an arithmetic device is provided, wherein the method of operation comprises: a step S101 of obtaining a plurality of first fixed points and a plurality of first exponents of a binary representation corresponding to a plurality of first floating points and a plurality of second floating points of a plurality of first floating points, and a plurality of second fixed points and a plurality of second exponents of a binary representation corresponding to a plurality of second floating points, wherein the plurality of first floating points correspond one-to-one with the plurality of second floating points, and each fixed point among the plurality of first fixed points and the plurality of second fixed points each includes a sign bit and a first preset quantity of fixed point number bits; and a step S102 of obtaining a fixed point multiplication value of each first fixed point among the plurality of first fixed points and a second fixed point corresponding to the first fixed point and a corresponding fixed point multiplication exponent. The method may include: a step S103 of obtaining a fixed-point dot product calculation result of a first vector and a second vector based on a fixed-point product exponent corresponding to each of a plurality of fixed-point product values ​​corresponding to each of a plurality of first fixed-point values; and a step S104 of obtaining a floating-point dot product calculation result of a floating-point data format corresponding to the fixed-point dot product calculation result based on the fixed-point dot product calculation result.

[0027] Thus, floating-point data input to the arithmetic unit can be converted into fixed-point data through operations within the arithmetic unit, and related operations can be completed based on the fixed-point data, thereby implementing a senseless operation of converting floating-point data into fixed-point data and saving labor costs and computational resources.

[0028] FIG. 2 illustrates a structural block diagram of a vector multiplication device of an embodiment of the present invention.

[0029] In some embodiments, the vector multiplication device (200) as illustrated in FIG. 2 may be a single arithmetic unit placed on a chip, and the arithmetic method may be implemented by the vector multiplication device (200). Specifically, the vector multiplication device (200) includes a data extraction module (210), a fixed-point multiplication module (220), an exponent comparison module (230), a digit shift module (240), an adder (250), and an inverse fixed-point module (260).

[0030] According to the IEEE 754 standard, single-precision floating-point data is represented as {S, E, M}, which includes a 1-bit sign bit S, an 8-bit exponent bit E, and a 23-bit floating-point number bit M, and the actual value N of the floating-point data can be calculated using the following formula.

[0031]

[0032] In some embodiments, first, a first vector A {a[0], a[1], … , a[k-1]} and a second vector B {b[0], b[1], … , b[k-1]} are input in parallel to a data extraction module (210) in a vector multiplication device (200), where k is a positive integer greater than 2. Each first floating-point a[i] in the first vector A and each second floating-point b[i] in the second vector B are all single-precision floating-point data (where i∈[0,k-1]), and the storage format is as described above.

[0033] In some embodiments, each first floating-point a[k-1] and each second floating-point b[k-1] can be processed first through a data extraction module (210) to convert them into a first fixed-point and a second fixed-point, respectively.

[0034] In some embodiments, as illustrated in FIG. 3, the step of obtaining a plurality of first fixed points and a plurality of first exponents of a binary representation corresponding to a plurality of first floating points, and a plurality of second fixed points and a plurality of second exponents of a binary representation corresponding to a plurality of second floating points, comprises performing the following operations for each floating point among the plurality of first floating points and the plurality of second floating points, wherein the operations include: step S301 of extracting a sign bit, an exponent bit, and a floating point number bit among the floating points, wherein the floating points are represented in binary; step S302 of extracting an upper number bit of a second preset quantity of the upper number bits among the floating point number bits; step S303 of determining a fixed point corresponding to the floating point based on the upper number bit and the sign bit of the floating point; and step S304 of determining an exponent corresponding to the floating point based on the exponent bit of the floating point.

[0035] Thus, since the input floating-point number can be converted to a fixed-point number through the computing device of the present invention, a senseless conversion of floating-point data into fixed-point data is implemented, and while improving the user experience, computational resources are saved by performing data fixed-point conversion processing without separate resources.

[0036] First, the sign bit, exponent bit, and floating-point number bit of each first floating-point a[i] and each second floating-point b[i] are extracted in parallel through the data extraction module (210), wherein the sign bit is recorded as a[i]

[31] and b[i]

[31] , the exponent bit is recorded as a[i][30:23] and b[i][30:23], and the floating-point number bit is recorded as a[i][22:0] and b[i][22:0].

[0037] Next, the upper number bits of the most significant second preset quantity are each extracted from the floating-point number bits among them.

[0038] In some embodiments, the first fixed-point and second fixed-point to be converted may be 16-bit fixed-point data, and correspondingly, the second preset quantity may be 14 bits, that is, for a[i][22:0] and b[i][22:0] respectively, the most significant 14-bit data is extracted, and the obtained upper number bits are a[i][22:9] and b[i][22:9] respectively.

[0039] In some embodiments, based on the following formula, a fixed point corresponding to each floating point can be obtained based on the upper number bit and the sign bit.

[0040]

[0041]

[0042] Here, a.us[i] and b.us[i] represent the first fixed-point and second fixed-point corresponding to the first floating-point a[i] and the second floating-point b[i], respectively, and also, among the two-digit data supplemented before a[i][22:9] and b[i][22:9], is the sign bit, and It is used to fill the integer before the omitted decimal point when storing floating-point data.

[0043] The first exponent ae[i] and the second exponent be[i], corresponding to the first fixed-point a.us[i] and the second fixed-point b.us[i] respectively, can be obtained based on the following formula.

[0044]

[0045]

[0046] In some embodiments, as illustrated in FIG. 4, the step of obtaining a fixed-point multiplication value and a corresponding fixed-point multiplication exponent of each first fixed point among a plurality of first fixed points and a second fixed point corresponding to the first fixed point may include: a step S401 of obtaining a plurality of first complements corresponding to a plurality of first fixed points and a plurality of second complements corresponding to a plurality of second fixed points; a step S402 of obtaining a fixed-point multiplication value by calculating the product of each first complement among a plurality of first complements and the second complement corresponding to the first complement; and a step S403 of obtaining a fixed-point multiplication exponent by calculating the sum of a first exponent corresponding to each first fixed point among a plurality of first fixed points and a second exponent corresponding to the second fixed point corresponding to the first fixed point.

[0047] Thus, by converting the fixed-point number to its complement and proceeding with subsequent operations so that the sign bit can also participate in the operation, there is no need to calculate the sign bit separately, thereby saving computational resources.

[0048] In some embodiments, based on the sign bits a[i]

[31] and b[i]

[31] corresponding to the first floating-point a[i] and the second floating-point b[i], respectively, the first complement as[i] and the second complement bs[i] corresponding to the first fixed-point a.us[i] and the second fixed-point b.us[i], respectively, can be obtained through the following formula.

[0049]

[0050]

[0051] Thus, through the data extraction module (210), the first fixed point a.us[i] and its complement as[i] and the second fixed point b.us[i] and its complement bs[i] are obtained correspondingly for each first floating point a[i] in the first vector A and each second floating point b[i] in the second vector B, respectively, and are transmitted to the fixed point multiplication module (220) through the data extraction module (210). At the same time, the data extraction module can transmit each first exponent ae[i] and each second exponent be[i] to the exponent comparison module (230).

[0052] In some embodiments, the fixed-point multiplication module (220) may include a fixed-point multiplication submodule of 16-bit fixed-point data of k groups, and the input data of the fixed-point multiplication module (220) may be each first fixed-point a.us[i] and each second fixed-point b.us[i], and the corresponding first fixed-point and second fixed-point are multiplied to obtain a corresponding fixed-point multiplication value, and the sign of the fixed-point multiplication value is determined through the sign bits of the first floating-point and second floating-point corresponding to the first fixed-point and the second fixed-point.

[0053] In some embodiments, the fixed-point multiplication module (220) may include a fixed-point multiplication submodule of k groups of 16-bit fixed-point data, and the input data of the fixed-point multiplication module (220) may also be a first complement as[i] corresponding to each first fixed point and a second complement bs[i] corresponding to each second fixed point. Additionally, a fixed-point multiplication value m[i] is obtained by multiplying the corresponding first complement as[i] and second complement bs[i] of each pair, and the specific formula is as follows.

[0054]

[0055] Here, the number of bits for each fixed-point multiplication value m[i] is 32 bits.

[0056] At the same time, correspondingly, through the exponent comparison module (230), a fixed-point product exponent ab.e[i] corresponding to each fixed-point product value m[i] can be obtained, and the specific formula is as follows.

[0057]

[0058] Since the exponents corresponding to each fixed-point multiplication value are different, before calculating the sum of each fixed-point multiplication value, each fixed-point multiplication value must first be unified into the same number range. For example, if the number bits of two fixed-point multiplication values ​​are both 1001, but the numbers corresponding to their exponents are 2 and -2, respectively, their actual values ​​are 100.1 and 0.01001, respectively (i.e., by shifting and aligning these two fixed-point multiplication values ​​based on the exponent 0 to unify them into the same number range), then the result obtained by calculating the sum again is accurate.

[0059] In related technologies, typically, before inputting data, the maximum value of each vector is obtained first, and the numerical range of each data point within the vector is unified based on the said maximum value. This method results in a loss of relatively high precision after the data is converted to fixed-point data.

[0060] In some embodiments, the step of obtaining the fixed-point inner product calculation result of a first vector and a second vector based on a fixed-point product exponent corresponding to each of a plurality of fixed-point product values ​​corresponding to each of a plurality of first fixed-point product values, comprises: a step of performing an arithmetic shift on the fixed-point product value based on a fixed-point product exponent corresponding to each of a plurality of fixed-point product values ​​corresponding to each of a plurality of first fixed-point product values; and a step of obtaining the fixed-point inner product calculation result of the first vector and the second vector by calculating the sum of the plurality of fixed-point product values ​​after arithmetic shifting corresponding to the plurality of first fixed-point product values.

[0061] Thus, after obtaining each fixed-point multiplication value, and before calculating the fixed-point inner product result, a digit shift is performed on each fixed-point multiplication value to unify the numerical range of each fixed-point multiplication value, which can reduce the loss of data precision and improve the accuracy of the calculation compared to methods among related technologies.

[0062] In some embodiments, each fixed-point multiplication value and each fixed-point multiplication exponent are input into a corresponding shift module (240) to perform an arithmetic shift on each fixed-point multiplication value based on the corresponding fixed-point multiplication exponent, thereby aligning each fixed-point multiplication value to the same number range, and then inputting the fixed-point multiplication value after alignment into an adder (250) to obtain the result of calculating the fixed-point inner product of the first vector and the second vector.

[0063] In some embodiments, the step of performing an arithmetic shift on a fixed-point multiplication value based on a fixed-point multiplication exponent corresponding to each of a plurality of fixed-point multiplication values ​​corresponding to each of a plurality of first fixed-point multiplication values ​​comprises: determining a first fixed-point multiplication exponent from a plurality of fixed-point multiplication exponents corresponding to a plurality of first fixed-point points; obtaining an arithmetic shift value corresponding to each of a fixed-point multiplication value among a plurality of fixed-point multiplication values ​​based on each of a fixed-point multiplication exponent among a plurality of fixed-point multiplication exponents and the first fixed-point multiplication exponent; and performing an arithmetic shift on a fixed-point multiplication value based on an arithmetic shift value corresponding to each of a fixed-point multiplication value among a plurality of fixed-point multiplication values.

[0064] In some embodiments, a first fixed-point product exponent can be determined from all fixed-point product exponents and used as a basis for arithmetic shifts; and a difference value between another fixed-point product exponent and the first fixed-point product exponent can be calculated to determine a position shift distance corresponding to each fixed-point product value. This allows for further simplification of operations, thereby improving computational efficiency and saving computational resources.

[0065] In some embodiments, the step of determining a first fixed-point product index from a plurality of fixed-point product indices corresponding to a plurality of first fixed-point points includes the step of obtaining the fixed-point product index with the largest value among the plurality of fixed-point product indices corresponding to a plurality of first fixed-point points and using it as the first fixed-point product index; additionally, the step of obtaining an arithmetic shift value corresponding to each fixed-point product value among a plurality of fixed-point product values ​​based on each fixed-point product index among the plurality of fixed-point product indices and the first fixed-point product index includes the step of calculating the difference value between the first fixed-point product index and each fixed-point product index among the plurality of fixed-point product indices and using it as the arithmetic shift value corresponding to the fixed-point product value corresponding to the fixed-point product index.

[0066] Thus, by determining the reference exponent, which is the first fixed-point product exponent, as the maximum exponent among all exponents, relatively large loss of precision for data with relatively large actual values ​​can be prevented, and the accuracy of the calculation can be guaranteed.

[0067] In some embodiments, the maximum exponent among all fixed-point product exponents ab.e[i] is first obtained through the exponent comparison module (230) and used as the first fixed-point product exponent ab.e.max, and the specific formula is as follows.

[0068]

[0069] Next, the difference between the first fixed-point product exponent ab.e.max and each fixed-point product exponent ab.e[i] is calculated to obtain each arithmetic shift value sft[i], and the specific formula is as follows.

[0070]

[0071] Each arithmetic shift value sft[i] output from the exponent comparison module (230) and each fixed-point multiplication value m[i] output from the fixed-point multiplication module (220) are input into the digit shift module (240), and the digit shift module (240) performs an arithmetic right shift for each fixed-point multiplication value m[i] based on its corresponding arithmetic shift value sft[i] to obtain the fixed-point multiplication value ms[i] after the corresponding arithmetic shift, and the specific formula is as follows.

[0072]

[0073] In some embodiments, the fixed-point multiplication value ms[i] after each arithmetic shift output from the position shift module (240) is input to the adder (250), and the result of the calculation of the fixed-point multiplication value after the arithmetic shift, si, can be obtained by calculating the sum of the fixed-point multiplication values ​​after the arithmetic shift based on the adder (250), and the specific formula is as follows.

[0074]

[0075] Here, the fixed-point dot product result si is signed fixed-point data with (32+j) bits, and also, here j is determined through the following formula.

[0076]

[0077] Here, is a function for rounding data. By doing so, it prevents data overflow during the process of calculating the sum by setting more bits in the fixed-point dot product result si.

[0078] In some embodiments, the step of obtaining a floating-point dot product result of a floating-point data format corresponding to a fixed-point dot product result based on the fixed-point dot product result can be processed through an inverse fixed-point module (260).

[0079] The adder (250) transmits the fixed-point dot product result si to the inverse fixed-point module (260), and first converts the fixed-point dot product result si into data sf in an equivalent floating-point format through the inverse fixed-point module (260).

[0080] At the same time, the exponent comparison module (230) transmits the first fixed-point product exponent ab.e.max to the inverse fixed-point conversion module (260), converts the first fixed-point product exponent ab.e.max into floating-point format exponent data df through the inverse fixed-point conversion module (260), and the specific formula is as follows.

[0081]

[0082] Next, the multiplication of sf and df is calculated through a floating-point multiplier in the inverse fixed-point module (260) to obtain the result of the floating-point dot product calculation of the first vector and the second vector in the floating-point data format, res, and the specific formula is as follows.

[0083]

[0084] In some embodiments, the arithmetic device can calculate the inner product of two vectors of a first preset length within one operation cycle, and for a third vector and a fourth vector having the same length and both being greater than the first preset length, the method further comprises the steps of: dividing the third vector and the fourth vector into a plurality of first vectors and a plurality of second vectors, respectively, based on the first preset length, wherein the plurality of first vectors and the plurality of second vectors correspond one-to-one; calculating the floating-point inner product calculation results of the corresponding groups of first vectors and second vectors, respectively; and calculating the sum of the floating-point inner product calculation results of the corresponding groups of first vectors and second vectors to obtain the inner product calculation results of the third vector and the fourth vector.

[0085] Thus, by dividing the input vector, inputting it to the computing device in different operation cycles, and accumulating the results obtained in each operation cycle, it is possible to implement the calculation of the inner product of a larger-scale vector.

[0086] FIG. 5 illustrates a structural block diagram of a vector multiplication device of another embodiment of the present invention.

[0087] In some embodiments, the vector multiplication device (500) as illustrated in FIG. 5 may be a single arithmetic unit placed on a chip, and the arithmetic method may be implemented by the vector multiplication device (500). Specifically, the vector multiplication device (500) includes a data extraction module (510), a fixed-point multiplication module (520), an exponent comparison module (530), a digit shift module (540), an adder (550), an inverse fixed-point module (560), and an accumulation module (570).

[0088] Here, the operation of modules (510) to (560) in the vector multiplication device (500) is similar to the operation of modules (210) to (260) in the vector multiplication device (200), so it is not described in detail here.

[0089] In some embodiments, a first vector and a second vector of a first preset length are sequentially extracted for a third vector and a fourth vector input to a vector multiplication device (500) within each operation cycle through a data extraction module (510), and a floating-point dot product calculation result of the first vector and the second vector is obtained in the operation cycle and stored in an accumulation module (570). After obtaining each floating-point dot product calculation result, the sum of all floating-point dot product calculation results is calculated through the accumulation module (570) to obtain a floating-point format dot product calculation result of the third vector and the fourth vector.

[0090] In some embodiments, as illustrated in FIG. 6, a method of operation performed by an operation device that calculates matrix multiplication is further provided, the method comprising: step S601 of obtaining a first matrix and a second matrix, wherein the first matrix includes a row vector of a first quantity and the second matrix includes a column vector of a second quantity, and the vector lengths of the row vector and the column vector are the same; and step S602 of obtaining the result of the inner product of each row vector in the first matrix and each column vector in the second matrix, respectively, according to a method of operation for calculating the inner product of the vectors, thereby obtaining the result of the inner product of the first matrix and the second matrix.

[0091] FIG. 7 illustrates a structural block diagram of a matrix multiplication device of an embodiment of the present invention.

[0092] In some embodiments, a matrix multiplication device as illustrated in FIG. 7 may be a single arithmetic unit placed on a chip, and the arithmetic method for calculating matrix multiplication may be implemented through the matrix multiplication device.

[0093] Here, the first matrix contains m row vectors A(0), A(1), ..., A(m-1), and the second matrix contains n column vectors B(0), B(1), ..., B(n-1). A matrix multiplication device as shown in FIG. 7 includes m×n vector multiplication devices, and each of the vector multiplication devices may be a vector multiplication device (200) shown in FIG. 2 or a vector multiplication device (500) shown in FIG. 5, and each of the vector multiplication devices can calculate the vector inner product of row vector A and column vector B through the operation method of calculating the inner product of vectors, so that the result of the calculation of each vector multiplication device can be obtained to obtain the inner product result matrix of the first matrix and the second matrix.

[0094] In some embodiments, as illustrated in FIG. 8, an arithmetic device (800) is provided, wherein the arithmetic device (800) is configured to obtain a plurality of first fixed points and a plurality of first exponents of a binary representation corresponding to a plurality of first floating points and a plurality of second floating points of a plurality of second floating points, based on a plurality of first floating points of a first vector input to the arithmetic device and a plurality of second floating points of a second vector, wherein each fixed point among the plurality of first fixed points and the plurality of second fixed points each includes a sign bit and a first preset quantity of fixed point number bits; and a multiplier (820) configured to obtain a fixed point multiplication value and a corresponding fixed point multiplication exponent of each first fixed point among the plurality of first fixed points and a second fixed point corresponding to the first fixed point. It may include a second acquisition unit (830) configured to acquire a fixed-point dot product calculation result of a first vector and a second vector based on a fixed-point product exponent corresponding to each of a plurality of fixed-point product values ​​corresponding to each of a plurality of first fixed-point products; and a third acquisition unit (840) configured to acquire a floating-point dot product calculation result of a floating-point data format corresponding to the fixed-point dot product calculation result based on the fixed-point dot product calculation result.

[0095] Here, the operation of the unit (810)-unit (840) in the arithmetic device (800) is similar to the operation of steps S101 to S104 of the arithmetic method performed by the arithmetic device, so it is not described in detail here.

[0096] In some embodiments, the second acquisition unit may include a shifter configured to perform an arithmetic shift on a fixed-point multiplication value based on a fixed-point multiplication exponent corresponding to each of a plurality of fixed-point multiplication values ​​corresponding to a plurality of first fixed-point multiplication values; and an adder configured to obtain the result of calculating the fixed-point inner product of the first vector and the second vector by calculating the sum of a plurality of fixed-point multiplication values ​​after arithmetic shifting corresponding to a plurality of first fixed-point multiplication values.

[0097] In some embodiments, the shifter may include: a determination module configured to determine a first fixed-point product index from a plurality of fixed-point product indices corresponding to a plurality of first fixed-point points; an acquisition module configured to obtain an arithmetic shift value corresponding to each fixed-point product value among a plurality of fixed-point product values ​​based on each fixed-point product index among the plurality of fixed-point product indices and the first fixed-point product index; and a position shift module configured to perform an arithmetic shift on the fixed-point product value based on the arithmetic shift value corresponding to each fixed-point product value among the plurality of fixed-point product values.

[0098] In some embodiments, the determination module may be configured to obtain the largest fixed-point product exponent among a plurality of fixed-point product exponents corresponding to a plurality of first fixed-point products and use it as the first fixed-point product exponent; additionally, the acquisition module may be configured to calculate the difference between the first fixed-point product exponent and each fixed-point product exponent among the plurality of fixed-point product exponents and use it as the corresponding arithmetic shift value of the fixed-point multiplication value corresponding to the fixed-point product exponent.

[0099] In some embodiments, the first acquisition unit may be configured to perform the operation of the following sub-unit for each of the plurality of first floating-point numbers and the plurality of second floating-point numbers, wherein the first extraction sub-unit is configured to extract a sign bit, an exponent bit, and a floating-point number bit among the floating-point numbers, wherein the floating-point numbers are represented in binary; the second extraction sub-unit is configured to extract the upper number bit of the second preset quantity of the uppermost floating-point number bits; the first determination sub-unit is configured to determine a fixed-point number corresponding to the floating-point number based on the upper number bit; and the second determination sub-unit is configured to determine an exponent corresponding to the floating-point number based on the exponent bit of the floating-point number.

[0100] In some embodiments, the multiplier may include: an acquisition sub-unit configured to obtain a plurality of first complements corresponding to a plurality of first fixed points and a plurality of second complements corresponding to a plurality of second fixed points based on a plurality of first floating-point sign bits corresponding to a plurality of first fixed points and a plurality of second floating-point sign bits corresponding to a plurality of second fixed points; a first calculation sub-unit configured to obtain a fixed-point multiplication value by calculating the product of each first complement among the plurality of first complements and the second complement corresponding to the first complement; and a second calculation sub-unit configured to obtain a fixed-point product exponent by calculating the sum of a first exponent corresponding to each first fixed point among the plurality of first fixed points and a second exponent corresponding to the second fixed point corresponding to the first fixed point.

[0101] In some embodiments, the arithmetic device may calculate the inner product of two vectors of a first preset length within one operation cycle, and for a third vector and a fourth vector having the same length and both being greater than the first preset length, the device may further include: a partitioning unit configured to divide the third vector and the fourth vector into a plurality of first vectors and a plurality of second vectors, respectively, based on the first preset length, wherein the plurality of first vectors and the plurality of second vectors correspond one-to-one; a first computing unit configured to calculate the floating-point inner product calculation results of corresponding groups of first vectors and second vectors, respectively; and a second computing unit configured to calculate the sum of the floating-point inner product calculation results of corresponding groups of first vectors and second vectors to obtain the inner product calculation results of the third vector and the fourth vector.

[0102] In some embodiments, as illustrated in FIG. 9, a calculation device (900) is further provided, wherein the calculation device (900) is configured to obtain a first matrix and a second matrix, wherein the first matrix includes a row vector of a first quantity and the second matrix includes a column vector of a second quantity, and the vector lengths of the row vector and the column vector are the same, a fourth acquisition unit (910); and a fifth acquisition unit (920) configured to obtain the inner product result matrix of the first matrix and the second matrix by obtaining the inner product result of each row vector in the first matrix and each column vector in the second matrix, respectively, according to a calculation method performed by the calculation device for calculating the inner product of the vectors described above.

[0103] Here, the operation of the unit (910)-unit (920) in the arithmetic device 900 is similar to the operation of steps S601 to S602 of the arithmetic method performed by the arithmetic device that calculates the matrix multiplication, so it is not described here.

[0104] In some embodiments, a chip is provided that includes at least one of an arithmetic device for calculating the inner product of the vector and an arithmetic device for calculating the inner product of the matrix.

[0105] In some embodiments, an electronic device including the chip is provided.

[0106] In some embodiments, an electronic device is provided comprising: at least one processor; and a memory connected to the communication of at least one processor; wherein the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor so that at least one processor can perform an operation method for calculating the inner product of a vector or an operation method for calculating the inner product of a matrix.

[0107] In some embodiments, a non-transient computer-readable storage medium is provided in which computer instructions are stored, wherein the computer instructions cause the computer to perform an operation method for calculating the inner product of the vector or an operation method for calculating the inner product of the matrix.

[0108] In some embodiments, a computer program product is provided that includes a computer program, wherein the computer program, when executed by a processor, implements a method of operation for calculating the inner product of the vector or a method of operation for calculating the inner product of the matrix.

[0109] According to an embodiment of the present invention, an electronic device, a readable storage medium, and a computer program product are further provided.

[0110] Referring to FIG. 10, a structural block diagram of an electronic device (1000) that can be used as a server or client of the present invention is to be described, which is an example of a hardware device applicable to each embodiment of the present invention. The electronic device refers to various forms of digital electronic computer devices such as laptop computers, desktop computers, operating platforms, personal information terminals, servers, blade servers, large computers, and other suitable computers. The electronic device may also refer to various forms of mobile devices such as personal digital processing, cellular phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely exemplary and do not limit the implementation of the present invention described and / or requested herein.

[0111] As illustrated in FIG. 10, the electronic device (1000) includes a computing unit (1001) capable of performing various appropriate operations and processing according to a computer program stored in a read-only memory (ROM) (1002) or a computer program loaded into a random access memory (RAM) (1003) from a storage unit (1008). The RAM (1003) also stores various programs and data necessary for the operation of the electronic device (1000). The computing unit (1001), ROM (1002), and RAM (1003) are connected to each other via a bus (1004). An input / output (I / O) interface (1005) is also connected to the bus (1004).

[0112] A plurality of components including an input unit (1006), an output unit (1007), a storage unit (1008), and a communication unit (1009) among the electronic device (1000) are connected to an I / O interface (1005). The input unit (1006) may be any type of device capable of inputting information to the electronic device (1000), and the input unit (1006) may be used to receive input numeric or character information and to generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keypad, touchscreen, track panel, trackball, joystick, microphone and / or remote control. The output unit (1007) may be any type of device for displaying information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator and / or printer. The storage unit (1008) may include, but is not limited to, a magnetic disk or an optical disk. The communication unit (1009) enables the electronic device (1000) to exchange information / data with other devices through a computer network such as the Internet and / or various electronic communication networks, and includes, but is not limited to, a modem, an internet card, an infrared communication device, a wireless communication transceiver and / or a group of chips, e.g., a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMAX device, a cellular communication device and / or an analogue.

[0113] The computing unit (1001) may be various general-purpose and / or dedicated processing components having processing and computing functions. Some examples of the computing unit (1001) include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, a computing unit executing various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit (1001) performs each of the above-described methods and processing, such as a method of calculating the inner product of the vector or a method of calculating the inner product of the matrix. For example, in some embodiments, the method of calculating the inner product of the vector or the method of calculating the inner product of the matrix may be implemented by a computer software program, which is tangibly contained in a machine-readable medium such as a storage unit (1008). In some embodiments, part or all of the computer program may be loaded and / or installed in the electronic device (1000) via a ROM (1002) and / or a communication unit (1009). When a computer program is loaded into RAM (1003) and executed by a computing unit (1001), it may perform one or more steps of the operation method for calculating the inner product of the vector or the operation method for calculating the inner product of the matrix described above. Alternatively, in another embodiment, the computing unit (1001) may be configured to perform the operation method for calculating the inner product of the vector or the operation method for calculating the inner product of the matrix by any other suitable method (e.g., with the help of firmware).

[0114] Various embodiments of the systems and technologies described in the text may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), dedicated integrated circuits (ASICs), dedicated standard products (ASSPs), systems on chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementation in one or more computer programs, said one or more computer programs may be executed and / or interpreted in a programmable system comprising at least one programmable processor, said programmable processor may be a dedicated or general-purpose programmable processor, may receive data and instructions from a storage system, at least one input device, and at least one output device, and may transmit data and instructions to said storage system, said at least one input device, and said at least one output device.

[0115] Program code implementing the method of the present invention may be written in any combination of one or more programming languages. Such program code is provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram can be implemented. The program code may be executed entirely on a machine, partially on a machine, or as a standalone software package on a machine, and may be partially executed on a remote machine or entirely on a remote machine or server.

[0116] In the context of the present invention, a machine-readable medium may be a tangible medium capable of containing or storing a program for use by a command execution system, device, or apparatus, or for use in combination with a command execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatus, or any suitable combination of the above. More specific examples of a machine-readable storage medium include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, CD-ROMs, optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0117] In order to provide interaction with a user, the system and technology described herein may be implemented on a computer, said computer having a display device for displaying information to a user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a directional device (e.g., a mouse or a trackball), said computer providing input to the computer through said keyboard and said directional device. Other types of devices may also provide interaction with a user, for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and may receive input from the user in any form (sound input, voice input, or tactile input).

[0118] The systems and technologies described herein may be implemented in a computing system comprising a backend component (e.g., a data server), a computing system comprising a middleware component (e.g., an application server), a computing system comprising a frontend component (e.g., a user computer equipped with a graphical user interface or a web browser, through which the user may interact with embodiments of the systems and technologies described herein), or any combination of such backend component, middleware component, or frontend component. Components of the systems may be connected to each other through digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0119] A computer system may include clients and servers. Clients and servers are typically located far apart from each other and generally interact through a communication network. The relationship between the client and server is established through computer programs that run on corresponding computers and also have a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server combined with a blockchain.

[0120] It should be understood that steps can be rearranged, added, or deleted using the various forms of processes described above. For example, each step described in this invention may be performed simultaneously, sequentially, or in a different order, and the text is not limited thereto as long as it enables the realization of the result intended by the technical solution disclosed in this invention.

[0121] Although embodiments or examples of the present invention have been described with reference to the drawings, it should be understood that the methods, systems, and devices described herein are merely exemplary embodiments or examples, and the scope of the present invention is not limited to these embodiments or examples, but is limited only to the claims and their equivalents thereof. Various elements in the embodiments or examples may be omitted or replaced with equivalent elements. Additionally, each step may be performed in a different order than described in the present invention. Furthermore, various elements in the embodiments or examples may be combined in various ways. Importantly, as technology advances, many elements described herein may be replaced by equivalent elements that appear later in the present invention.

Claims

Claim 1 A method of operation performed by an arithmetic device, comprising the step of acquiring a first matrix and a second matrix, wherein the first matrix includes a row vector of a first quantity and the second matrix includes a column vector of a second quantity, and the vector lengths of the row vector and the column vector are the same; The method comprises the step of determining each row vector of the first matrix and each column vector of the second matrix as the first vector and the second vector, respectively, and performing an inner product calculation operation on the first vector and the second vector to obtain the inner product results of each row vector of the first matrix and each column vector of the second matrix, respectively. The inner product calculation operation on the first vector and the second vector comprises: obtaining a plurality of first fixed points and a plurality of first exponents of a binary representation corresponding to the plurality of first floating points and a plurality of second floating points of a binary representation corresponding to the plurality of second floating points, based on a plurality of first floating points of the first vector and a plurality of second floating points of the second vector input to the computing device, wherein the plurality of first floating points correspond one-to-one with the plurality of second floating points, and each fixed point among the plurality of first fixed points and the plurality of second fixed points includes a sign bit and a first preset quantity of fixed point number bits. Step; obtaining a fixed-point multiplication value and a corresponding fixed-point multiplication exponent of each of the first fixed-points among the plurality of first fixed-points and a second fixed-point corresponding to the first fixed-point; and obtaining a fixed-point dot product calculation result of the first vector and the second vector based on a fixed-point multiplication exponent corresponding to each of the fixed-point multiplication values ​​among the plurality of fixed-point multiplication values ​​corresponding to each of the plurality of first fixed-points.A method of operation comprising: a step of obtaining a floating-point dot product result of a floating-point data format corresponding to the fixed-point dot product result based on the fixed-point dot product result; and a step of obtaining a matrix of dot product results of the first matrix and the second matrix based on the dot product result of each row vector in the first matrix and each column vector in the second matrix. Claim 2 In claim 1, the step of obtaining the result of calculating the fixed-point dot product of the first vector and the second vector based on a fixed-point product exponent corresponding to each of the fixed-point multiply values ​​among the plurality of fixed-point multiply values ​​corresponding to each of the plurality of first fixed-point values ​​comprises: a step of performing an arithmetic shift on the fixed-point multiply value based on a fixed-point product exponent corresponding to each of the fixed-point multiply values ​​among the plurality of fixed-point multiply values ​​corresponding to each of the plurality of first fixed-point values; and a step of obtaining the result of calculating the fixed-point dot product of the first vector and the second vector by calculating the sum of the plurality of fixed-point multiply values ​​after the arithmetic shift corresponding to the plurality of first fixed-point values. Claim 3 In claim 2, the step of performing an arithmetic shift on a fixed-point multiplication value based on a fixed-point multiplication exponent corresponding to each of a plurality of fixed-point multiplication values ​​corresponding to each of the plurality of first fixed-point multiplication values ​​comprises: a step of determining a first fixed-point multiplication exponent from a plurality of fixed-point multiplication exponents corresponding to the plurality of first fixed-point points; a step of obtaining an arithmetic shift value corresponding to each of the fixed-point multiplication values ​​among the plurality of fixed-point multiplication values ​​based on each of the fixed-point multiplication exponents and the first fixed-point multiplication exponent; and a step of performing an arithmetic shift on the fixed-point multiplication value based on an arithmetic shift value corresponding to each of the fixed-point multiplication values ​​among the plurality of fixed-point multiplication values. Claim 4 In claim 3, the step of determining a first fixed-point product index from a plurality of fixed-point product indices corresponding to a plurality of first fixed-point points includes the step of obtaining the fixed-point product index with the largest value among the plurality of fixed-point product indices corresponding to a plurality of first fixed-point points and using it as the first fixed-point product index; and the step of obtaining an arithmetic shift value corresponding to each fixed-point product value among the plurality of fixed-point product values ​​based on each fixed-point product index among the plurality of fixed-point product indices and the first fixed-point product index includes the step of calculating the difference value between the first fixed-point product index and each fixed-point product index among the plurality of fixed-point product indices and using it as the arithmetic shift value corresponding to the fixed-point product value corresponding to the fixed-point product index. Claim 5 A method of operation according to claim 1, wherein the step of obtaining a plurality of first fixed points and a plurality of first exponents of binary representations corresponding to a plurality of first floating points, and a plurality of second fixed points and a plurality of second exponents of binary representations corresponding to a plurality of second floating points, comprises the step of performing the following operation for each floating point among the plurality of first floating points and the plurality of second floating points, wherein the operation comprises: a step of extracting a sign bit, an exponent bit, and a floating point number bit among the floating points, wherein the floating points are represented in binary; a step of extracting the upper number bit of a second preset quantity of the uppermost number bit among the floating point number bits; a step of determining a fixed point corresponding to the floating point based on the upper number bit; and a step of determining an exponent corresponding to the floating point based on the exponent bit of the floating point. Claim 6 In claim 5, the step of obtaining a fixed-point multiplication value and a corresponding fixed-point multiplication exponent of each of the plurality of first fixed-points and a second fixed-point corresponding to the first fixed-point is a method of operation comprising: a step of obtaining a plurality of first complements corresponding to the plurality of first fixed-points and a plurality of second complements corresponding to the plurality of second fixed-points based on a sign bit of a plurality of first floating-points corresponding to each of the plurality of first fixed-points and a sign bit of a plurality of second floating-points corresponding to each of the plurality of second fixed-points; a step of obtaining the fixed-point multiplication value by calculating the product of each of the first complements among the plurality of first complements and the second complement corresponding to the first complement; and a step of obtaining the fixed-point multiplication exponent by calculating the sum of a first exponent corresponding to each of the first fixed-points among the plurality of first fixed-points and a second exponent corresponding to the second fixed-point corresponding to the first fixed-point. Claim 7 In claim 1, the arithmetic device can calculate the inner product of two vectors of a first preset length within one operation cycle, and for a third vector and a fourth vector having the same length and both being greater than the first preset length, the method further comprises the step of dividing the third vector and the fourth vector into a plurality of first vectors and a plurality of second vectors based on the first preset length, wherein the plurality of first vectors and the plurality of second vectors correspond one-to-one; the step of calculating the floating-point inner product calculation results of the corresponding groups of first vectors and second vectors, respectively; and the step of calculating the sum of the floating-point inner product calculation results of the corresponding groups of first vectors and second vectors to obtain the inner product calculation results of the third vector and the fourth vector. Claim 8 A fourth acquisition unit configured to acquire a first matrix and a second matrix as a computing device, wherein the first matrix includes a row vector of a first quantity and the second matrix includes a column vector of a second quantity, and the vector lengths of the row vector and the column vector are the same; The apparatus includes a fifth acquisition unit configured to determine each row vector of the first matrix and each column vector of the second matrix as the first vector and the second vector, respectively, perform an inner product calculation operation on the first vector and the second vector to obtain the inner product results of each row vector of the first matrix and each column vector of the second matrix, respectively, and obtain an inner product result matrix of the first matrix and the second matrix based on the inner product results of each row vector of the first matrix and each column vector of the second matrix, wherein the inner product calculation operation for the first vector and the second vector is implemented by the following units, and the units: obtain a plurality of first fixed-point numbers and a plurality of first exponents of a binary representation corresponding to the plurality of first floating-point numbers and a plurality of second fixed-point numbers and a plurality of second exponents of a binary representation corresponding to the plurality of second floating-point numbers, based on a plurality of first floating-point numbers of the first vector and a plurality of second floating-point numbers of the second vector input to the arithmetic device. A first acquisition unit configured such that the plurality of first floating-point numbers correspond one-to-one with the plurality of second floating-point numbers, and each fixed-point number among the plurality of first fixed-point numbers and the plurality of second fixed-point numbers includes a sign bit and a first preset quantity of fixed-point number bits; and a multiplier configured to acquire a fixed-point multiplication value and a corresponding fixed-point multiplication exponent of each first fixed-point number among the plurality of first fixed-point numbers and a second fixed-point number corresponding to the first fixed-point number.A computing device comprising: a second acquisition unit configured to acquire a fixed-point dot product calculation result of the first vector and the second vector based on a fixed-point product exponent corresponding to each of a plurality of fixed-point product values ​​corresponding to each of the plurality of first fixed-point values; and a third acquisition unit configured to acquire a floating-point dot product calculation result of a floating-point data format corresponding to the fixed-point dot product calculation result based on the fixed-point dot product calculation result. Claim 9 In claim 8, the second acquisition unit comprises: a shifter configured to perform an arithmetic shift on the fixed-point multiplication value based on a fixed-point multiplication exponent corresponding to each of the multiple fixed-point multiplication values ​​corresponding to each of the multiple first fixed-point multiplication values; and an adder configured to obtain the result of calculating the fixed-point inner product of the first vector and the second vector by calculating the sum of the multiple fixed-point multiplication values ​​after arithmetic shifting corresponding to the multiple first fixed-point values. Claim 10 In claim 9, the shifter comprises: a determination module configured to determine a first fixed-point product index from a plurality of fixed-point product indices corresponding to the plurality of first fixed-point points; an acquisition module configured to obtain an arithmetic shift value corresponding to each fixed-point product value among the plurality of fixed-point product values ​​based on each fixed-point product index among the plurality of fixed-point product indices and the first fixed-point product index; and a position shift module configured to perform an arithmetic shift on the fixed-point product value based on the arithmetic shift value corresponding to each fixed-point product value among the plurality of fixed-point product values. Claim 11 In claim 10, the above-mentioned determination module is configured to obtain the largest fixed-point product index among a plurality of fixed-point product indices corresponding to the plurality of first fixed-point points and use it as the first fixed-point product index, and the above-mentioned acquisition module is configured to calculate the difference value between the first fixed-point product index and each fixed-point product index among the plurality of fixed-point product indices and use it as the arithmetic shift value corresponding to the fixed-point multiplication value corresponding to the fixed-point product index. Claim 12 In claim 8, the first acquisition unit is configured to perform the operation of the following sub-unit for each of the plurality of first floating-point numbers and the plurality of second floating-point numbers, wherein the first extraction sub-unit is configured to extract a sign bit, an exponent bit, and a floating-point number bit among the floating-point numbers, wherein the floating-point numbers are represented in binary; the second extraction sub-unit is configured to extract the upper number bit of the second preset quantity of the uppermost floating-point number bits; the first determination sub-unit is configured to determine a fixed-point number corresponding to the floating-point number based on the upper number bits; and the second determination sub-unit is configured to determine an exponent corresponding to the floating-point number based on the exponent bit of the floating-point number. Claim 13 In claim 12, the multiplier comprises: an acquisition sub-unit configured to obtain a plurality of first complements corresponding to a plurality of first fixed points and a plurality of second complements corresponding to a plurality of second fixed points based on a plurality of first floating-point sign bits corresponding to a plurality of first fixed points and a plurality of second floating-point sign bits corresponding to a plurality of second fixed points; a first calculation sub-unit configured to obtain a fixed-point multiplication value by calculating the product of each first complement among the plurality of first complements and the second complement corresponding to the first complement; and a second calculation sub-unit configured to obtain a fixed-point multiplication exponent by calculating the sum of a first exponent corresponding to each first fixed point among the plurality of first fixed points and a second exponent corresponding to the second fixed point corresponding to the first fixed point. Claim 14 In claim 8, the arithmetic device can calculate the inner product of two vectors of a first preset length within one operation cycle, and for a third vector and a fourth vector having the same length and both being larger than the first preset length, the device is configured to divide the third vector and the fourth vector into a plurality of first vectors and a plurality of second vectors, respectively, based on the first preset length, wherein the plurality of first vectors and the plurality of second vectors correspond one-to-one; a dividing unit configured to calculate the floating-point inner product calculation results of several corresponding groups of first vectors and second vectors, respectively; and a second computing unit configured to calculate the sum of the floating-point inner product calculation results of several corresponding groups of first vectors and second vectors to obtain the inner product calculation results of the third vector and the fourth vector. Claim 15 A chip comprising a device according to any one of claims 8 to 14. Claim 16 An electronic device comprising a chip according to paragraph 15. Claim 17 An electronic device comprising: at least one processor; and a memory connected via communication to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to perform a method according to any one of claims 1 to 7. Claim 18 A non-transient computer-readable storage medium in which computer instructions are stored, wherein the computer instructions are used to enable the computer to perform a method according to any one of claims 1 through 7. Claim 19 A computer program stored on a computer-readable storage medium, wherein the computer program includes instructions, and when the instructions are executed by at least one processor, the computer program implements a method of operation according to any one of claims 1 to 7. Claim 20 delete Claim 21 delete

Citation Information

Patent Citations

  • Multiplication and accumulation circuits

    KR1020210057158A

  • Parallel processing method of vector product

    JP2009199430A

  • Convolutional neural network hardware configuration

    JP2022058660A

  • Method and apparatus for quantizing parameter of neural network

    KR1020190043849A

  • Neural network device for neural network operation, operating method of neural network device and application processor comprising neural network device

    KR1020210124894A