Arbitrary precision computing accelerator, integrated circuit device, board card and method
By generating pattern vectors and accumulated sequences through inner product operation processing components, the problem of low efficiency in arbitrary precision calculations in existing technologies is solved, and a high-efficiency, low-energy-consumption computing scheme is realized.
Patent Information
- Application Number
- CN202210990132.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2041-10-20
AI Technical Summary
Existing technologies cannot efficiently handle the variable-length operations required for arbitrarily precise calculations, resulting in low computational efficiency and frequent memory accesses. Existing solutions such as pure computation and approximate computation cannot meet the requirements for high precision.
The processing unit employing inner product operations includes a conversion unit, an inner product unit, and a synthesis unit. By generating a pattern vector and an accumulation unit accumulation sequence, it achieves efficient calculation of the inner product result and processes different bit streams in parallel to reduce redundant calculations.
It achieves efficient and low-energy-consumption arbitrary-precision calculations, reduces redundant calculations and memory accesses, and improves computational efficiency.
Smart Images

Figure CN115437602B_ABST
Abstract
Description
[0001] This application is a divisional application of application number 202111221317.4, filed on October 20, 2021, entitled "Arbitrary Precision Computing Accelerator, Integrated Circuit Device, Board and Method". Technical Field
[0002] This invention generally relates to the field of computers. More specifically, this invention relates to arbitrary-precision computing accelerators, integrated circuit devices, boards, and methods. Background Technology
[0003] Arbitrarily precise computation, which uses arbitrary numbers of bits to represent operands, is crucial in many technological fields, such as supernova simulation, climate simulation, atomic simulation, artificial intelligence, and planetary orbit calculation. These fields require processing data with hundreds, thousands, or even millions of bits, a range of data bits that far exceeds the hardware capabilities of traditional processors.
[0004] Even with high-bit-width processors, current technologies cannot handle the variable lengths required for arbitrarily precise computations because the optimal bit width varies significantly across different algorithms, and even subtle differences in bit width can lead to substantial cost variations. Furthermore, existing technologies have proposed numerous techniques to improve architecture-level computational efficiency, primarily effective-only computation and approximate computation. The former performs only basic computations, skipping or eliminating invalid calculations such as sparsity reduction and duplicate data, while the latter uses less accurate data, such as low-bit-width or quantized data, to replace computations with the original accurate data. However, finding duplicate data is extremely difficult and expensive for effective-only computation, and approximate computation intuitively contradicts the goal of arbitrarily precise computation, which requires precise calculations to achieve high accuracy. Finally, these existing technologies inevitably lead to a large number of inefficient memory accesses.
[0005] Therefore, an efficient and arbitrarily accurate calculation scheme is urgently needed. Summary of the Invention
[0006] In order to at least partially solve the technical problems mentioned in the background art, the present invention provides an arbitrary precision computing accelerator, integrated circuit device, board and method.
[0007] In one aspect, the present invention discloses a processing component for inner productting a first vector and a second vector, comprising: a conversion unit, a plurality of inner product units, and a synthesis unit. The conversion unit generates a plurality of pattern vectors based on the length and bit width of the first vector. Each inner product unit, using the data vector of the second vector in the length direction as an index, accumulates a specific pattern vector from the plurality of pattern vectors to form a unit accumulation sequence. The synthesis unit sums the plurality of unit accumulation sequences to obtain the inner product result.
[0008] In another aspect, this invention discloses an arbitrary-precision computing accelerator connected to off-chip memory. The arbitrary-precision computing accelerator includes a core memory proxy, a core controller, and a processing array. The core memory proxy reads multiple operands from off-chip memory. The core controller splits the multiple operands into multiple vectors, including a first vector and a second vector. The processing array includes multiple processing units, which perform an inner product of the first and second vectors based on their lengths to obtain an inner product result. The core controller integrates the inner product result into a computation result of the multiple operands, and the core memory proxy stores the computation result in off-chip memory.
[0009] In another aspect, the present invention discloses an integrated circuit device including the aforementioned arbitrary-precision computing accelerator, processing device, and off-chip memory. The processing device is used to control the arbitrary-precision computing accelerator, and the off-chip memory includes an LLC. The arbitrary-precision computing accelerator and the processing device are connected via the LLC.
[0010] In another aspect, the present invention discloses a board that includes the aforementioned integrated circuit device.
[0011] In another aspect, the present invention discloses a method for an inner product of a first vector and a second vector, comprising: generating multiple pattern vectors based on the length and bit width of the first vector; accumulating specific pattern vectors among the multiple pattern vectors based on the data vector of the second vector in the length direction as an index to form multiple unit accumulation sequences; and summing the multiple unit accumulation sequences to obtain the inner product result.
[0012] In another aspect, the present invention discloses an arbitrary precision calculation method, comprising: reading multiple operands from off-chip memory; splitting the multiple operands into multiple vectors, the multiple vectors including a first vector and a second vector; performing an inner product of the first vector and the second vector according to their lengths to obtain an inner product result; integrating the inner product result into a calculation result of the multiple operands; and storing the calculation result in off-chip memory.
[0013] In another aspect, the present invention discloses a computer-readable storage medium having stored thereon computer program code for arbitrary precision calculations, which, when run by a processing device, executes the aforementioned method.
[0014] This invention proposes a scheme for handling arbitrary precision calculations, processing different bit streams in parallel. It deploys a complete bit-serial data path for flexible and elastic high-precision computation. This invention fully utilizes simple hardware configurations, reduces redundant calculations, and thus achieves arbitrarily precise computations with low energy consumption. Attached Figure Description
[0015] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts. Wherein:
[0016] Figure 1 This is a structural diagram of a board card according to an embodiment of the present invention;
[0017] Figure 2 is a structural diagram illustrating an integrated circuit device according to an embodiment of the present invention;
[0018] Figure 3 This is a schematic diagram illustrating the internal structure of a computing device according to an embodiment of the present invention;
[0019] Figure 4 This is a schematic diagram illustrating an exemplary multiplication operation;
[0020] Figure 5 This is a schematic diagram illustrating the conversion unit according to an embodiment of the present invention;
[0021] Figure 6 This is a schematic diagram illustrating the generation unit of an embodiment of the present invention;
[0022] Figure 7 This is a schematic diagram illustrating the inner product unit of an embodiment of the present invention;
[0023] Figure 8 This is a schematic diagram illustrating the synthesis unit of an embodiment of the present invention;
[0024] Figure 9 This is a schematic diagram illustrating a full adder assembly according to an embodiment of the present invention;
[0025] Figure 10 This is a flowchart illustrating arbitrary precision calculation according to another embodiment of the present invention; and
[0026] Figure 11 This is a flowchart illustrating the inner product of a first vector and a second vector according to another embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] It should be understood that the terms "first," "second," "third," and "fourth," etc., in the claims, specification, and drawings of this invention are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" used in the specification and claims of this invention indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.
[0029] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.
[0030] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection."
[0031] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0032] Arbitrary-precision computation plays a crucial role in many scientific and technological fields. For example, the seemingly mundane equation x... 3 +y 3 +z 3 =3, solving this problem using a computer would require more than 200 bits of precision; in Ising theory, calculating the integral requires more than 1000 bits of precision; and calculating the volume of the knot complement in hyperbolic space involves as much as 60,000 bits of precision. Even a very small precision error can lead to huge differences in the calculation results, therefore arbitrary precision computation is a very serious technical problem in the field of computer science.
[0033] This invention proposes an efficient arbitrary-precision computing accelerator architecture, which mainly references the computational form of inner product operation and highlights the intra-parallelism and inter-parallelism of the accelerator architecture to realize operand multiplication operations.
[0034] Figure 1 A schematic diagram of the structure of a board 10 according to an embodiment of the present invention is shown. Figure 1 As shown, board 10 includes chip 101, which is a system-on-chip (SoC) integrating one or more combined processing units. These combined processing units are artificial intelligence computing units used to support various deep learning and machine learning algorithms, meeting the intelligent processing needs of complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in cloud intelligence. A significant characteristic of cloud intelligence applications is the large volume of input data, placing high demands on the platform's storage and computing capabilities. Board 10 in this embodiment is suitable for cloud intelligence applications, possessing massive off-chip storage, on-chip storage, and powerful computing capabilities.
[0035] Chip 101 is connected to external device 103 via external interface device 102. External device 103 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 103 to chip 101 via external interface device 102. The calculation results from chip 101 can be transmitted back to external device 103 via external interface device 102. Depending on the application scenario, external interface device 102 may have different interface forms, such as a PCIe interface.
[0036] The board 10 also includes a storage device 104 for storing data, which includes one or more memory cells 105. The storage device 104 is connected to and transmits data with the controller 106 and the chip 101 via a bus. The controller 106 in the board 10 is configured to regulate the state of the chip 101. Therefore, in one application scenario, the controller 106 may include a microcontroller (MCU).
[0037] Figure 2 is a structural diagram illustrating the combined processing device in chip 101 of this embodiment. As shown in Figure 2, the combined processing device includes a computing device 201, a processing device 202, off-chip memory 203, a communication node 204, and an interface device 205. In this embodiment, several integration schemes can be used to coordinate the operation of the computing device 201, the processing device 202, and the off-chip memory 203, among which... Figure 2AThis demonstrates the LLC integration scheme. Figure 2B This demonstrates the SoC integration solution. Figure 2C This illustrates the I / O integration scheme.
[0038] The computing device 201 is configured to execute user-specified operations, primarily implemented as a multi-core intelligent processor for performing deep learning or machine learning computations. It can interact with the processing device 202 to jointly complete the user-specified operations. The computing device 201 includes the aforementioned arbitrary-precision computing accelerator for processing linear computations, more specifically, operand multiplication operations applied in operations such as convolution.
[0039] The processing device 202, as a general-purpose processor, performs basic control functions including but not limited to data transfer, starting and / or stopping the computing device 201, and nonlinear calculations. Depending on the implementation, the processing device 202 can be one or more types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), or other general-purpose and / or special-purpose processors. These processors include, but are not limited to, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. When the computing device 201 and the processing device 202 are considered together, they are regarded as forming a heterogeneous multi-core structure.
[0040] Off-chip memory 203 is used to store data to be processed and data that has been processed. Its hierarchy, based on increasing latency, can be divided into: Level 1 cache (L1), Level 2 cache (L2), Level 3 cache (L3, also known as LLC), and physical memory. Physical memory is DDR, typically 16GB or larger. When computing device 201 or processing device 202 wants to read data from off-chip memory 203, L1 is usually accessed first because it is the fastest. If the data is not stored in L1, L2 is accessed next; if the data is not stored in L2 either, L3 is accessed; and if the data is still not stored in L3, DDR is accessed last. The cache hierarchy of off-chip memory 203 speeds up data access by storing the most frequently accessed data in the cache. Compared to the cache, DDR is considerably slower. As the cache level increases (L1→L2→LLC→DDR), the access latency increases, but the storage space increases.
[0041] Communication node 204 is a routing node or router in a network-on-chip (NoC) network. When computing device 201 or processing device 202 generates a data packet, it sends it to communication node 204 through a specific interface. Communication node 204 reads the address information in the header microchip of the data packet, calculates the optimal routing path using a specific routing algorithm, and establishes a reliable transmission path to deliver the data packet to the destination node (e.g., off-chip memory 203). Similarly, when computing device 201 or processing device 202 needs to read a data packet from off-chip memory 203, communication node 204 also calculates the optimal routing path and sends the data packet from off-chip memory 203 to computing device 201 or processing device 202.
[0042] Interface device 205 is the input / output interface of the combined processing device. When the combined processing device exchanges information with external devices, since there are many types of external devices and each device has different requirements for the information to be transmitted, interface device 205 will perform tasks such as setting up data buffer to solve the inconsistency caused by the speed difference between the two, setting up signal level conversion, setting up information conversion logic to meet the requirements of their respective formats, setting up timing control circuit to synchronize the work of the sender and receiver, and providing address transcoding, according to the requirements of the sender and receiver.
[0043] Figure 2A LLC integration refers to the connection between computing device 201 and processing device 202 via LLC. Figure 2B The SoC integration is achieved by integrating the computing device 201, the processing device 202, and the off-chip memory 203 through the communication node 204. Figure 2C The I / O integration is achieved by integrating the computing device 201, the processing device 202, and the off-chip memory 203 through the interface device 205. These three integration methods are merely examples, and the present invention does not limit the integration methods.
[0044] This embodiment preferably chooses an LLC integration scheme. Since the core of deep learning and machine learning is the convolution operator, and the foundation of the convolution operator is the inner product operation, which is a combination of multiplication and addition, the main task of the computing device 201 is a large number of low-level operations such as multiplication and addition. During the training and inference of the neural network model, the computing device 201 and the processing device 202 require intensive interaction. Integrating the computing device 201 and the processing device 202 into an LLC allows for data sharing, thus achieving lower interaction costs. Furthermore, since high-precision data may have millions of bits, the capacity of L1 and L2 caches is limited, and interaction through L1 and L2 caches can lead to insufficient capacity. The computing device 201 utilizes the relatively large capacity of the LLC to cache high-precision data, saving time spent on repeated accesses.
[0045] Figure 3 A schematic diagram of the internal structure of a computing device 201 is shown, which includes a core memory agent 301, a core controller 302, and a processing array 303.
[0046] The core memory agent 301 acts as the management terminal for the computing device 201 to access the off-chip memory 203. When the core memory agent 301 reads operands from the off-chip memory 203, the starting address of the operand is set in the core memory agent 301. The core memory agent 301 reads multiple operands simultaneously, continuously, and serially by incrementing the address. The reading method is to read from the least significant bit to the most significant bit of these operands one by one. For example, when three operands need to be read, the least significant bit (512 bits) of the first operand is read serially according to the starting address of each operand, then the least significant bit (512 bits) of the second operand is read serially, and then the least significant bit (512 bits) of the third operand is read serially. After the least significant bit is read, the address is incremented by 512 bits, and then the least significant bit (512 bits) of each subsequent operand is read serially, and so on until the most significant bit of these three operands is read. When the core memory proxy 301 stores the calculation results back to the off-chip memory 203, it sends them in parallel. For example, if the core memory proxy 301 needs to send three calculation results to the off-chip memory 203, it sends the least significant bit of each of the three calculation results simultaneously, then sends the next least significant bit, and so on, until the most significant bit of each of the three calculation results has been sent simultaneously. Generally, these operands are represented in the form of matrices or vectors.
[0047] Based on the computing power and number of processing units in the processing array 303, the core controller 302 controls the splitting of each operand into multiple data segments, or multiple vectors, so that the core memory agent 301 sends them to the processing array 303 in units of data segments.
[0048] Processing array 303 is used to perform multiplication of two operands. For example, the first operand can be divided into 8 data segments, x0 to x7, and the second operand can be divided into 4 data segments, y0 to y3. When the first operand and the second operand are multiplied, the algorithm expands as follows: Figure 4 As shown, the processing array 303 calculates the inner product by splitting the first operand and the second operand, and then shifting, aligning and summing the intermediate results 401, 402, 403 and 404 to obtain the result of the multiplication operation.
[0049] To clearly illustrate the technical solution, the above data segments will be uniformly represented as vectors. Multiplying two data segments is equivalent to taking the inner product of two vectors (the first vector and the second vector), where the first vector comes from the first operand and the second vector comes from the second operand.
[0050] The processing array 303 includes multiple processing units 304 arranged in an array. The figure exemplarily shows 4×8 processing units 304, but the number of processing units 304 is not limited in this invention. Each processing unit 304 is used to perform an inner product of the first vector and the second vector based on the lengths of the first and second vectors to obtain an inner product result. Finally, the core controller 302 controls the memory proxy 301 to integrate or reduce the inner product result into a calculation result of multiple operands and send it to the core memory proxy 301. The core memory proxy 301 stores the calculation result in off-chip memory 203.
[0051] Specifically, the computing device 201 employs a recursive decomposition algorithm in its control. When the computing device 201 receives an instruction from the processing device 202 to perform arbitrary-precision calculations, the core controller 302 divides the multiplication operands into multiple vectors on an average basis and sends them to the processing array 303 for calculation. Each processing unit 304 is responsible for calculating a set of vectors, such as the inner product of the first and second vectors. In this embodiment, each processing unit 304 further divides a set of vectors into smaller inner product calculation units based on its own hardware resources to facilitate inner product calculation. The computing device 201 uses a multi-bit stream on the data path, that is, each operand is imported from the core memory broker 301 to the processing unit 303 at a rate of 1 bit per cycle, but multiple operands are transmitted in parallel simultaneously. After the calculation is completed, the processing unit 304 sends the inner product result to the core memory broker 301 in a bit-serial manner.
[0052] As the core computing unit of computing device 201, the main task of processing unit 304 is inner product calculation. Processing unit 304 divides the process based on the inner product of bit index vectors into three stages: the first stage is the pattern generation stage, the second stage is the pattern indexing stage, and the third stage is the weighted synthesis stage.
[0053] With the first vector With the second vector Taking the inner product as an example, suppose the first vector With the second vector The sizes are N×p x With N×p y , where N is the first vector With the second vector The length of p, more specifically the number of row elements. x The first vector bit width, p y For the second vector The bit width. In this embodiment, the first vector is to be performed. With the second vector The inner product of the first vector is first... Transpose, then with the second vector Do the inner product, i.e. (p x ×N)·(N×p y ), to generate p x ×p y The inner product result.
[0054] This embodiment will use the second vector Disassembled into:
[0055]
[0056] Where K is a fixed value of size N×2. N Binary matrix, B col It is a size of 2 N ×p y A binary matrix, C is p y Weighted vector.
[0057] First vector The arrangement of elements along the length direction has 2 N The pattern, taking N=2 as an example, is the first vector. The length of K is 2, and K is based on the first vector. The length is divided into 2 N Given _n_ unit vectors, we can arrange them into all possible unit vectors of length 2. Therefore, K is a vector of size 2×2. 2 A binary matrix is used to cover all possibilities of combinations of elements of length 2. Combinations of elements of length 2 are: There are 4 possibilities, therefore the fixed form of K is:
[0058]
[0059] In other words, once the first vector With the second vector Once the length of K is determined, the size of K and the values of its elements are also determined.
[0060] B col It is a one-hot vector where each column has only one element that is 1, and the rest are 0. Which element is 1 depends on the second vector. Which column of K does this column correspond to? For ease of explanation, let's exemplarily define the first vector. With the second vector for:
[0061]
[0062]
[0063] The second vector Comparing it with K, we can find that the second vector The first column The fourth column of K, the second vector The second column The third column of K, the second vector The third column The fourth column of K, the second vector The fourth column Since it is the first column of K, therefore when the second vector With K·B col When used to represent, B col For a size of 2 2 The index matrix of ×4 is as follows:
[0064]
[0065] B col The first column has only the fourth element that is 1, indicating the second vector. The first column is the fourth column of K; B col The second column has only one third element that is 1, indicating the second vector. The second column is the third column of K; B col The third column has only one fourth element that is 1, indicating the second vector. The third column is the fourth column of K; B col The fourth column has only one element that is 1, indicating the second vector. The fourth column is the first column of K. In conclusion, once K is determined, B... col The element values are also determined.
[0066] C is p y The weighted vector is used to reflect the second vector. The power of p, which is the bit width. y The value is 4, indicating the second vector. Since the power of is 4, C is:
[0067]
[0068] This embodiment decomposes the second vector in the manner described above. Make the second vector The elements in can be represented by K and B. col It is represented by two binary matrices. In other words, this embodiment will... The inner product operation is converted to The operation.
[0069] Processing unit 304 is used to implement the vector inner product based on the aforementioned transformation. During the pattern generation phase, processing component 304 obtains... The various possibilities, i.e., generating pattern vectors During the pattern indexing stage, processing component 304 calculates... During the weighted composition stage, processing unit 304 accumulates the index pattern according to the weight C. This design allows operands of any precision to be converted into index patterns for inner product execution, reducing redundant calculations and avoiding the high bandwidth requirements of arbitrary precision computation.
[0070] Figure 3 A further schematic diagram of the processing unit 304 is shown. To achieve the aforementioned three stages, the processing unit 304 includes a processing unit memory proxy unit 305, a processing unit control unit 306, a conversion unit 307, multiple inner product units 308, and a synthesis unit 309.
[0071] The processing unit memory proxy unit 305 serves as the interface for the processing unit 304 to access the core memory proxy 301, and is used to receive two vectors that need to be used for inner product operations, such as the aforementioned first vector. With the second vector
[0072] The processing unit control unit 306 is used to coordinate and manage the work of each unit in the processing unit 304.
[0073] The conversion unit 307 is used to implement the pattern generation stage. The self-processing unit memory proxy unit 305 receives the first vector. And implement the binary matrix K in hardware, and execute... To generate multiple pattern vectors Figure 5 The diagram shows a conversion unit 307, which includes: N bitstream input terminals 501, a generation component 502, and 2... N 503 is a bitstream output terminal.
[0074] The N bitstream input terminals 501 are used to correspond to the first vector. The length of the array is N, and it receives N data vectors respectively. Figure 5 With the first vector The length of the first vector is 4. It includes four data vectors: x0, x1, x2, and x3, each with a bit width of p. x That is, each data vector has p x Single digits.
[0075] Component 502 is generated for execution. The core component. The response K has 2N Unit vectors, generating component 502 includes 2 N There are 2 generating units, each simulating a unit vector, to generate 2... N pattern vectors like Figure 5 As shown, the first vector The data is split into four data vectors: x0, x1, x2, and x3, and these vectors are input to the left side of the generator component 502 in parallel. Since the inner product operation in binary is essentially bit addition, the generator component 502 directly simulates all unit vectors in K in hardware, adding them sequentially to the bits of x0, x1, x2, and x3. More specifically, in each cycle, the corresponding bits of x0, x1, x2, and x3 are input simultaneously. For example, in the first cycle, the least significant bit of x0, x1, x2, and x3 is input simultaneously; in the second cycle, the second least significant bit of x0, x1, x2, and x3 is input simultaneously; and so on until the p-th cycle. x The cycle continues until the most significant bits of x0, x1, x2, and x3 are input simultaneously. The required bandwidth is only N bits per cycle; in this example, the required bandwidth is only 4 bits per cycle.
[0076] First vector When the length is 4, the generation component 502 includes 16 generation units, which respectively simulate 16 unit vectors in K. These unit vectors are (0000), (0001), (0010), (0011), (0100), (0101), (0110), (0111), (1000), (1001), (1010), (1011), (1100), (1101), (1110), and (1111).
[0077] Figure 6 A schematic diagram of a generation unit 504 with a unit vector of (1011) is shown. Taking generation unit 504 as an example, it simulates the unit vector (1011), so generation unit 504 includes three element registers 601, an adder 602, and a carry register 603. The three element registers 601 receive and temporarily store the bit values of the data vector corresponding to the simulated unit vector, that is, the bit values of x0, x1, and x3, and directly ignore the bit value of x2. This structure is used to implement the following:
[0078]
[0079] The value in temporary register 601 is sent to adder 602 for accumulation. If a carry occurs after accumulation, the carry value is temporarily stored in carry temporary register 603 and added to the bit values of x0, x1, and x3 input in the next cycle, until the p-th bit. xThe cycle continues until the most significant bits of x0, x1, and x3 are added together. Each generation unit is designed based on the same technical logic, which those skilled in the art can apply to this process. Figure 6 The structure of the generator unit 504, which implements the unit vector (1011), can be easily derived from the structure of other generator units without any creative effort, so it will not be described in detail. It should be noted that some generator units do not need to set up adders 602 and carry registers 603, such as the generator units that simulate unit vectors (0000), (0001), (0010), (0100), and (1000). These generator units have only one input in the same cycle, and there is no addition operation or carry.
[0080] Back Figure 5 ,2 N Each bitstream output terminal 503 is connected to the output of the adder 602 of each generation unit to output 2. N pattern vectors exist Figure 5 Since N is 4, the 503 output terminal outputs a total of 16 mode vectors for the 16 bitstreams. These pattern vectors The bit width could be p x (If the most significant bits are added without a carry), or p x +1 (carry if the most significant bits are added together). From Figure 5 It can be seen that the pattern vector For all possible combinations of addition operations on x0, x1, x2, and x3, that is:
[0081] z0 = 0
[0082] z1=x0
[0083] z2=x1
[0084] z3=x0+x1
[0085] z4 = x2
[0086] z5 = x0 + x2
[0087] z6=x1+x2
[0088] z7 = x0 + x1 + x2
[0089] z8 = x3
[0090] z9=x0+x3
[0091] z 10 =x1+x3
[0092] z 11 =x0+x1+x3
[0093] z 12 =x2+x3
[0094] z 13 =x0 + x2 + x3
[0095] z 14 =x1+x2+x3
[0096] z 15 =x0+x1+x2+x3
[0097] Pattern vector The vector is sent to inner product unit 308. This embodiment has multiple inner product units 308, each equivalent to a processor core, used to implement the pattern indexing stage and the weighted synthesis stage. This invention does not limit the number of inner product units 308. The inner product unit 308 receives the second vector from the processing unit memory proxy unit 305. With the second vector The data vector along the length direction is used as the index, and based on each index, all pattern vectors are used... Select the corresponding specific mode vector, accumulate these specific mode vectors, generate one bit of intermediate result in each cycle, and continuously p x or p x +1 cycles form a unit cumulative sequence. The above operation is performing...
[0098] Figure 7 A schematic diagram of the inner product unit 308 in this embodiment is shown. To achieve... Inner product unit 308 includes p y Multiplexer 701 and p y -1 serial full adder 702.
[0099] p y Multiplexers 701 are used to implement the mode indexing stage. Each multiplexer 701 receives all mode vectors. (z0 to z) 15 According to the second vector The data vectors in the same position along the length direction make all pattern vectors The specific pattern vector in the vector passes through. Because of the second vector... The length of the second vector is N, therefore... It can be decomposed into N data vectors. Since N is 4, the second vector... It can be decomposed into four data vectors: y0, y1, y2, y3, etc., and each data vector has a bit width of p. y Therefore, these data vectors can be decomposed into p bits from the perspective of the same bit position.y Each data vector has a corresponding bit. For example, the highest bit of the four data vectors y0, y1, y2, and y3 forms the highest bit corresponding bit data vector 703, the second highest bit of the four data vectors y0, y1, y2, and y3 forms the second highest bit corresponding bit data vector 704, and so on, with the lowest bit of the four data vectors y0, y1, y2, and y3 forming the lowest bit corresponding bit data vector 705.
[0100] Multiplexer 701 determines which unit vector in the binary matrix K the input corresponding data vector is identical to, and outputs the specific pattern vector corresponding to the identical unit vector. For example, the highest bit corresponding data vector 703 is used as a selection signal input to the first multiplexer. Assuming the highest bit corresponding data vector 703 is (0101), it is... Figure 5 If the unit vector 505 is the same as the unit vector 505, then the first multiplexer will output a specific mode vector z5 corresponding to the unit vector 505. For example, the second-highest bit co-occurring data vector 704 is used as a selection signal input to the second multiplexer. Assuming the second-highest bit co-occurring data vector 704 is (0010), and... Figure 5 If the unit vector 506 is the same as the unit vector 506, then the second multiplexer will output a specific mode vector z2 corresponding to the unit vector 506. Finally, the least significant bit data vector 705 is used as a selection signal input to the p-th multiplexer. y The multiplexer, assuming the least significant bit data vector 705 is (1110), and... Figure 5 If the unit vector 507 is the same, then the p-th... y The multiplexer will output a specific mode vector z corresponding to the unit vector 507. 14 This concludes the process. The operation.
[0101] The serial full adder 702 implements the weighted synthesis stage. y - One serial full adder 702 is connected serially as shown in the figure. It receives specific pattern vectors output by multiplexer 701 and sequentially accumulates these specific pattern vectors to obtain a unit accumulation sequence. It is particularly important to note that, to ensure accumulation from the least significant bit and carry-over (if any) to the next bit, the specific pattern vector corresponding to the least significant bit's corresponding data vector 705 must be input to the outermost serial full adder 702. This ensures that the specific pattern vectors corresponding to the least significant bit's corresponding data vector are accumulated first. The specific pattern vectors corresponding to higher-order corresponding data vectors are input to the innermost serial full adder 702. Similarly, the specific pattern vector corresponding to the most significant bit's corresponding data vector 703 must be input to the innermost serial full adder 702, ensuring that the specific pattern vectors corresponding to higher-order corresponding data vectors are accumulated later. This ensures the correctness of the accumulation, i.e., according to p... yThe weighted vector C reflects the second vector The power of . The unit cumulative sequence is in Based on this, we further implement the weighting of C. Thus, we obtain... Figure 4 Intermediate results 401, 402, 403, and 404.
[0102] Synthesis unit 309 is used to perform, for example Figure 4 The summation calculation in section 405. Synthesis unit 309 receives unit accumulation sequences from each inner product unit 308, each unit accumulation sequence being like... Figure 4 The intermediate results 401, 402, 403, and 404 are aligned in the inner product unit 308. Then, the synthesis unit 309 sums these aligned unit cumulative sequences to obtain the first vector. With the second vector The inner product result.
[0103] Figure 8 A schematic diagram of the synthesis unit 309 of this embodiment is shown. The synthesis unit 309 in the figure exemplarily receives the outputs of eight inner product units 308, namely, unit accumulation sequences 801 to 808. These unit accumulation sequences 801 to 808 are a first vector. With the second vector After being split into 8 data segments, the intermediate results are obtained by inner product calculations in 8 inner product units 308. The synthesis unit 309 includes 7 full adder groups 809 to 815. Since the least significant bit operation 816 and the most significant bit operation 817 have only one intermediate result, the least significant bit operation 816 and the most significant bit operation 817 do not require adder groups, as... Figure 4 The x0y0 (least significant bit) and x7y3 (most significant bit) values do not need to be added to other intermediate results and can be output directly. In other words, only the operations from the second least significant bit to the second most significant bit require a full adder group to perform operations such as... Figure 4 The summation shown is 405.
[0104] Figure 9A schematic diagram of full adder groups 810 to 815 is shown. Full adder groups 810 to 815 include a first full adder 901 and a second full adder 902. The first full adder 901 and the second full adder 902 each include multiplexers 903 and 904, respectively. The input of multiplexer 903 is connected to the carry output of the adder and the value 0, while the input of multiplexer 904 is connected to the carry output of the adder and the value 1. The values 0 and 1 are used to simulate the sum of the intermediate results of the previous digit without carry and carry-over, respectively. Therefore, the first full adder 901 generates the sum of the intermediate results of the previous digit without carry-over, and the second full adder 902 generates the sum of the intermediate results of the previous digit with carry-over. This structure eliminates the need to wait for the intermediate result of the previous digit to determine whether to carry. This embodiment adopts a design that simultaneously calculates the carry-over and carry-over, reducing computational latency. The full adder group 810 to 815 also includes a multiplexer 905. The sum of the two intermediate results is input to the multiplexer 905. The multiplexer 905 selects to output the sum of the intermediate results with or without carry, depending on whether the calculation result of the previous digit has a carry. The accumulated output 818 is the first vector. With the second vector The inner product result.
[0105] Back Figure 8 Since the operation of the least significant bit cannot generate a carry, the second least significant bit full adder group 809 only includes the first full adder 901, which directly generates the intermediate result without carry, without the need to set up the second full adder 902 and the multiplexer 905.
[0106] according to Figure 8 , Figure 9 According to the related description, when the synthesis unit 309 of this embodiment wants to sum up M unit cumulative sequences, it will be configured with M-1 full adder groups, including M-1 first full adders 901, M-2 second full adders 902 and M-2 multiplexers 905.
[0107] In other cases, the synthesis unit 309 can flexibly select to enable or disable the operation of the full adder group, such as the first vector. With the second vector When the generated unit cumulative sequence is less than M, a certain number of full adder groups can be turned off to flexibly support various possible split numbers and expand the application scenarios of the synthesis unit 309.
[0108] Back Figure 3 The first vector is obtained in the synthesis unit 309. With the second vector After obtaining the inner product result, it is sent to the processing unit memory proxy unit 305. The processing unit memory proxy unit 305 receives the inner product result and sends it to the core memory proxy 301. The core memory proxy 301 integrates the inner product results of all processing units 304 to generate the calculation result, which is then sent to the off-chip memory 203 to complete the product operation of the first operand and the second operand.
[0109] Based on the above structure, the computing device 201 of this embodiment performs different numbers of inner product operations depending on the length of the operands. Furthermore, the processing array 303 can control the sharing of indices among the vertical processing units 304 and control the sharing of pattern vectors among the horizontal processing units 304 to perform operations efficiently.
[0110] In data path management, this embodiment employs a two-level architecture: a core memory proxy 301 and a processing unit memory proxy unit 305. The starting address of the operand in the LLC is recorded in the core memory proxy 301, which reads multiple operands from the LLC simultaneously, continuously, and serially using an auto-incrementing address. Since the source address is auto-incrementing, the order of data blocks is deterministic. The core controller 302 determines which processing units 304 receive data blocks, and the processing unit control unit 306 then determines which inner product units 308 receive these data blocks.
[0111] Another embodiment of the present invention is an arbitrary precision calculation method, which can be implemented using the hardware structure of the foregoing embodiments. Figure 10 A flowchart illustrating this embodiment is shown.
[0112] In step 1001, multiple operands are read from off-chip memory. When reading operands from off-chip memory, the starting address of the operand is set in the kernel memory proxy. The kernel memory proxy reads multiple operands simultaneously, continuously, and serially by incrementing the address. The reading method is to read from the low-order bits of these operands to the high-order bits one by one.
[0113] In step 1002, multiple operands are split into multiple vectors, including a first vector and a second vector. Based on the computing power and number of processing units in the processing array, the core controller controls the splitting of each operand into multiple data segments, i.e., multiple vectors, so that the core memory agent sends them to the processing array in units of data segments.
[0114] In step 1003, the first vector and the second vector are doubly multiplied according to their lengths to obtain the doubly product result. The processing array includes multiple processing units arranged in an array. Each processing unit does the doubly product of the first vector and the second vector according to their lengths to obtain the doubly product result. More specifically, in this step, the pattern generation stage is executed first, followed by the pattern indexing stage, and finally the weighted synthesis stage.
[0115] With the first vector With the second vector Taking the inner product as an example, suppose the first vector With the second vector The sizes are N×p x With N×p y , where N is the first vector With the second vector The length of p x The first vector bit width, p y For the second vector The bit width. This embodiment also uses the second vector... Disassembled into:
[0116]
[0117] Where K is a fixed value of size N×2. N Binary matrix, B col It is a size of 2 N ×p y A binary matrix, C is p y Weighted vectors, K, B col The definition of C is the same as in the previous embodiment, so it will not be repeated. This embodiment decomposes the second vector in the manner described above. Make the second vector The elements in can be represented by K and B. col It is represented by two binary matrices. In other words, this embodiment will... The inner product operation is converted to The operation.
[0118] During the pattern generation phase, this embodiment obtains The various possibilities, i.e., generating pattern vectors In the schema indexing phase, this embodiment calculates... In the weighted synthesis stage, the index pattern is accumulated according to the weight C. This design allows operands of any precision to be converted into index patterns for inner product execution, reducing redundant calculations and avoiding the high bandwidth requirements of arbitrary precision calculations. Figure 11The flowchart of the inner product of the first and second vectors is further shown.
[0119] In step 1101, multiple pattern vectors are generated based on the length and bit width of the first vector. First, these are mapped to the first vector. The length N is given, and N data vectors are received respectively. Then the response K has 2 N Each unit vector is simulated in hardware to generate 2 unit vectors. N pattern vectors Since the inner product operation in binary is essentially the addition of each bit, the generation component in this embodiment directly simulates all unit vectors in K and the first vector. The bits of the data vector are added sequentially. More specifically, the first vector is input simultaneously in each cycle. The corresponding bits of the data vector are input simultaneously, for example, the least significant bit of the data vector is input simultaneously in the first cycle, the second least significant bit of the data vector is input simultaneously in the second cycle, and so on until the p-th cycle. x The cycle continues until the most significant bit of the input data vector is reached. The required bandwidth is only N bits per cycle.
[0120] When simulating a unit vector, the bit values of the data vector corresponding to that unit vector are first received and temporarily stored. These bit values are accumulated. If a carry occurs after accumulation, the carry value is temporarily stored in a carry register and added to the bit values of the data vector input in the next cycle, until the p-th cycle. x The cycle continues until the most significant bit value of the data vector is added together.
[0121] Finally, the accumulated result is received, which is the pattern vector. In summary, pattern vectors The first vector All possible combinations of addition operations on the data vector.
[0122] In step 1102, based on the second vector Using the data vector along the length direction as an index, specific pattern vectors from multiple pattern vectors are accumulated to form multiple unit accumulation sequences. This step implements the pattern indexing stage and the weighted synthesis stage. Using the second vector... The data vector along the length direction is used as the index, and based on each index, all pattern vectors are used... Select the corresponding specific mode vector, accumulate these specific mode vectors, generate one bit of intermediate result in each cycle, and continuously p x or p x +1 cycles form a unit cumulative sequence. The above operation is performing...
[0123] More specifically, according to the second vector The data vectors in the same position along the length direction make all pattern vectors The specific pattern vector in the vector passes through. Because of the second vector... The length of the second vector is N, therefore... It can be decomposed into N data vectors, each data vector having a bit width of p. y Therefore, these data vectors can be decomposed into p bits from the perspective of the same bit position. y A number of identical data vectors.
[0124] Next, determine which unit vector in the binary matrix K the input data vector in the same position is identical to, and output the specific pattern vector corresponding to the identical unit vector. This completes the process. The operation.
[0125] Finally, these specific pattern vectors are sequentially summed to obtain a unit summation sequence. It is particularly important to ensure the accuracy of the summation, that is, according to p... y The weighted vector C reflects the second vector The power of . The unit cumulative sequence is in Based on this, we further implement the weighted summation of C. Each unit of the cumulative sequence is like... Figure 4 The intermediate results 401, 402, 403, and 404 in the data have been aligned.
[0126] In step 1103, multiple unit summation sequences are added together to obtain the inner product result. To achieve simultaneous calculation, this embodiment uses the first vector... With the second vector After splitting the data into multiple segments, the intermediate results obtained by performing inner product calculations on each segment are as follows. Since the operations on the least significant bit and the most significant bit only have one intermediate result, the operations on the least significant bit and the most significant bit do not require addition, just like... Figure 4 The least significant bit (x0y0) and most significant bit (x7y3) in the result do not need to be added to other intermediate results and can be output directly. In other words, only the operations from the second least significant bit to the second most significant bit require addition.
[0127] This embodiment employs a design that simultaneously calculates both carry-free and carry-free results to reduce computational latency. The sum of intermediate results for both carry-free and carry-free results is obtained simultaneously. Then, based on whether the calculation result of the previous digit carries, the output is chosen to be either the sum of intermediate results with carry or the sum of intermediate results without carry. The accumulated output is the first vector. With the second vector The inner product result.
[0128] Back Figure 10In step 1004, the inner product result is integrated into the calculation result of multiple operands. The core controller controls the memory agent to integrate or reduce the inner product result into the calculation result of multiple operands and send it to the core memory agent.
[0129] In step 1005, the calculation results are stored in off-chip memory. The core memory agent sends the calculation results in parallel, first sending the least significant bit of these calculation results simultaneously, then sending the next least significant bit simultaneously, and so on until the most significant bit of these calculation results has been sent simultaneously.
[0130] Another embodiment of the present invention is a computer-readable storage medium storing computer program code for arbitrary-precision calculations. When the computer program code is run by a processor, it performs the following... Figure 10 or Figure 11 The method. In some implementation scenarios, the integrated unit described above can be implemented as a software program module. If implemented as a software program module and sold or used as an independent product, the integrated unit can be stored in a computer-readable storage memory. Based on this, when the solution of the present invention is embodied in the form of a software product (e.g., a computer-readable storage medium), the software product can be stored in a memory, which may include several instructions to cause a computer device (e.g., a personal computer, a server, or a network device, etc.) to execute some or all of the steps of the method described in the embodiments of the present invention. The aforementioned memory may include, but is not limited to, various media capable of storing program code, such as USB flash drives, flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0131] This invention proposes a novel architecture for efficiently handling arbitrary precision computations. Regardless of the precision of the operands, this invention can decompose the operands and use indexes to process fixed-length bit streams in parallel, avoiding bit-level redundancy such as sparsity or redundant computations. It achieves flexible application and large-bit-width computation without requiring high-bit-width hardware.
[0132] Depending on the application scenario, the electronic devices or apparatus of the present invention may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, IoT terminals, mobile terminals, mobile phones, dashcams, navigators, sensors, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, autonomous driving terminals, vehicles, home appliances, and / or medical devices. The vehicles include airplanes, ships, and / or vehicles; the home appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, lights, gas stoves, and range hoods; the medical devices include MRI scanners, ultrasound machines, and / or electrocardiographs. The electronic devices or apparatus of the present invention can also be applied in fields such as the Internet, IoT, data centers, energy, transportation, public management, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and healthcare. Furthermore, the electronic devices or apparatus of the present invention can also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing, such as cloud computing, edge computing, and terminal computing. In one or more embodiments, the high-computing-power electronic devices or apparatuses according to the present invention can be applied to cloud devices (e.g., cloud servers), while the low-power electronic devices or apparatuses can be applied to terminal devices and / or edge devices (e.g., smartphones or cameras). In one or more embodiments, the hardware information of the cloud devices and the hardware information of the terminal devices and / or edge devices are compatible with each other, so that suitable hardware resources can be matched from the hardware resources of the cloud devices to simulate the hardware resources of the terminal devices and / or edge devices based on the hardware information of the terminal devices and / or edge devices, so as to complete the unified management, scheduling and collaborative work of end-to-cloud or cloud-edge-end integration.
[0133] It should be noted that, for the sake of brevity, this invention describes some methods and their embodiments as a series of actions and combinations thereof. However, those skilled in the art will understand that the solution of this invention is not limited to the order of the described actions. Therefore, based on the disclosure or teachings of this invention, those skilled in the art will understand that some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art will understand that the embodiments described in this invention can be considered as optional embodiments, that is, the actions or modules involved are not necessarily essential for the implementation of one or more solutions of this invention. In addition, depending on the solution, the description of some embodiments of this invention also has different emphases. In view of this, those skilled in the art will understand that parts not described in detail in a certain embodiment of this invention can also refer to the relevant descriptions of other embodiments.
[0134] In terms of specific implementation, based on the disclosure and teachings of this invention, those skilled in the art will understand that the several embodiments disclosed herein can also be implemented in other ways not disclosed herein. For example, regarding the various units in the electronic device or device embodiments described above, this document has divided them based on logical functions, but in actual implementation, there may be other ways of division. As another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. Regarding the connection relationship between different units or components, the connection discussed above in conjunction with the accompanying drawings can be a direct or indirect coupling between units or components. In some scenarios, the aforementioned direct or indirect coupling involves a communication connection utilizing an interface, wherein the communication interface can support electrical, optical, acoustic, magnetic, or other forms of signal transmission.
[0135] In this invention, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same position or distributed across multiple network units. Furthermore, depending on actual needs, some or all of the units can be selected to achieve the purpose of the solution described in the embodiments of this invention. Additionally, in some scenarios, multiple units in the embodiments of this invention may be integrated into one unit or each unit may exist physically independently.
[0136] In other implementation scenarios, the integrated units described above can also be implemented in hardware, i.e., as specific hardware circuits, which may include digital circuits and / or analog circuits. The physical implementation of the circuit's hardware structure may include, but is not limited to, physical devices, which may include, but are not limited to, transistors or memristors. Therefore, the various devices described herein (e.g., computing devices or other processing devices) can be implemented using appropriate hardware processors, such as central processing units, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage units or storage devices can be any suitable storage medium (including magnetic storage media or magneto-optical storage media), such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), ROM, and RAM.
[0137] The embodiments of the present invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An arbitrary precision computing accelerator connected to an off-chip memory, the arbitrary precision computing accelerator comprising: a core memory agent configured to read a plurality of operands from the off-chip memory; a core controller configured to split the plurality of operands into a plurality of vectors, the plurality of vectors comprising a first vector and a second vector; and a processing array comprising a plurality of processing units configured to inner product the first vector and the second vector to obtain an inner product result according to a length of the first vector and the second vector, wherein each processing unit comprises: a conversion unit configured to generate a plurality of pattern vectors according to the length of the first vector and a bit width; a plurality of inner product units, each inner product unit being indexed based on a data vector of the second vector in the length direction, accumulating a particular pattern vector of the plurality of pattern vectors to form a unit accumulation series; and a synthesis unit configured to sum up a plurality of unit accumulation series to obtain the inner product result; wherein the core controller integrates the inner product result into a computation result of the plurality of operands, and the core memory agent stores the computation result to the off-chip memory; a data vector in the length direction of the second vector is identical to a unit vector of a binary matrix K, and a pattern vector corresponding to the unit vector of the binary matrix K is output from the conversion unit, and this pattern vector is referred to as a particular pattern vector.
2. The arbitrary precision computing accelerator of claim 1, wherein a starting address of the plurality of operands is set in the core memory agent, and the core memory agent serially reads the plurality of operands by increasing the address.
3. The arbitrary precision computing accelerator of claim 2, wherein the core memory agent reads the plurality of operands from low bits to high bits of the plurality of operands at a time.
4. The arbitrary precision computing accelerator of claim 1, wherein the core memory agent sends the computation result to the off-chip memory in parallel.
5. An integrated circuit device comprising: the arbitrary precision computing accelerator of any one of claims 1 to 4; a processing device configured to control the arbitrary precision computing accelerator; and an off-chip memory comprising an LLC; wherein the arbitrary precision computing accelerator and the processing device are connected through the LLC.
6. A board card comprising the integrated circuit device of claim 5.
7. An arbitrary precision computing method comprising: reading a plurality of operands from an off-chip memory; splitting the plurality of operands into a plurality of vectors, the plurality of vectors comprising a first vector and a second vector; and According to the length of the first vector and the second vector, the first vector and the second vector are inner multiplied to obtain an inner product result, wherein a plurality of mode vectors are generated according to the length of the first vector and a bit width, a specific mode vector in the plurality of mode vectors is accumulated based on a data vector in the length direction of the second vector to form a plurality of unit accumulation sequences, and the plurality of unit accumulation sequences are summed to obtain the inner product result; wherein the same bit data vector in the length direction of the second vector and the unit vector of the binary matrix K, the mode vector corresponding to the corresponding unit vector of the binary matrix K generated from the conversion unit is output as the specific mode vector; The inner product result is integrated into the calculation result of the plurality of operands; and The calculation result is stored to the off-chip memory.
Citation Information
Patent Citations
Operational accelerator
CN109213962A