Method and apparatus for managing weight data access of neural network processor
By introducing a DMA system and a scheduler-controlled neural processing unit into the neural network processor, the memory access operation of the weight matrix is optimized, solving the problems of memory bandwidth and processor resource utilization, and improving processor performance and efficiency.
Patent Information
- Application Number
- CN202480044932.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-21
- Filing Date
- 2024-05-30
- Publication Date
- 2026-01-30
AI Technical Summary
In the prior art, neural network processors consume a lot of power and memory bandwidth during memory access operations, resulting in performance degradation, and there is a lack of effective ways to optimize the use of memory bandwidth and processor resources.
The neural processing unit (NPU) employs a direct memory access (DMA) system and a scheduler and sequence processor (SSP) control, which reduces the frequency of external data access and improves processor utilization by efficiently managing memory access operations of the weight matrix.
It optimizes the use of memory bandwidth and processor resources, improves the performance and efficiency of neural network processors, and reduces memory access latency and power consumption.
Smart Images

Figure CN121444104A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to the field of digital processing circuits. In particular, but not exclusively, the invention discloses a digital circuit design, control system and mode of operation for managing on-chip and off-chip weight matrix data access of digital circuits performing neural network processing operations. BACKGROUND
[0002] Computer system designers are always trying to design faster computer systems. Faster computer systems allow extremely complex models of computation to be performed faster, such as weather forecasting, protein folding, astrodynamics, artificial intelligence, and complex three-dimensional video rendering. In addition, the models of computation being simulated can be made more detailed, thus presenting more accurate results.
[0003] To design faster computer systems, many different techniques are used. One of the simplest techniques is to increase the clock speed at which the computer system operates, although increasing clock speed is becoming much more difficult due to the physical properties of current transistor materials. Processing wider data structures can also improve computer performance, but this only helps for certain types of computational tasks that can utilize wider data structures. Two of the current popular techniques for improving processing speed are using parallel processing techniques, such as implementing multiple computational cores within a computer processor, or combining thousands of different computer systems on a network to collaborate on a single computational problem.
[0004] One of the areas where specialized processors are most needed to improve performance is the field of artificial intelligence (AI). Artificial intelligence is being used more and more for a variety of complex tasks, such as image recognition, high performance computing (HPC), scientific computing, machine learning, data mining, speech recognition, and self-driving cars. Artificial intelligence applications tend to rely heavily on matrix computations from the field of mathematics from linear algebra. Specifically, matrix operations are often needed to implement an artificial neural network (ANN) that learns from a set of training data, and then stores the learning in the form of neural network weight values. The neural network can then later apply the learning stored within the neural network weight values to new input data to make logical inferences on the new input data.
[0005] Due to the heavy use of matrix operations in neural networks, artificial intelligence is a computationally intensive field of computation that is in dire need of computational optimization. One of the most popular techniques to improve artificial intelligence application performance is to create specialized processing circuits for performing the matrix operations needed to implement a neural network. Specialized matrix processors take advantage of the inherent parallelism in matrix operations, and thus efficiently perform the matrix computations commonly used in artificial intelligence.
[0006] Artificial intelligence processing systems perform a large number of linear algebra matrix computations. The matrix computations performed by artificial intelligence systems are typically performed repeatedly using the same set of matrix weights but different input data vectors. Similarly, the data vectors can need to be processed through several neural network layers that require generation of many intermediate results through many matrix computations before a final output result is computed.
[0007] All of these complex matrix computations required by neural network based artificial intelligence applications involve moving large amounts of data from memory storage and then into and out of specialized neural network processing circuitry. In particular, neural network matrix processing operations require large weight matrices to be loaded into the matrix processing circuitry. The memory access operations for the large weight matrices required by neural networks can consume a large amount of power, consume a large amount of memory bandwidth, and cause latency. Without good coordination, all of these memory access operations for the weight matrices can slow down the performance of specialized neural network processors. Therefore, it is desirable to develop new techniques for organizing and scheduling the memory access operations for the weight matrices used within neural network processing in a manner that optimizes the use of memory bandwidth and processor resources. BRIEF DESCRIPTION OF DRAWINGS
[0008] In the drawings, which are not necessarily drawn to scale, like numerals describe substantially similar components throughout the several views. Like numerals having different letter suffixes
[0009] Figure 1A Illustrates a conceptual diagram of a single layer artificial neural network.
[0010] Figure 1B Illustrates a conceptual diagram of a three layer artificial neural network.
[0011] Figure 1C Illustrates a conceptual diagram of a two layer artificial neural network that does not have a dependency between each input data vector value and each output data vector value.
[0012] Figure 2A Illustrates a block diagram of one embodiment matrix processor circuit that can be used to perform matrix computations.
[0013] Figure 2B Illustrates Figure 2A a conceptual diagram of the matrix processor circuit of FIG. 1, where a 4x4 weight matrix composed of 16 weight value elements W[0,0] through W[3,3] is stored in an SRAM memory system.
[0014] Figure 2C Illustrates a block diagram of an abstracted matrix processor circuit that can be used to perform matrix computations.
[0015] Figure 2D This is a block diagram illustrating an abstract matrix processor circuit with bus interfaces on all sides.
[0016] Figure 3A This is a block diagram illustrating a matrix processor array, also known as a Neural Processing Unit (NPU).
[0017] Figure 3B illustrate Figure 3A An embodiment of a matrix processor array.
[0018] Figure 4 This is a timing diagram illustrating the typical processing of a neural network.
[0019] Figure 5A A nine-layer neural network is conceptually described, which receives a set of input data, processes the input data through nine neural network layers, and outputs a set of result data.
[0020] Figure 5B A simplified timing diagram illustrating how the weight matrix can be loaded into a neural processor.
[0021] Figure 5C illustrate Figure 5B The timing diagram, where work segments allow simultaneous operation of work from multiple different neural network layers.
[0022] Figure 5D This is a timing diagram illustrating a method for loading a weight matrix into a neural processor while maximizing neural processor utilization.
[0023] Figure 5E This section presents a timing diagram illustrating a method for loading a weight matrix into a neural processor while maximizing neural processor utilization and minimizing memory usage.
[0024] Figure 6 Explain how the weight matrix can be pre-extracted into the timing diagram of the neural network processor. Detailed Implementation
[0025] The following detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings illustrate illustrations based on exemplary embodiments. These embodiments (also referred to herein as "exemplaries") have been described in sufficient detail to enable those skilled in the art to practice the invention. It will be apparent to those skilled in the art that the specific details of the exemplary embodiments are not necessary for practicing the invention. For example, although some exemplary embodiments are disclosed with reference to specific matrix processor circuit embodiments, the disclosed techniques can be used in conjunction with any other embodiments of matrix processor circuits. Exemplary embodiments may be combined, other embodiments may be utilized, or structural, logical, and electrical changes may be made without departing from the scope of the claims. Therefore, the following detailed description should not be considered limiting, and the scope is defined by the appended claims and their equivalents.
[0026] In this document, the term "a (or an)" is used as commonly found in patent documents to include one or more. In this document, the term "or" is used to refer to a non-exclusive "or," such that "A or B" includes "A but not B," "B but not A," and "A and B," unless otherwise indicated. Furthermore, all publications, patents, and patent documents referenced in this document are incorporated herein by reference in their entirety, as if individually incorporated. In the event of any inconsistency between this document and those documents incorporated by reference, the use in the incorporated reference shall be considered supplementary to this document; in the case of irreconcilable inconsistency, the use in this document shall prevail.
[0027] Neural network overview
[0028] One of the core technologies of artificial intelligence (AI) is the use of artificial neural networks (ANNs). Artificial neural networks first learn from training data and then later use it to perform logical reasoning from new input data. Artificial neural networks were initially designed to resemble biological neural networks in the animal brain.
[0029] Figure 1A This is a conceptual diagram illustrating a very simple single-layer, four-input artificial neural network (100). (Reference) Figure 1A During the training session, training data (consisting of inputs 101 to 104) is provided to the input data vector, and new input data is then provided when performing inference using the artificial neural network. The input data vector (consisting of inputs 101 to 104) is processed with weights in weight matrix 120 to create an output data vector (consisting of outputs 161 to 164). Many different types of data processing can be performed using weight matrix 120 (e.g., Hadamard product, Frobenius inner product, matrix addition, etc.), however, this document will focus on the well-known matrix product.
[0030] After processing the input data vector (consisting of inputs 101 to 104) with weighting matrix 120, the system creates an output data vector (consisting of outputs 161 to 164). The output data vector (consisting of outputs 161 to 164) can be combined with output function 170 to create the final output 191 of the artificial neural network 100. Output function 170 may be referred to as an activation function. During a training session, the output data can be compared with a desired target output (not shown), and the difference between the output data and the desired target output can be used to adjust the weight data within weighting matrix 120 to improve the accuracy of the artificial neural network 100.
[0031] It should be noted that Figure 1A The four-input artificial neural network is just one example of a simple small artificial neural network (ANN). Artificial neural networks can be constructed to be much wider than just four inputs. Multiple independent artificial neural networks can be used in parallel, and the outputs of independent artificial neural networks can be combined.
[0032] Artificial neural networks can include many layers of weight matrices, enabling them to perform very complex computational analyses of input data. For example, Figure 1B A three-layer artificial neural network is described, wherein the input data vector (consisting of inputs 101 to 104) is processed by a first weighting matrix 121 to create a first intermediate data vector (consisting of data 141 to 144). Next, the first intermediate data vector (consisting of data 141 to 144) is processed by a second weighting matrix 122 to create a second intermediate data vector (consisting of data 151 to 154). Then, the second intermediate data vector (consisting of data 151 to 154) is processed by a third weighting matrix 123 to create an output data vector (consisting of outputs 161 to 164). The output data vector (consisting of outputs 161 to 164) can then be processed by an output function 170 to create a final output 191. Alternatively (or otherwise), the output data vector (consisting of outputs 161 to 164) can also be used as intermediate data, which is fed into additional layers of the artificial neural network (not shown), allowing the creation of very complex hierarchical artificial neural networks.
[0033] Parallel operation
[0034] It is well known that due to the use of matrices, the operations to be performed have a high degree of parallelism within each layer, which can be utilized to enable parallel execution of operations. Specifically, matrix multiplication requires many independent multiplication operations that can be executed in parallel. However, artificial neural networks can also contain inherent parallelism that can be utilized between different layers of the artificial neural network.
[0035] Figure 1CA two-layer neural network is described, which receives an input tensor 100, processes the input tensor 100 with a first weight matrix 121 in a first layer to create an intermediate tensor 140, and then processes the intermediate tensor 140 with a second weight matrix 122 in a second layer to create an output tensor 150. However, in Figure 1C In the two-layer neural network, not every input data value in the input tensor 100 affects any intermediate value in the intermediate tensor 140. Similarly, not every intermediate data value in the intermediate tensor 140 affects any output value in the output tensor 150. Therefore, some operations in the neural network layers can be performed before all input values from the input tensors are ready.
[0036] For example, intermediate value 141 depends only on input data values 101 and 102. Similarly, intermediate value 142 depends only on input data value 101. Therefore, intermediate values 141 and 142 can be computed before input values 103 and 104 are available, and thus can be computed in parallel with the computation required to compute input values 103 and 104. Furthermore, output value 151 depends only on intermediate values 141 and 142. Therefore, the computation of output value 151 can be performed simultaneously with the computation required to compute input values 103 and 104.
[0037] In some embodiments, a grouping system is used to create individual work segments such that each individual work segment can be distributed as soon as the required input data becomes available. In this way, individual work segments can be created for the computations required to create intermediate values 141, 142, and output value 151. Those work segments can be executed as long as input values 101 and 102 are available, and in parallel with the computations required to create input values 103 and 104. An example of a grouping system is disclosed in U.S. Patent Application No. 17 / 970,450, filed October 20, 2022, entitled “METHOD AND APPARATUS FOR USING A PACKET ARCHITECTURE TO PROCESS NEURAL NETWORKS IN ANEURAL PROCESSING UNIT,” which is incorporated herein by reference. The teachings of this document are ideally implemented in such systems to best utilize the parallelism inherent in the neural network being processed.
[0038] Example matrix processor circuit
[0039] For reference Figure 1A , 1BAs explained in 1C, artificial intelligence relies on a large number of computationally intensive matrix operations to initially learn and adjust the weights in the weight matrix using training data. Later, those adjusted weight matrices are used to perform complex matrix operations on a new set of input data to make inferences about the new input. Fortunately, the linear algebraic matrix operations used in artificial neural networks allow for many performance optimizations because of the high degree of parallelism in the required matrix operation tasks. Therefore, many dedicated processors have been created for artificial intelligence applications. These dedicated AI processors can use a single-instruction multiple-data (SIMD) architecture, where a wide data vector is processed with each instruction, making matrix operations efficient.
[0040] To provide optimal processing for artificial intelligence tasks, a dedicated matrix processor can be used. A matrix processor is a digital processing circuit designed to help perform AI computational tasks efficiently. Specifically, a matrix processor is designed to rapidly read input data vectors, output data vectors, and matrix weight data in a parallel format to achieve high throughput. In this way, a matrix processor can be used for forward propagation inference and backward propagation AI learning.
[0041] Figure 2A A block diagram illustrating one embodiment of the example matrix processor circuit 200 is provided. It should be noted that the matrix processor circuit can process data vectors having many more or fewer data elements in each data vector.
[0042] Figure 2A The matrix processor circuit 200 includes a local wide static random access memory (SRAM) bank 230. The SRAM bank 230 can be configured to access data across an entire wide row in a single memory cycle. In this way, an entire input data vector or a whole row of weight values from a weight matrix can be read from or written to the SRAM bank 230 in a single memory cycle. The matrix processor circuit 200 also includes an operand register file 210 for storing the input data vector and other data vectors that can be used as operands during computation.
[0043] A wide SRAM bank 230, an operand register file 210, and an operand bus 221 are coupled to a set of multiplexers 240, which provide operand data to a set of multiply-accumulate (MAC) units 260. A local control system 205 within the matrix processor circuitry 200 controls all these individual circuit elements to perform the required data vector processing operations. Therefore, the local control system 205 selects between data stored in the wide SRAM 230, data in the operand register file 210, and data in the operand bus 221 to provide to the multiply-accumulate (MAC) units 260 for data vector processing.
[0044] The computational outputs from the set of multiplication and accumulation (MAC) units 260 can be stored in the result register file 250. These outputs can be output in parallel in their raw form using the result bus 291. Alternatively (or in addition to the raw output data), the results in the result register file 250 can be combined with the reduction tree 270 to provide a single output on the reduction bus 295. It should be noted that the reduction tree 270 can be implemented externally to the matrix processor circuitry 200.
[0045] It should be noted that for some operations, the results stored in the result register file 250 can be used as input operands in subsequent data vector calculations. To handle such calculations, there exists a data path from the result register file 250 back to the multiplication and accumulation (MAC) unit 260. The local control system 205 is used to precisely control how the multiplication and accumulation (MAC) unit 260 selects the input data to be processed and how the multiplication and accumulation (MAC) unit 260 processes the input data.
[0046] Figure 2B This conceptually illustrates how a 4×4 weight matrix consisting of elements W[0,0] to W[3,3] can be stored within a wide SRAM group 230. It should be noted that the weight values in the weight matrix are stored aligned with the row structure of the underlying SRAM 230 memory, allowing an entire row of weight values to be read from the wide SRAM group 230 in a single memory cycle. For example, the weight values W[0,0], W[0,1], W[0,2], and W[0,3] from the first row of the weight matrix can be read from the wide SRAM group 230 in a single memory cycle and simultaneously provided in parallel to individual multiplication and accumulation (MAC) units in MAC group 260. Other input operands used for computation can be obtained in parallel from operand register file 210 or from the operand bus (…). Figure 2B (Not shown in the text) Go to MAC group 260.
[0047] Figure 2A and 2B The matrix processor circuit 200 illustrates only one possible embodiment of the matrix processor circuit. Figure 2A and 2B Details of the matrix processor circuit 200 can be found in U.S. Patent Application No. 16 / 149,054, entitled "Methods and Apparatus for Constructing Digital Circuits for Performing Matrix Operations," which is incorporated herein by reference. However, the matrix processor circuit can be implemented in many different ways and in many different sizes.
[0048] Abstracted matrix processor circuit
[0049] Matrix processor circuits can be implemented in many different sizes and in many different ways. However, to further process matrix operations efficiently, multiple matrix processor circuits can be combined in an efficient manner, enabling the controlled network of the matrix processor circuits to perform a wide variety of matrix operations. Therefore, for the sake of simplicity in this disclosure, reference will be made to... Figure 2C The abstract matrix processor circuit is publicly disclosed.
[0050] Figure 2C This illustrates the first block diagram of the abstract processor circuit 201. The abstract matrix processor circuit 201 receives input data on one or more operand buses. Figure 2C In a particular embodiment, there are two operand buses: an operand bus 221T from the top and an operand bus 221L from the left. Data received on the operand buses can be used directly by the processing logic 267 or stored in the memory bank 230 for later use. The received data may include weight matrix data, input data operand vectors, or other data. The memory bank 230 may also contain a register file tightly coupled to the processing logic 267. It should be noted that the operand and result buses can be placed on all different sides of the matrix processing array circuitry. Figure 2D This describes an example with operand buses (221T, 221B, 221L, and 221R) and result buses (291T, 291B, 291L, and 291R) on all sides.
[0051] Return to reference Figure 2C The matrix processor circuit 201 also receives commands on the command bus 207. The control system 205 within the matrix processor circuit 201 parses the received commands and uses them to determine how the processing logic 267 will be used to process the data. The processing logic 267 can be implemented in many different ways, as long as the matrix processor 201 performs the desired matrix operations and outputs the correct matrix operation results. For example, the processing logic 267 can be used as follows: Figure 2A and 2B The single instruction multiple data (SIMD) processor, digital signal processor (DSP), conventional central processing unit (CPU) core, highly parallelized dedicated matrix processor circuit 200, or any other implementation that performs the desired matrix operations, as described herein.
[0052] The abstract matrix processor circuit 201 can be designed to operate using many different types of data formats and levels of precision. For example, the abstract matrix processor circuit 201 can handle integers, 16-bit floating-point numbers, 32-bit floating-point numbers, or any other data format. Many different matrix operations can be implemented in the abstract matrix processor circuit 201. Two well-known matrix operations that can be included are the matrix dot product and the matrix cross product.
[0053] The control system 205 instructs processing logic 267 to output the result of the requested matrix operation on one or more result buses 291. In some embodiments, the matrix processor 205 includes reduction logic that outputs the reduced form of the result on the reduction bus 295. As will be described later, the reduction logic may also be implemented outside the abstract matrix processor circuitry 201.
[0054] Operand buses 221T and 221L can be parallel buses, allowing the entire input data vector to be loaded into the abstract matrix processor circuit 201 in a single loop. Similarly, the entire row of the weight matrix from the weight matrix can be read into the memory bank 230 of the abstract matrix processor circuit 201 in a single loop. Likewise, result buses 291R and 291B can be parallel buses, allowing the entire output data vector to be output from the abstract matrix processor circuit 201 in a single loop.
[0055] As previously explained, the memory bank 230 of the abstract matrix processor circuit 201 can be wide and deep to optimize performance. The memory bank 230 can be wide because the entire data vector can be written to or read from the memory bank 230 in a single loop. For example, in a large matrix processor circuit 201 that processes a 16×16 element matrix where each element is a 16-bit floating-point value, the memory bank 230 can be configured to read 256-bit values such that the entire 16-element data vector of 16-bit data values can each be read from the memory bank 230 in a single loop.
[0056] The memory bank 230 can be constructed large enough to store multiple sets of different weight matrices. In this way, the matrix processor circuit 201 can be used to perform matrix operations on multiple different artificial neural network layers without reloading different matrix weight values. For example, if the matrix processor circuit 201 cannot perform an operation on a particular neural network layer because the required input data vector is not yet available, it can instead perform matrix operations on other neural network layers or other neural networks. And as previously mentioned, the operations on neural network layers can be divided into individual working segments, such that some segments of later layers can be executed before working segments from earlier layers, provided the required input data is available. The deep memory bank 230 allows for very efficient use of the matrix processor 201 because it can handle a steady stream of requested matrix operations from many different neural networks without loading new weight matrix data. Loading weight matrix data can be one of the most time-consuming (and energy-intensive) tasks of the matrix processor circuit 201.
[0057] In addition to storing the weight values of multiple different weight matrices, memory bank 230 can also be used to store other information that may be needed, such as input data vectors, output data vectors, error vectors, etc. Intermediate result data vectors from the forward propagation operation can be stored in memory bank 230 and then accessed later when performing the relevant backpropagation operation. Another very important type of data that can be stored in memory bank 230 is the matrix weight gradient. The matrix weight gradient includes an adjustment matrix of the weight matrix that can be periodically used to update the weight matrix.
[0058] Combining matrix processors into an array
[0059] Figure 2C The abstract matrix processor circuit 201 described herein can be used alone to perform simple matrix operations very quickly. For example, the matrix processor circuit 201 can be used to fully process... Figure 1A The very small artificial neural network described herein. It can also be used to implement [the system] by sequentially performing the required matrix calculations for all three artificial neural network layers 121, 122, and 123. Figure 1B The small three-layer artificial neural network described in the paper.
[0060] However, most artificial neural networks must handle more than Figure 1A and 1B The example described herein uses a very small artificial neural network with much larger data input and output vectors. Therefore, it is desirable to combine the computational power of many different matrix processor circuits 201 to process wider and multi-layered artificial neural networks. In this way, much larger multi-layered artificial neural networks for performing useful artificial intelligence tasks can be processed very efficiently.
[0061] Figure 3A A block diagram illustrating a first embodiment of an architecture that uses multiple matrix processor circuits in a coordinated manner to efficiently process wide, multi-layered artificial neural networks. Figure 3A In this context, each individual matrix processor circuit is labeled as the "MP" of the matrix processor. For example... Figure 3A As illustrated in the example, the matrix processor circuitry is arranged in a grid array format. Between the individual matrix processor circuits in the matrix processor array is bus routing and combinational logic 399, which couples all individual matrix processor circuits in the array to buffers carrying input data vectors, matrix weight values, and result data vectors. The matrix processor array and the bus routing and combinational logic 399 may be referred to as matrix processor array 397. The bus routing and combinational logic 399 can be implemented in different ways to achieve different objectives.
[0062] In one embodiment, to provide input data vectors to the matrix processor array 397, the vector scalar processor (VSP) 371 is coupled to the operand bus of each individual matrix processor circuit in the array bus wiring 399. This can be achieved by coupling the operand bus 221L, as follows: Figure 2C As explained in the documentation, data vectors can be loaded into individual matrix processor circuits within the array in this manner. Data vectors may include weight matrix rows, input data vectors, or any other data required for the processing operations to be performed. It should be noted that these data vector loading operations can be performed in parallel due to the presence of multiple buses.
[0063] Similarly, the result bus of each individual matrix processor circuit in the array is coupled to an accumulation buffer (Acc buffer) 375 on the bottom of the matrix processor array 397 using bus routing and combinational logic 399. This can be achieved by... Figure 2C The result bus 291B is coupled to the bottom accumulator buffer 375 (Acc buffer) to achieve this, such as Figure 3A As described in the documentation, the accumulator buffer 375 contains both a storage device for storing the result data vector and processing logic for performing various vector processing operations on the received result data vector. For example, the accumulator buffer 375 can combine partial result data vectors from multiple different matrix processor circuits into a single complete output data vector result.
[0064] Each individual matrix processor circuit in the matrix processor array 397 receives commands on its individual command bus. In this way, each individual matrix processor circuit in the array can be individually controlled. For example, an individual matrix processor circuit can be notified when input data is available on its operand bus and what operation will be performed. Therefore, different parts of the same neural network layer or segments of work from different network layers can be executed simultaneously by different matrix processor circuits. By carefully controlling each individual matrix processor circuit in the matrix processor array 397 in a coordinated manner, the matrix processor array 397 becomes a very powerful system for efficiently processing many segments of matrix operations in neural network applications in parallel. Specifically, the matrix processor array 397, together with all supporting circuitry (accumulator buffer 375, vector scalar processor (VSP) 371, direct memory access (DMA) unit 381, etc.), can be referred to as a neural processing unit (NPU) 300.
[0065] Neural processing unit external data management overview
[0066] In order to Figure 3A The neural network is processed within the matrix processor array 397, and the neural processing unit (NPU) 300 must receive data from an external source. For example... Figure 3A As described, a neural processing unit (NPU) 300 has an input / output (I / O) interface 395 controlled by direct memory access (DMA) 381, which can receive and send data from outside the NPU 300. For most efficient operation, the NPU 300 must utilize its various internal memories and external data sources coupled to the I / O interface 395 as efficiently as possible. Ideally, Figure 3A The Neural Processing Unit (NPU) 300 typically stores the data required for performing matrix operations as frequently as possible in its internal local memory, allowing processing to occur uninterrupted. When accessing external data sources via the Input / Output (I / O) interface 395, such external access operations should be performed as infrequently and efficiently as possible. Therefore, the Neural Processing Unit (NPU) 300 should attempt to retain critical data that will be reused in its internal memory until it is no longer needed.
[0067] Neural processing unit overview
[0068] refer to Figure 3AThe Neural Processing Unit (NPU) 300 is controlled by a scheduler and sequence processor (SSP) 350. The scheduler and sequence processor (SSP) 350 is responsible for sending all loop commands to the Vector Scalar Processor (VSP) 371, all individual matrix processor (MP) circuits within the matrix processor array 397, and the accumulator buffer 375. (Connections are not shown for clarity.)
[0069] The scheduler and sequence processor (SSP) 350 may include a tree traverser (TW) 351 and a row sorter (RS) 353. The tree traverser (TW) 351 traverses the neural network tree and is responsible for obtaining the data slices required for processing. The row sorter (RS) 353 can process more than one row at a time, combining multiple rows into a single row sequence. The row sorter (RS) 353 is responsible for implementing all loop commands for each data slice. Each vector scalar processor (VSP) 371 follows the received loop commands in each operation loop. The same applies to the matrix processors, accumulator buffers 375, and direct memory access (DMA) units 381 within the matrix processor array 397. Each resource needs to be carefully ordered for proper operation by the neural processing unit (NPU) 300. Therefore, a set of loop commands needs to be generated for each operation loop.
[0070] The Direct Memory Access (DMA) unit 381 can issue multiple requests in parallel. This document will use the term Direct Memory Access (DMA) unit, which can be used to access external DRAM memory. However, DMA unit 381 should be considered a general-purpose memory access system that can be used to access any type of memory system (DRAM, SRAM, flash memory) with any type of memory interface (serial memory bus, network access, parallel bus, etc.). Additional information about DMA unit 381 will be provided in later sections.
[0071] The computer system can utilize numerous neural processing units 300 within an artificial intelligence computer system. Different neural processing units can be controlled collaboratively on the same neural network problem. (Alternatively, a single neural processing unit can be partitioned into multiple regions and simultaneously handle completely different computational problems within the same neural processing unit.) When several different neural processing units collaborate on the same computational problem, a DMA system 381 can be used to transfer data from one neural processing unit to another. This allows different neural processing units to solve different stages or layers of the same neural network computational problem.
[0072] Matrix processor array 397 is responsible for processing convolutional and fully connected (FC) neural network layers. Matrix processor array 397 can also perform group-by-group convolution. Matrix processor array 397 can perform two-dimensional reuse. The first dimension is input data broadcasting, where data is broadcast from vector scalar processor (VSP) 371. Partial summing units may exist within bus wiring and combinational logic 399, which can combine data values en route to accumulation buffer 375.
[0073] The accumulator buffer 375 is responsible for accumulating results from various matrix processors in the matrix processor array 397. The accumulator buffer 375 can also perform quantization. The accumulator buffer 375 can also execute activation functions for neural networks. For example, ReLU (Modified Linear Unit), PReLU (Paragon ReLU), Drain ReLU, and other well-known activation functions can be executed within the accumulator buffer circuit system 375. (Some activation functions can also be executed in the Vector Scalar Processor (VSP) 371.)
[0074] The Vector Scalar Processor (VSP) 371 is responsible for operations not performed within the matrix processor array 397 or the accumulator buffer 375. The VSP 371 can execute pooling functions, such as max pooling and average pooling functions. The VSP 371 can also execute data shaping functions. For example, data can be changed from one format to another to increase or decrease data precision. The VSP 371 has its own dedicated memory 372, which is described as the memory block to the left of the VSP 371.
[0075] Each matrix processor (MP) in the matrix processor array 397 contains its own local memory (as shown in the diagram). Figure 2D (Local memory group 230 in the matrix processor). This local matrix processor memory is primarily used to describe neural network weights. (It should be noted that matrix weights can also be loaded on the fly from the vector scalar processor (VSP) 371, but to reduce power consumption and memory port bandwidth usage, matrix weights are typically stored in local matrix process memory.) This local storage of matrix weights allows for the reuse of those matrix weights as much as possible. The local matrix processor memory can also be used to store temporary results.
[0076] Neural processing unit data access
[0077] Figure 3A and 3BThe neural processing unit 300 uses a direct memory access (DMA) 381 coupled to the input / output port 395 to send and receive data from outside the neural processing unit 300. It should be noted that this document will specifically refer to the use of "direct memory access." However, the term direct memory access in this document can be applied to any movement of data outside the neural processing unit 300. For example, the DMA unit 381 may access memory on different integrated circuit dies, or may access memory from different partitioned regions of the same integrated circuit chip.
[0078] In one embodiment, the Direct Memory Access System (DMA) 381 will be used to access a high-speed memory system, such as an external Double Data Rate (DDR) memory system. However, all disclosed techniques relating to access to data outside the Neural Processing Unit 300 are applicable to any type of interface system used for accessing data and for accessing any type of data storage system. For example, the external interface system 395 of the DMA unit 381 may include a Fast Peripheral Component Interconnect (PCIe) bus to access data. Similarly, the external interface system 395 may include a Mobile Industry Processor Interface (MIPI) data bus. Another type of data bus for the interface is the Advanced Scalable Interface (AXI) data bus, which is typically used in conjunction with ARM (Advanced RISC Machine) based processor systems.
[0079] Beyond computer data bus systems, the technology can also be used to access any computer network system to access external data. For example, a direct memory access (DMA) system 381 can access a well-known Ethernet interface to access data available on a computer network. Alternatively, a DMA system can use on-chip data structures to access data stored on the same chip. For example, as previously explained, a chip may contain multiple different neural processing units 300 that work together to perform neural network processing tasks. Therefore, the direct memory access (DMA) system 381 can use on-chip data structures to transfer data between different neural processing units 300 on the same chip.
[0080] The Direct Memory Access (DMA) system 381 can operate in conjunction with all different types of master and slave interface systems. For a master-type interface, the master device controls the interface. Therefore, for a master-type interface, the DMA system 381 can initiate data transfer at any time. For a slave-type interface, the DMA system 381 can only initiate data transfer when the slave interface receives permission from the master device of the interface system. The slave device can send an "indication status" message to the master device, informing the master device whether the slave device can currently receive or send data. Then, the master device of the interface system can notify the slave device when the slave device can send data, so that the DMA system 381 on the slave interface can respond with data transfer.
[0081] External interface systems based on the master device typically require less data buffering capacity because the master device has the ability to initiate data transfer when necessary. Slave interfaces may require more data buffering because the slave device lacks control over the data interface and therefore can only transfer data when the slave device receives permission from the interface master device to transfer the data.
[0082] The Direct Memory Access (DMA) system 381 can perform data compression and decompression to reduce the amount of memory used when storing data and the amount of memory bandwidth used when loading or storing data on the input / output port 395 coupled to external memory. This can be very useful when processing potentially very large neural network weight matrices. Therefore, when the DMA system 381 needs to store the weight matrix, it will first compress the weight matrix, thereby saving bandwidth on the input / output port 395 and storage space in external memory (not shown). When the weight matrix is subsequently retrieved, the DMA system 381 will decompress the compressed weight matrix before use.
[0083] Conventional management overview
[0084] For simple convolutional neural networks, current neural network processing systems can be operated in a very simple and straightforward manner. To describe the operation of a typical neural network processing system, reference will be made to... Figure 4 Presentation processing Figure 1B The example of a three-layer neural network is described in the text.
[0085] refer to Figure 4 The first step is to load all the data required for the desired computation into the neural processing unit. This is in Figure 4 The above describes the input data and weight matrix data 410 being loaded onto the external interface. When sufficient data has been loaded into the neural processing unit's memory to begin computation, the system can begin processing tensor data. This is described as the system processing layer 1 431. During operation, the neural processing unit can store intermediate results 415 and reload them into external memory. The neural processing system then processes all convolutional neural network layers, such as layers 2 432 and 3 433, until the entire three-layer neural network is fully processed and has a final output. When the output data is available, the neural processing system can then begin sending out the output data, such as... Figure 4 The output data 470 is described as being sent out on the external interface.
[0086] This traditional system has several problems. One of the biggest drawbacks is its significant waste of the input / output bandwidth available on the external I / O interface. Specifically, the system first uses I / O bandwidth to load input data 410, but then largely allows the external interface's I / O bandwidth to remain idle while the data is processed through all neural network layers, except for potentially storing intermediate results 415. After all matrix operations in layers 1 431, 2 432, and 3 433, the system then uses the external interface to send out the final result data 470. Therefore, in such systems, I / O bandwidth may need to be excessively supplied to expensive high-speed interfaces for rapid loading of source data and unloading of result data to minimize latency. Furthermore, the high-speed external interface does nothing most of the time.
[0087] Figure 4 Another drawback of the technique described is that it is a very high latency solution. To be specific, Figure 4 The technique described herein makes almost no use of the natural parallelism inherent in computational problems. A small amount of parallelism exists because processing in layer 1431 can begin before all input data 410 has been received, and the resulting output data 470 can begin to be sent out before processing in the final layer (layer 3 433) is complete. However, apart from this, Figure 4 All processing operations performed in the system are done serially, resulting in high latency.
[0088] Figure 4 A final additional problem with this traditional solution is that such a traditional system requires a large amount of memory within the neural processor unit to store all input data, all matrix weight data, intermediate results created during neural processing, and the final output data. Therefore, improvements are expected to enhance performance and increase resource utilization efficiency.
[0089] Matrix processor array with improved external data access
[0090] To improve upon traditional systems, this disclosure employs data loading and storage techniques that more efficiently utilize the external storage interface of the neural processing system. Specifically, the proposed system makes fuller use of the external storage interface bandwidth when loading the weight matrix. This reduces neural network processing latency, improves resource utilization, and saves power.
[0091] When performing neural network processing, three main types of data are loaded and stored: matrix weight data, input data, and output data. In a typical neural network, intermediate result data can be both output data from the current neural network layer and input data from subsequent neural network layers. This disclosure presents techniques for improving the data transfer efficiency of matrix weight data.
[0092] Reducing weight matrix data into a neural processor
[0093] The weight matrix data represents a large amount of data that must be loaded into the neural processor to perform neural network computations. One technique used by the disclosed system is to load more than one set of weight matrices, allowing the neural processor to focus on processing the same layer, several different neural network layers, or even several different neural network segments without needing to reload the weight matrix data.
[0094] Figure 5 conceptually illustrates a nine-layer neural network that receives a set of input data 501, processes the input data 501 through nine neural network layers (511 to 519), and outputs a set of result data 509. To reduce the loading and reloading of weight matrix data, the system may load several sets of weight matrices and process several layers at a time. In the example of Figure 5, the nine-layer neural network has been divided into three "partitions" of three neural network layers: partition A 505, partition B 506, and partition C 507. Each partition of the neural network layer can be processed as a group. Therefore, a work segment from any of the layers in a partition can be processed, provided that the required input data is available for that work segment.
[0095] For example, to process partition A 505, the neural processor loads weight matrices 521, 522, and 523 to process neural network layers 1 511, 2 512, and 3 513, respectively. In this way, the neural processor can process working segments of neural network layers 1 511, 2 512, and 3 513 by loading only those weight matrices once. It should be noted that in embodiments using working segments, working segments from any of those three neural network layers can be processed, provided the required input data for those working segments is available. Therefore, working segments do not need to be processed in the traditional neural network processing order. After the system has completed processing partition A 505, the system can then load the weight matrices (524, 525, and 526) for the next partition of the neural network layers, namely partition B 506, which consists of neural network layers 4 514, 5 515, and 6 516.
[0096] Figure 5B A simplified timing diagram illustrating how to load the subsequent weight matrix in a system that processes neural network layers one at a time instead of using working fragments. (Reference) Figure 5BOnce the final layer 3 operation 580 is completed, the neural processor can then begin loading the weight matrices of the three layers of partition B506 using weight matrices 524, 525, and 526 respectively via loading operations 574, 575, and 576. Once the layer 4 weights 524 are loaded, the system can begin layer 4 operations in stage 581. In this embodiment without working segments, all layer 4 operations in stage 581 must be completed before the system begins layer 5 operations in stage 582. Figure 5B The method described may not be efficient in using the computational circuitry system because the computation available for work is limited to the single layer of the neural network being processed. For example, during stage 581, the system may only operate on layer 4 operations.
[0097] To make more efficient use of the available computing circuitry, the use of working segments allows the system to operate on operations from multiple different layers simultaneously. Figure 5C This describes a timing diagram for a system that allows simultaneous processing of any work segments from the same partition. (Reference) Figure 5C Once the final layer 3 operation 583 is completed, the neural processor can begin loading the weight matrices of the three layers of partition B 506 using weight matrices 524, 525, and 526 respectively via loading operations 571, 572, and 573. Once layer 4 weights 524 are loaded in stage 571, the system can begin processing layer 4 work segments in stage 584. However, immediately following the loading of layer 5 weights 525 in stage 572, the system can then begin processing work segments for both layers 4 and 5 in stage 585. Similarly, immediately following the loading of layer 6 weights 526 in stage 573, the system can then begin processing work segments for all layers (layers 4, 5, and 6) of partition B 506 in stage 586. In this way, the neural processor can utilize the computational circuitry more efficiently because more different work segments can be selected for processing in stage 586. Therefore, the neural processor is less likely to stall due to a lack of work segments ready for computation.
[0098] refer to Figure 5B and 5C Before loading the weight matrix for the next partition, both systems wait until the previous partition of the layer is complete. Therefore, this method introduces some latency into the neural network computation process. Specifically, this allows the neural processor to idle, such as... Figure 5B and 5C This is illustrated in the neural processor's computation timeline.
[0099] To operate more efficiently, this disclosure proposes preparing for the next partition of a neural network layer before the current partition has been processed. (See references.) Figure 5DBefore the neural processor completes the operations of layers 1 to 3 of partition A 505, it begins loading the weight matrix for the next partition of the neural network layer. Specifically, the neural processor's external interface performs a loading operation 534 to load the layer 4 weights 524 of partition B 506 before partition A 505 is completed. In this way, after the layer 4 weights 524 have been loaded into the neural processor, the neural processor can then immediately begin processing the work segments from layer 4 in stage 544. It should be noted that during stage 544, the system can continue processing the work segments from partition A 505 (layers 1, 2, and 3).
[0100] Then, after load operation 534, there can be load operation 535 for layer 5 weights 525 and load operation 536 for layer 6 weights 526. In this way, the neural processor can begin processing the remaining working segments in partition B 506. Specifically, after load operation 535 for layer 5 weights 525 is completed, the neural processor can then process any remaining working segments from partition A 505, as well as working segments from layers 4 and 5, during stage 545. Then, after load operation 536 for layer 6 weights 526, the neural processor can then process all working segments of all partitions B 506 (layers 4, 5, and 6) during stage 546. It should be noted that any working segments remaining from partition A 505 may still be in progress (not shown).
[0101] like Figure 5D As can be seen, the neural processing timeline demonstrates that the neural processor is not idle because the weight matrices for the next partition are reduced into the neural processor before they are needed. In this way, the neural processor always has work segments available for processing, greatly improving its utilization.
[0102] Reducing weight matrix data out of a neural processor
[0103] In addition to maximizing neural processor utilization and efficiently using the interface with external memory by reducing the weight matrix, the system can also minimize memory usage and maximize utilization by reducing the weight matrix from the neural processor. (As described in the previous paragraphs...) Figure 5D In this technique, the neural processor loads a new set of weight matrices, while the existing set of weight matrices remains in the neural processor's memory, thus exhausting the neural processor's precious memory resources. To improve this, Figure 5E One implementation is described, in which the neural processor reduces the weight matrix by discarding some weight matrices from the first partition before loading a new matrix for the next partition.
[0104] refer to Figure 5EThe timing diagram shows that once the neural processor completes its last layer 1 operation 561 from partition A 505, it begins preparing for the next partition of the layer by discarding 551 layer 1 weights 521, as those weights are no longer needed. This frees up memory space for the load operation 574 to load layer 4 weights 524. During this period, the neural processor continues processing the working segments of layers 2 and 3 of partition A 505 in stage 592.
[0105] As layer 2 512 is completed in stage 562, the system also discards layer 2 weights 522 in stage 552. During this time, the system may complete the work segment from layer 3 in stage 593. However, once layer 4 weights 524 are loaded, the neural processor can then process work segments from layer 3 513 in partition A 505 and layer 4 514 in partition B 506 in stage 594. It should be noted that the neural processor is able to process work segments from two different partitions (partition A 505 and partition B 506) because it reduces the weight matrix from the first partition and reduces it into the weight matrix from the next partition. After discarding layer 2 weights 522 in stage 552, the system loads layer 5 weights 525 in stage 575. It should be noted that memory can be saved by not loading new weights until previously used weights are discarded.
[0106] Approximately simultaneously, the system can complete the working segment of layer 3 at stage 563, and therefore discard layer 3 weights 523 at stage 553. Then, the system can begin loading the layer 6 weights 526 at stage 576. It should be noted that after discarding layer 3 weights 523 at stage 553, the system will focus on the working segment of layer 4 514 in partition B 506 at stage 595. However, once the layer 5 515 weights 525 become available after the loading operation 575, the neural processor can then process the working segments of layers 4 and 5 at stage 596. Finally, when the layer 6 weights 526 in partition B 506 become available after the loading operation 576, the neural processor can then process any of the working segments (layers 4, 5, and 6) from partition B 506 at stage 597, thus completing a smooth transition from partition A 505 to partition B 506 without causing the neural processor to stall or become idle, as described by... Figure 5E The neural processor's processing timeline at the bottom illustrates this.
[0107] Instead of spending time discarding previously needed weight matrices, the neural processor can further optimize its operation by immediately loading the new weight matrix into the memory location previously occupied by the weight matrix that was no longer needed. In this highly optimized embodiment, no additional memory is required at all, because the neural processor shrinks the new weight matrix into the memory location of the previous weight matrix. Specifically, the newly loaded matrix is placed in the same memory location as the discarded matrix.
[0108] Neural processors may have context switching, where they switch between different neural network processing jobs. During such context switching, the techniques disclosed herein for reducing weight matrices in and out of the neural processor can be used to maximize utilization during such context switching.
[0109] Pre-fetching weight matrix data into a neural processor
[0110] Return to reference Figure 5D and 5E Loading the weight matrix data for partitions of a neural network layer can consume significant bandwidth on the external memory interface. If the neural processor always waits until switching to a new partition before loading the weight matrix required for that partition, there may be too much traffic on the external memory interface, causing the required weight matrices to not be loaded before they are needed. This can lead to the neural processor stalling or becoming idle. To prevent this, the weight matrices are prefetched into the neural processor whenever bandwidth becomes available on the external memory interface. This will consume some memory in the neural processor, but it helps ensure that the neural processor does not stall or leave its matrix processor idle.
[0111] Figure 6 This sequence diagram illustrates how the weight matrix can be pre-extracted into the neural network processor. Initially, the neural network processor processes layers 1, 2, and 3 in stage 691. During this processing, the external interface is used for many different types of load and store operations. However, the neural network processor can also issue load operations for layers of weight matrices. When available bandwidth exists on the external interface, the system issues a load layer 4 weights operation 674 to load layer 4 weights 524. These weights can be stored in the neural processor until they are needed.
[0112] When the neural processor completes layer 1 of the first partition, it can decide to start using the layer 4 weights. Therefore, after the neural processor completes layer 1 in stage 661, it can switch from processing layers 1, 2, and 3 during stage 691 to processing layers 2, 3, and 4 during stage 692.
[0113] Throughout this period, load and store operations will occur on the external interface. When idle bandwidth exists on the external interface, the DMA unit can issue a load operation 675 to load layer 5 weights 525. Again, these weights will be stored until the neural processor decides to begin processing the work segment from layer 5. Again, this can be triggered by completing the previous layer. Therefore, in stage 662, the neural processor completes the layer 2 work segment, and then begins processing layers 3, 4, and 5 during stage 692.
[0114] like Figure 6 As illustrated in the timing diagram, prefetching allows for the opportunistic execution of memory-bandwidth-intensive tasks that load large weight matrices. For example, weight matrix prefetching can be performed on a lower priority basis. In this way, external memory bandwidth usage can be optimized.
[0115] The foregoing disclosure is intended to be illustrative and not restrictive. For example, the above embodiments (or one or more aspects thereof) may be used in combination with each other. Other embodiments will be apparent to those skilled in the art upon reading the above description. Therefore, the scope of the claims should be determined by reference to the full scope of the appended claims and the equivalents to which such claims are given. In the appended claims, the terms “comprising” and “in which” are used as common English equivalents to the corresponding terms “including” and “wherein”. Furthermore, in the following claims, the terms “comprising” and “including” are open-ended, meaning that a system, apparatus, article, or process that comprises elements other than those listed after such terms in the claims is still considered to be within the scope of the claims. Additionally, in the following claims, the terms “first,” “second,” and “third,” etc., are used merely as labels and are not intended to impose numerical requirements on their objects.
[0116] An abstract is provided to comply with 37 CFR §1.72(b), which requires that the abstract allow the reader to quickly determine the nature of the technical disclosure. The understanding of submitting an abstract is that it will not be used to interpret or limit the scope or meaning of the claims. Furthermore, in the detailed description above, various features may be grouped together to simplify this disclosure. This should not be construed as meaning that any unclaimed disclosed feature is essential to any claim. Rather, the subject matter of the invention may lie in fewer than all features of a particular disclosed embodiment. Therefore, the following claims are incorporated herein by reference, wherein each claim stands independently as a separate embodiment.
Claims
1. A method for processing a multi-layer neural network with a neural network processor by reducing in weight matrix data from a memory coupled to the neural network processor, the method comprising the steps of: partitioning the multi-layer neural network into subsets of neural network layers, wherein each subset is to be processed as a group; each of the subsets of neural network layers is referred to as a partition; partitioning each neural network layer of each of the partitions into a group of work segments, each work segment comprising a subset of operations of the neural network layer; grouping the group of work segments of each partition into work segment subsets that can be processed concurrently; loading a first work segment subset of a first partition from the memory into the neural network processor; loading a first subset of weight matrix data from the external memory for the first work segment subset of the first partition into the neural network processor; starting processing of the first work segment subset when the first subset of weight matrix data is available; while processing the first work segment subset of the first partition, loading a second subset of weight matrix data from the external memory for a second work segment subset of the first partition into the neural network processor if not already loaded; loading the second work segment subset of the first partition from the external memory into the neural network processor; and processing the second work segment subset of the first partition when the second subset of weight matrix data for the second work segment subset is available.
2. The method for processing a multi-layer neural network with a neural network processor and managing access to memory of claim 1, wherein the work segment subsets contain work segments from different neural network layers in the first partition.
3. The method for processing a multi-layer neural network with a neural network processor and managing access to memory of claim 1, wherein work segments can be processed out of order such that later neural network layers can be processed before previous neural network layers.
4. The method for processing a multi-layer neural network with a neural network processor and managing access to memory of claim 1, further comprising: decompressing the first subset of weight matrix data loaded from the external memory.
5. The method for processing a multi-layer neural network with a neural network processor and managing access to memory of claim 1, wherein the first subset of weight matrix data can comprise one of a number of different data precisions.
6. The method for processing a multi-layer neural network with a neural network processor and managing access to memory of claim 1, further comprising: reloading the first subset of weight matrix data from the external memory after a context switch of the neural network processor.
7. A method for processing a multi-layer neural network with a neural network processor and reducing out weight matrices from the neural network processor, the method comprising the steps of: partitioning the multi-layer neural network into subsets of neural network layers, wherein each subset is to be processed as a group; The subset of neural network layers is referred to as a partition; each network layer in each partition is divided into a set of work segments; each partition's set of the work segments is grouped into work segment subsets that can be processed concurrently; a first work segment subset of a first partition is loaded from the external memory into the neural network processor; a first weight matrix is loaded from the external memory for the first partition's set of work segment layers; processing of the work segments of the first partition's neural network layers is initiated; after processing a final work segment of the first network layer, the first weight matrix of the first neural network segment of the first partition is discarded to free up memory resources; and while processing of the work segments of the first partition is completed, a second weight matrix of a neural network layer in a subsequent partition is loaded into the neural network processor.
8. The method of processing a multi-layer neural network with a neural network processor and reducing weight matrix data from the neural network processor of claim 7, the method further comprising: decompressing the first weight matrix loaded from the external memory.
9. The method of processing a multi-layer neural network with a neural network processor and managing access to memory of claim 7, wherein the first weight matrix can include one of several different data precisions.
10. The method of processing a multi-layer neural network with a neural network processor and managing access to memory of claim 7, further comprising: after a context switch of the neural network processor, reloading the first weight matrix from the external memory.
11. The method of processing a multi-layer neural network with a neural network processor and managing access to memory of claim 7, wherein the partitions can belong to different neural networks.
12. A method of processing a multi-layer neural network with a neural network processor and pre-fetching weight matrix data from an external memory, the method comprising the steps of: dividing the multi-layer neural network into subsets of neural network layers, wherein each subset will be processed as a group; the subset of neural network layers is referred to as a partition; each network layer in each partition is divided into a set of work segments; each partition's set of the work segments is grouped into work segment subsets that can be processed concurrently; a first work segment subset of a first partition is loaded from the external memory into the neural network processor; a first weight matrix is loaded from the external memory for the first partition's set of work segment layers; processing of the work segments of the first partition's neural network layers is initiated; while memory bandwidth is available to the external memory, a second weight matrix of a neural network layer in a subsequent partition is pre-fetched from the external memory into the neural network processor while processing the work segments of the first partition; and the second weight matrix is stored in the neural network processor until the subsequent partition is triggered and the second weight matrix is needed for processing.
13. The method of processing a multi-layer neural network with a neural network processor and prefetching weight matrix data from external memory of claim 12, wherein the prefetching is performed with lower priority than other accesses to the external memory.
14. The method of processing a multi-layer neural network with a neural network processor and prefetching weight matrix data from external memory of claim 12, wherein the partitions can belong to different neural networks.
15. The method of processing a multi-layer neural network with a neural network processor and prefetching weight matrix data from external memory of claim 12, wherein the work segments can be executed out of order.
16. The method of processing a multi-layer neural network with a neural network processor and prefetching weight matrix data from external memory of claim 12, wherein the partitions can belong to different neural networks.
17. The method of processing a multi-layer neural network with a neural network processor and spooling weight matrix data out of the neural network processor of claim 12, the method further comprising: decompressing the second weight matrix that is prefetched from the external memory.
Citation Information
Patent Citations
Methods and apparatus for constructing digital circuits for performing matrix operations
US11983616B2
Method and apparatus for using a packet architecture to process neural networks in a neural processing unit
US20240152761A1