Vector operation acceleration with convolution computation units

By designing a hybrid convolution-vector computing processing system, the parallel execution of convolution and vector computing is realized by using hardware resources to reuse and control signal configuration, the problems of inefficiency and high cost in the prior art are solved, and the efficiency of neural network computing is improved.

CN120226017APending Publication Date: 2025-06-27MOZI INT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380066171.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-14
Filing Date
2023-09-05
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When handling convolution and vector operations in neural network computing, existing hardware architectures are inefficient and have high cost and design complexity, especially the parallel processing efficiency of vector operations is not high.

Method used

A hybrid convolution-vector computing processing system is designed to realize parallel execution of convolutional operations and vector computing by reusing hardware resources, using multiple weight selectors, activation input interfaces and multiplier accumulation (MAC) circuits. The MAC channel is configured to perform addition or subtraction operations through control signals, supporting convolutional calculations, vector reduction and pooling operations.

Benefits of technology

It improves the efficiency of neural network computing, reduces hardware cost and design complexity, and realizes efficient parallel processing of convolution and vector operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226017A_ABST
    Figure CN120226017A_ABST
Patent Text Reader

Abstract

Hybrid hardware accelerators, systems, and apparatuses are described for performing various computations in neural network applications using the same set of hardware resources. An example accelerator may include a weight selector, an activation input interface, and a plurality of multiplier accumulation (MAC) circuits organized into a plurality of MAC channels. Each of the plurality of MAC channels may be configured to: receive a control signal to indicate whether to perform a convolution operation or a vector operation; receiving one or more weights according to the control signal; receiving one or more activations according to the control signal; and generating output data based on the one or more weights and the one or more input activations according to the control signal, and feeding the output data into an output buffer. Each of the plurality of MAC channels includes a plurality of multiplier circuits and a plurality of addition and subtraction circuits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to a hardware design for improving the computational efficiency of neural networks, and more particularly to a hybrid convolution-vector operation PE (processing entity) cluster for processing neural network computations such as convolutions and vector operations. Background Art

[0002] Neural network (NN) computations involve convolution computations and various vector operations such as vector reduction operations (e.g., reduce max / min / sum / average or reduce index) or pooling operations (e.g., max pooling, average pooling). In addition to processing units for handling convolution computations, existing hardware architectures rely on SIMD (single instruction multiple data) vector processors or dedicated hardware (e.g., tensor processing units (TPUs) or other ASIC designs) to implement these vector operations. Vector processors are capable of processing one vector at a time, but are inefficient at processing multiple vectors in parallel. Additionally, prior to performing reduction operations, SIMD processors may require special hardware to perform operations across multiple elements within the same vector for pooling operations or preliminary operations (e.g., transpose operations). Further, the requirement to install a separate hardware vector processor to handle vector operations in neural network computations may increase the cost and design complexity of neural network processing units. Summary of the Invention

[0003] Various embodiments of this specification may include hardware accelerators, PE clusters, and systems for using the same set of hardware to process convolution computations and vector operations.

[0004] In some aspects, the techniques described herein relate to a vector operation accelerator for neural network computations. The accelerator may include a plurality of weight selectors configured to obtain weights, a plurality of activation input interfaces configured to obtain activations, and a plurality of multiplier-accumulator (MAC) circuits organized into a plurality of MAC channels. In some embodiments, each of the plurality of MAC channels may be configured to: receive a control signal indicating whether to perform a convolution operation or a vector operation; receive one or more weights from at least one of the plurality of weight selectors according to the control signal; receive one or more activations from at least one of the plurality of activation input interfaces according to the control signal; and generate output data based on the one or more weights and the one or more input activations according to the control signal and feed the output data into an output buffer, wherein: each of the plurality of MAC channels includes a plurality of first circuits for performing multiplication operations, and a plurality of second circuits for performing addition or subtraction operations according to the control signal.

[0005] In some aspects, the plurality of second circuits within a MAC channel are organized as a tree, and the second circuits at the leaf level of the tree are configured to receive data from the plurality of first circuits.

[0006] In some aspects, each of the plurality of second circuits is configured to: receive a first input and a second input; determine whether to perform addition or subtraction based on a control signal; generate a sum or average of the first input and the second input in response to the control signal indicating addition; and generate a minimum or maximum value between the first input and the second input in response to the control signal indicating subtraction.

[0007] In some aspects, each of the first input and the second input includes a vector having the same number of dimensions, and in order to generate the minimum value between the first input and the second input, each of the plurality of second circuits is further configured to: generate an output vector that includes the minimum value of the vectors on each corresponding dimension.

[0008] In some aspects, the accelerator may further include: a weight matrix generation circuit configured to generate weights for vector reduction operations, where the vector reduction operations include one or more of reduction mean, reduction minimum, reduction maximum, reduction average, reduction addition, or pooling.

[0009] In some aspects, each of the plurality of weight selectors includes a multiplexer coupled to the weight matrix generation circuit and a weight cache.

[0010] In some aspects, each of the plurality of weight selectors is configured to: obtain weights from the weight cache in response to a control signal indicating execution of a convolution calculation; and obtain weights from the weight matrix generation circuit in response to a control signal indicating execution of a vector calculation.

[0011] In some aspects, the accelerator may further include an adder-subtractor circuit external to the tree corresponding to the MAC channels, where the adder-subtractor circuit is configured to receive data from the second circuit at the root level of the MAC channels and write the data to an output buffer.

[0012] In some aspects, the adder-subtractor circuit is further configured to: during a first calculation iteration, write a first set of data received from the second circuit at the root level of the MAC channels to the output buffer; and during a second calculation iteration: receive a temporary data set from the second circuit at the root level of the MAC channels, retrieve the first set of data from the output buffer, calculate a second set of data based on the temporary data set, the first set of data, and a control signal indicating whether to perform a convolution calculation or a vector operation, and write the second set of data to the output buffer.

[0013] In some aspects, the plurality of MAC channels are configured to respectively receive a plurality of weight vectors generated by the weight matrix generation circuit to perform a plurality of vector operations in parallel.

[0014] In some aspects, a first subset of the plurality of MAC channels is configured to receive weights from a weight cache, a second subset of the plurality of MAC channels is configured to receive weights generated by a weight matrix generation circuit, and the first subset of the plurality of MAC channels is further configured to perform convolution calculations, the second subset of the plurality of MAC channels is further configured to perform vector operations, and the convolution calculations and the vector operations are performed in parallel.

[0015] In some aspects, the techniques described herein relate to a hybrid convolution-vector operation processing system. The system may include: a plurality of weight selectors configured to obtain weights; a plurality of activation input interfaces configured to obtain activations; and a plurality of multiplier-accumulator (MAC) circuits organized into a plurality of MAC channels. Each of the plurality of MAC channels is configured to: receive a control signal indicating whether to perform a convolution operation or a vector operation; receive one or more weights from at least one of the plurality of weight selectors according to the control signal; receive one or more activations from at least one of the plurality of activation input interfaces according to the control signal; and generate output data based on the one or more weights and the one or more input activations according to the control signal and feed the output data into an output buffer, wherein: each of the plurality of MAC channels includes a plurality of first circuits for performing multiplication operations and a plurality of second circuits for performing addition operations or subtraction operations according to the control signal.

[0016] After considering the following description and the appended claims in conjunction with the accompanying drawings, these and other features of the systems, methods, and non-transitory computer-readable media disclosed herein, as well as the methods of operation and functions of the related elements of the structures, combinations of components, and manufacturing economies, will become more apparent, all of the drawings forming a part of this specification, wherein like reference numerals represent corresponding parts in each of the drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended as a definition of the limitations of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 An exemplary system diagram for processing neural network computations in a hybrid PE array according to various embodiments is shown.

[0018] Figure 2A An exemplary architecture diagram of a PE array according to various embodiments is shown.

[0019] Figure 2B An exemplary architecture diagram of a hybrid PE array for processing convolution calculations and vector operations according to various embodiments is shown.

[0020] Figure 3 An exemplary internal structure diagram of a MAC (multiplier-accumulator) channel in a hybrid PE array according to various embodiments is shown.

[0021] Figure 4 An exemplary logic circuit design of accumulators and subtractors in each MAC channel according to various embodiments is shown.

[0022] Figure 5A An exemplary use case of performing vector reduction using a hybrid PE array according to various embodiments is shown.

[0023] Figure 5B Another exemplary use case of performing multiple vector reductions in parallel using a hybrid PE array according to various embodiments is shown.

[0024] Figure 5C Another exemplary use case of performing pooling using a hybrid PE array according to various embodiments is shown.

[0025] Figure 6 An exemplary system design of a hybrid PE array according to various embodiments is shown.

[0026] Figure 7 An example computer system in which any of the embodiments described herein can be implemented is shown. Detailed Description

[0027] The embodiments described herein provide hardware devices, systems, and PE (processing element) arrays that have the ability to perform various neural network computations in parallel by reusing hardware resources. Here, neural network computations involve convolutional computations and vector operations such as vector reduction operations (e.g., reducing the maximum / minimum / sum / average of values in a vector, or reducing the maximum / minimum of vector indices) and pooling operations (e.g., max pooling, average pooling), which almost constitute all the computations involved in neural network training and applications. In some embodiments, the hardware devices described herein include a PE array and other auxiliary logic circuits, and are capable of processing different types of neural network computations by reusing the same hardware resources. For simplicity, in the following designs, the hardware device may be referred to as a hybrid PE array. In some embodiments, the auxiliary logic circuits may include a weight matrix generation circuit configured to generate function weights (compared with the weights in a neural network).

[0028] While existing NN hardware designs use SIMD vector processors or separate / standalone hardware (such as a TPU) to process vector operations in addition to configuring processors for performing convolutional computations, the hybrid PE array described herein re-uses the same hardware resources, such as multiplier-accumulators (MACs), for performing convolutional operations while providing scalable parallel processing capabilities for vector operations. In particular, the MAC resources can be configured via control signals to perform summation or comparison, where the summation function can be triggered to implement convolutional computations, vector sum reduction (e.g., adding corresponding values of multiple input vectors and generating an output vector with the sum), vector mean reduction (e.g., adding the corresponding values of multiple input vectors and then dividing the sum by the number of input vectors to generate an output vector), or average pooling (e.g., computing the average of each patch of a feature map), and the comparison function can be triggered to implement vector maximum reduction, vector minimum reduction, or max pooling (e.g., finding the maximum value within individual patches of a feature map), etc.

[0029] In the following description, specific, non-limiting embodiments of the present invention will be described with reference to the accompanying drawings. Specific features and aspects of any embodiment disclosed herein can be used and / or combined with specific features and aspects of other embodiments disclosed herein. It should also be understood that these embodiments are provided by way of example only and illustrate only a few embodiments within the scope of the present invention. Various changes and modifications that are obvious to those skilled in the art to which the present invention pertains are considered to be within the spirit, scope, and contemplation of the present invention further defined in the appended claims.

[0030] Figure 1 An exemplary system diagram for processing neural network computations in a hybrid PE array 160 according to various embodiments is shown. Figure 1 The figure shows a hardware architecture that can be configured to perform common neural network computations, such as convolutional computations and vector operations, using the same hardware resources (such as the hybrid PE array 160). The embodiments described in this disclosure can be implemented as Figure 1 part of a neural network computation in or other suitable environments.

[0031] Typical neural networks, such as convolutional neural networks (CNNs), may involve various computations, such as convolutional operations and vector operations. For example, a convolutional layer within a neural network (e.g., a CNN) can typically perform a convolution based on one or more input feature maps (IFMs) (including activations) obtained from an input source (e.g., such as an input image) or a previous layer (e.g., a tensor output from a previous layer), and one or more weight tensors corresponding to a given layer (e.g., from a weight source 130, such as a weight cache or a weight generator). The weight tensors can be used to convolve the IFMs to extract various features (e.g., convolutional computations). The convolution process can be performed in parallel in the hybrid PE array 160. Each PE can refer to a processor with processing capabilities and storage capacity (e.g., a buffer or a cache). In some embodiments, each PE can include one or more logic gates or circuits configured as multipliers and accumulators (MACs). The PEs in the hybrid PE array 160 can be interconnected by wires and arranged in multiple channels called PE channels or MAC channels. The hybrid PE array 160 can be fabricated as a neural network accelerator or a data processing system.

[0032] As another example, a pooling layer in a CNN can provide a method for downsampling an IFM by summarizing the presence of features in patches of the feature map. Common pooling methods, such as average pooling and max pooling, can summarize the average presence of features and the most active presence of features, respectively. These pooling operations may involve vector operations different from convolutional computations. Specifically, convolutional computations include convolving a window through an IFM and performing multiplication and accumulation to extract features, while vector operations may include comparison or subtraction. Other types of vector operations are also common in neural networks, such as vector reduction operations. The vectors to be processed can be obtained from a vector memory (an input source 150), which can be the same source that obtains convolutional activations. In some embodiments, the MACs in the hybrid PE array 160 can be configured to switch between a "convolution mode" and a "vector mode" to perform convolutional computations and vector operations, serving the convolutional layer and the pooling layer, respectively.

[0033] The configuration of the hybrid PE array 160 can be based on control signals or instructions issued by the instruction decoder 110. The instruction decoder 110 can decode instructions from upper-layer processors such as a CPU or a GPU. The instruction decoder 110 can send corresponding control signals to different components based on the decoded instructions. For example, the instruction decoder 110 can send an input activation / vector load control signal to the input source 150, indicating whether to obtain an activation or a vector for the hybrid PE array 160. The instruction decoder 110 can send a weight load control signal to the weight source 130, which indicates whether the weights should be obtained from the weight cache or from the weight matrix generator ( Figure 2B(see more details in). Different weights play a crucial role in neural network computations. For example, the weights fetched from the weight cache may refer to the weights from the filters corresponding to the convolutional layer, which are configured to extract features from the IFM. The weights generated from the weight matrix generator can be designed to perform a specific type of vector operation. That is, depending on whether the current operation is a convolutional operation or a vector operation, the source of the weights may be different. As another example, if the weights have a certain pattern, the instruction decoder 110 can instruct the weight matrix generator to generate weights so as to generate weights internally instead of fetching weights from the weight cache. This can minimize external memory access, thereby improving the overall performance of neural network computations.

[0034] The instruction decoder 110 can also send a computation control signal to the hybrid PE array 160 to perform the required operations based on the input from the input source 150 and the weights from the weight source 130. The control signal can configure the MACs in the hybrid PE array 160 to perform summation (e.g., for convolution) or subtraction (e.g., for comparison in vector operations). After the hybrid PE array 160 completes the computation according to the control signal, the output data can be fed into the output buffer 170 for storing temporary activations or vectors. In some embodiments, the instruction decoder 110 can also control the output buffer to feed the data from the previous iteration back to the hybrid PE array 160 to participate in the current iteration.

[0035] Figure 2A An exemplary architecture diagram of a PE array according to various embodiments is shown. Figure 2A The arrangement of PEs in the PE array in is for illustrative purposes and can be implemented in other ways according to the use case.

[0036] As Figure 2AAs shown in the left part of FIG. 0, the PE array 200 may include a PE matrix. Each PE may include a plurality of multipliers (MUL gates). The multipliers within each PE may work in parallel, and the PEs within the PE array 200 may work in parallel. For ease of reference, in the following description, the number of columns 220 of PEs in the PE array 200 is denoted as X, the number of rows 210 of PEs in the PE array 100 is denoted as Y2, and the number of multipliers within each PE is denoted as Y1. Each row 210 of PEs may be referred to as a PE cluster, and each PE cluster may be coupled to Y1 adder trees 230 for aggregating the partial sums generated by the multipliers within the PE cluster. That is, the first multiplier in each PE within the PE cluster is coupled to the first adder tree 230 for aggregation, the second multiplier in each PE in the PE cluster is coupled to the second adder tree 230 for aggregation, and so on. The aggregation results of the adder trees 230 of all PE clusters (a total of Y1×Y2 adder trees) may be fed into the adder 250 for aggregation. The adder 250 may refer to a digital circuit that performs digital addition, and this digital circuit is part of the network-on-chip (NoC) subsystem.

[0037] Figure 2B FIG. 4 shows an exemplary architecture diagram of a hybrid PE array for processing convolution calculations and vector operations according to various embodiments. The example diagram includes a plurality of hardware components that interact with each other, such as a weight matrix generator 270, a plurality of weight selectors 272 configured to receive weights from the weight matrix generator 270 or a weight cache ( Figure 2B not shown in FIG. 6), a plurality of MAC (multiplier-accumulator circuit) channels 273 corresponding to the plurality of weight selectors 272 respectively, and a plurality of MACs 274 for receiving data from the plurality of MAC channels 273 respectively and outputting the data to an output buffer ( Figure 2B not shown in FIG. 8). These components are for illustrative purposes only, and according to the implementation, the architecture may include more, fewer, or alternative components. For example, the plurality of MACs 274 may be implemented as the last MAC in the plurality of MAC channels but with a different configuration from the other MACs in the MAC channels. The weight selector 272 may be implemented using a multiplexer coupled to the weight matrix generator 270 and the weight cache.

[0038] Figure 2B All the components shown in FIG. 12 are through instructions from a decoder (such as Figure 1is directly or indirectly controlled by the control signals of the instruction decoder 110). The control signals may include weight loading control and computation control for configuring the hybrid PE array to perform convolution computations or vector operations, or both, using multiple MAC channels 273. For example, if the upper application layer indicates to perform a convolution computation, the corresponding loading control may be sent to the multiple weight selectors 272 and they may be notified to select weights from the weight cache. The weights from the weight cache may be from the filters of the convolutional layer in the neural network and may be used to perform convolution through the IFM to extract features. If the upper application layer indicates to perform a vector operation, such as vector reduction (e.g., vector mean, which combines multiple vectors into one vector with the mean), the weight loading control may be sent to the weight matrix generator 270 to generate weights, and the weight selectors 272 may be configured to select the generated weights and block weights from the weight cache. These generated weights may be functionally different from the weights for convolution computations. For example, the generated weights may be used to calculate the average of the input vectors (e.g., for X vectors, each vector may be assigned a weight of 1 / X). In some embodiments, if the weights involved are highly organized (following a specific pattern), the weight matrix generator 270 may also be triggered to perform convolution computations. The weight matrix generator 270 may be instructed to generate these weights according to a specific pattern of the MAC channels. In this way, the MAC channels may avoid accessing the weight cache to obtain weights, thus avoiding memory access latency. The weight selectors 272 may then forward the received weights to the MAC channels for corresponding computations.

[0039] Similarly, multiple MAC channels 273 can be controlled by computational control signals. The computational control signals can instruct the MAC channels to receive input activations for convolutional computations or vectors for other vector operations. In some embodiments, in parallel, some MAC channels (also referred to as a first subset of MAC channels) can be configured for convolutional computations, while other MAC channels (also referred to as a second subset of MAC channels) can be configured for vector operations. This means that the hybrid PE array can perform convolutional computations and vector operations simultaneously (e.g., during the same iteration). The computational control signals can also configure the MACs in each channel according to the specific workload to be executed. For example, for convolutional computations on a MAC channel, as part of the convolutional computation, the multiplier in the MAC channel can perform multiplication operations, and the accumulator within the MAC channel can perform summation operations. For vector reduction, such as vector max reduction or max pooling in a neural network pooling layer, the multiplier in the MAC channel can perform multiplication (using weights generated by a weight matrix generator), and the accumulator of the MAC channel can be configured to perform subtraction to achieve a comparison function. The comparison may help determine the maximum value of each dimension of the input vector (for vector max reduction) or the maximum value within each patch of the feature map (for max pooling). In some embodiments, the accumulator in the MAC channel is designed to be configurable to perform summation or subtraction. A more detailed circuit design of the hybrid accumulator can be found in Figure 4 found.

[0040] In some embodiments, the last - layer MACs 274 can act as a bridge between the MAC channels and the output buffers. For example, the MACs 274 can receive data (e.g., partial sums, temporary activations, temporary vector outputs) from the MAC channels and save the data to the corresponding output buffers. In some cases, after the MACs 274 store the data in the output buffers during the first iteration, they can also read data from the output buffers and use it together with the new data received from the MAC channels 273 as part of the next - iteration computation.

[0041] Figure 3 An exemplary internal structure diagram of a MAC (multiplier - accumulator) channel 300 in a hybrid PE array according to various embodiments is shown. Figure 3The figures in [the text] are for illustrative purposes only and, according to the embodiments, may include more, fewer, or alternative components. As described above, each MAC channel 300 may include a plurality of multipliers 310 and accumulators, which may be implemented as digital gates or circuits (e.g., an arithmetic logic unit (ALU)). According to specific calculation instructions from the upper layer, the accumulator can not only perform summation but also subtraction. In this way, the same piece of hardware (MAC channel 300) can be reused for different types of calculations without an additional special processor (e.g., a vector processor, TPU). In this application, the accumulator may also be referred to as an adder-subtractor circuit 320 because they have a hybrid function: summation (addition) and comparison (subtraction).

[0042] In some embodiments, the adder-subtractor circuits 320 in the MAC channel 300 may be organized as an adder tree 330. The leaf-level adder-subtractor circuits of the adder tree 330 may be configured to receive data from a plurality of multipliers 310. For example, each multiplier 310 may perform multiplication (for convolution or vector operations) based on an input vector / activation and weights. The multiplication results from two or more multipliers 310 may be fed into the leaf-level adder-subtractor circuits of the adder tree 330 to perform summation or subtraction (comparison) according to a control signal. For example, each adder-subtractor circuit 320 may be configured to receive a first input and a second input; determine whether to perform addition or subtraction based on a control signal; generate the sum or average of the first input and the second input in response to a control signal indicating addition; and generate the minimum or maximum value between the first input and the second input in response to a control signal indicating subtraction. If the control signal indicates performing a vector minimum reduction, each of the first input and the second input may include vectors having the same number of dimensions, and the adder-subtractor circuit may generate an output vector including the minimum values of the vectors for each corresponding dimension.

[0043] In some embodiments, the adder tree 330 may include a plurality of add-subtracter circuits 320 at the leaf level, one add-subtracter circuit 340 at the root level, and one or more intermediate levels. From one level to the next level, the number of add-subtracter circuits is reduced by half. In some embodiments, the add-subtracter circuit 340 at the root level may obtain the calculation result of the adder tree 330 (e.g., the sum, minimum value, or maximum value of values or indices), and send the result to an add-subtracter circuit 350 outside the adder tree 330 corresponding to the MAC channel. Except that the external add-subtracter circuit 350 can write to and read from the output buffer, the external add-subtracter circuit 350 may be similar to the add-subtracter circuits within the adder tree 330. In some embodiments, the external add-subtracter circuit 350 and the add-subtracter circuit 340 at the root level may be the same circuit and are part of the adder tree 330.

[0044] In some embodiments, during a first calculation iteration, the external add-subtracter circuit 350 may be configured to write a first set of data received from the add-subtracter circuit at the root level to the output buffer. During a second calculation iteration, the external add-subtracter circuit 350 may be configured to receive a second set of temporary data from the root-level add-subtracter circuit 35 again, retrieve the first set of data from the output buffer, calculate a second set of data based on the second set of temporary data, the first set of data, and a control signal indicating whether a convolution calculation or a vector operation is being performed, and write the second set of data to the output buffer.

[0045] Figure 4 An exemplary logic circuit design 400 of the accumulator and subtracter in each MAC channel according to various embodiments is shown. Design 400 shows one way to implement a hybrid circuit that can be configured to perform addition or subtraction. According to the implementation, the hybrid circuit can be implemented using other digital components.

[0046] As shown, the example circuit includes a plurality of multiplexers 410 and 440 for selecting signals from a plurality of inputs based on a control signal. The control signal may include a first signal 420 indicating whether addition or subtraction is being performed, and a second signal 430 indicating whether the minimum value or the maximum value is obtained from the input value (if subtraction is being performed). These signals can control the selection logic of the multiplexers 410 and 440 to select appropriate inputs.

[0047] Figure 5A An exemplary use case of performing vector reduction using a hybrid PE array according to various embodiments is shown. Although the hybrid PE array is used for convolution calculations in the convolutional layer of a neural network, it can also be used for other types of calculations, such as vector reduction that requires different logical processing. Figure 5AThe use case shown in [description] involves vector reduction on the channel dimension of the input tensor 520 that generates a single output vector 530. The single output vector 530 can include rows of vectors, denoted as V0(:,0) to V0(:,z-1), where z is the channel dimension index. The weight matrix 510 for vector reduction can be generated by a weight generation circuit. As shown, the weight matrix 510 can include a first row of 1s and all other elements 0, while the input tensor 520 includes multiple vectors. Note that even though Figure 5A the operator in [description] is a multiplication operator, it represents a vector operator rather than matrix multiplication. The vector operator can be defined as part of the calculation control signal, which can include vector sum reduction, vector minimum, or reduction, etc. For simplicity, the vector reduction operator can be represented as the "reduce0" function. Using the "reduce0" function, each column vector in the input tensor 520 can be used as an argument for the "reduce0" function. For example, the first vector V0(:,0) of the single output vector can be represented as reduce0(A(0,0), A(1,0), …, A(y-1,0)), where y represents one of the row or column dimensions of the input tensor 520, and "reduce0" can be sum, minimum, maximum, mean, etc. This example shows a simple application of the hybrid PE array in channel dimension vector reduction.

[0048] Figure 5B Another exemplary use case is shown for parallel execution of multiple vector reductions using a hybrid PE array according to various embodiments. Figure 5B The example in [description] shows that due to the rich dynamics in the weight matrix 540, the hybrid PE array has high flexibility. According to the requirements of the upper-layer application, the weight matrix generation circuit can generate multiple weight rows in the weight matrix 540 to implement parallel vector reduction operations. Two weight rows can include different weights to implement two different vector reductions. For example, row 0 of the weight matrix 540 that includes all 1s can correspond to vector sum reduction, which can be used to generate V0(:,0) = sum(A(0,0), A(1,0), … A(y-1,0)) with the input tensor 550, where y refers to one of the row or column dimensions of the input tensor 520; and row 1 that includes all 1 / y can correspond to vector mean reduction, which can be used to generate V1(:,0) = sum(A(0,0), A(1,0), … A(y-1,0)) / y. The output vectors V0(:,0) and V1(:,0) can be stored in the output vector tensor 560.

[0049] In some embodiments, the weight matrix 540 can include a first row of weights for vector operations generated by the weight matrix generation circuit and a second row of weights for convolution calculations fetched from a weight cache. In this way, one weight matrix 540 can be used to trigger vector operations and convolution calculations in parallel. In particular, Figure 5BThe "multiplication" operator in can include an array of operators that includes a vector reduction operator corresponding to a first row and a multiplication operator (for convolution) corresponding to a second row. This example shows that the hybrid PE array can be configured to perform multiple identical vector reductions, multiple different vector reductions, or a mix of vector reduction and convolution in a single computational cycle.

[0050] Figure 5C Another exemplary use case of performing pooling using a hybrid PE array according to various embodiments is shown. Figure 5C The use case in involves a pooling operation. Pooling operations are common in neural networks and are used to downsample a feature map by summarizing the feature presence in patches (e.g., patches of size 3*3) of the feature map. When the patches are represented as vectors, the pooling operation can be implemented as a vector operation. To this end, the hybrid PE array can first convert the patches into vectors and organize the vectors into an input tensor 580 according to corresponding control signals. Typically, the pooling process involves convolving the patches (e.g., 3*3 patches) by an activation tensor, but the convolution step may be smaller than the size of the patch. Thus, the vectors in the input tensor 580 can have overlapping elements. The weight matrix 570 can include weights generated from a weight generation circuit, where some weight rows can be configured to perform one type of pooling (e.g., sum pooling), while other weight rows can be configured to perform another type of pooling (e.g., max pooling). Using the weight matrix 570 and the input tensor 580, multiple pooling calculations can be performed using the same PE array in the same computational cycle. The output vectors can be stored in the output vector tensor 590.

[0051] Figure 6 An exemplary system design of a hybrid PE array according to various embodiments is shown. The hybrid PE array can be implemented as a hardware accelerator 600. Figure 6 The components of the accelerator 600 in are for illustrative purposes only. According to the implementation, the accelerator 600 can include more, fewer, or alternative components. In some embodiments, the accelerator 600 can include Figure 1 , Figure 2A and Figure 2B all or some of the components in , such as MAC channels. Each MAC channel in the accelerator 600 can include multiple multiplier and adder trees, as shown in Figure 3 . As shown in Figure 4 , the adder tree can include multiple multifunctional adder-subtractors.

[0052] From a functional perspective, in some embodiments, accelerator 600 may include a weight selection circuit 610, an activation selection circuit 620, a plurality of MAC channels 630, and a weight matrix generation circuit 640. In some embodiments, the weight selection circuit 610 may be implemented as a multiplexer and coupled to the weight matrix generation circuit 640 and a weight cache, where the weight cache represents two weight sources. The weight selection circuit 610 may be instructed to obtain weights from these two weight sources according to a control signal. For different types of computations, the weights may be obtained from different sources. The activation selection circuit 620 may be configured to obtain an activation or a vector based on another control signal according to a target computation (e.g., convolution, pooling, vector operation).

[0053] In some embodiments, each MAC channel 630 may be configured to receive a control signal indicating whether to perform a convolution operation or a vector operation; receive one or more weights from at least one of a plurality of weight selectors according to the control signal; receive one or more activations from at least one of a plurality of activation input interfaces according to the control signal; and generate output data based on the one or more weights and the one or more input activations according to the control signal and feed the output data into an output buffer, where: each of the plurality of MAC channels includes a plurality of first circuits for performing multiplication operations according to the control signal and a plurality of second circuits for performing addition operations or subtraction operations.

[0054] In some embodiments, the weight matrix generation circuit 640 may be configured to generate weights for vector reduction operations, where the vector reduction operations include one or more of reduction mean, reduction minimum, reduction maximum, reduction average, reduction addition, or pooling.

[0055] Figure 7 An example computing device is shown in which any of the embodiments described herein may be implemented. The computing device may be used to implement Figures 1 - 6 one or more components of the systems and methods shown therein. The computing device 700 may include a bus 702 or other communication mechanism for passing information, and one or more hardware processors 704 coupled to the bus 702 for processing information. The hardware processor 704 may be, for example, one or more general-purpose microprocessors.

[0056] The computing device 700 may also include a main memory 707, such as random access memory (RAM), a cache, and / or other dynamic storage devices, coupled to the bus 702 for storing information and instructions to be executed by the processor 704. The main memory 707 may also be used to store temporary variables or other intermediate information during the execution of instructions by the processor 704. When stored in a storage medium accessible to the processor 704, such instructions can cause the computing device 700 to appear as a special-purpose machine, customized to perform the operations specified in the instructions. The main memory 707 may contain non-volatile and / or volatile media. Non-volatile media may include, for example, optical or magnetic disks. Volatile media may include dynamic memory. Common forms of media may include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with a hole pattern, RAM, DRAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, or a networked version thereof.

[0057] The computing device 700 may implement the techniques described herein using customized hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic, which in combination with the computing device may cause the computing device 700 to be or be programmed as a special-purpose machine. According to one embodiment, the techniques herein are performed by the computing device 700 in response to one or more sequences of one or more instructions contained in the main memory 707 being executed by the processor 704. Such instructions may be read into the main memory 707 from another storage medium, such as the storage device 708. Execution of the instruction sequences contained in the main memory 707 may cause the processor 704 to perform the processing steps described herein. For example, the processes / methods disclosed herein may be implemented by computer program instructions stored in the main memory 707. When these instructions are executed by the processor 704, they may perform the steps as shown in the corresponding figures and described above. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.

[0058] The computing device 700 also includes a communication interface 710 coupled to the bus 702. The communication interface 710 may provide a two-way data communication coupling to one or more network links connected to one or more networks. As another example, the communication interface 710 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN (or a WAN component communicating with a WAN). A wireless link may also be implemented.

[0059] The execution of certain operations may be distributed among processors, residing not only within a single machine but also deployed across multiple machines. In some example embodiments, a processor or a processor-implemented engine may be located in a single geographical location (e.g., in a home environment, an office environment, or a server farm). In other example embodiments, a processor or a processor-implemented engine may be distributed across multiple geographical locations.

[0060] Each process, method, and algorithm described in the previous sections can be embodied in code modules executed by one or more computer systems or computer processors including computer hardware, and be fully or partially automated by them. These processes and algorithms can be implemented partially or fully in dedicated circuitry.

[0061] When the functions disclosed herein are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. The specific technical solutions (in whole or in part) disclosed herein or aspects contributing to the current technology can be embodied in the form of a software product. The software product can be stored in a storage medium including multiple instructions that cause a computing device (which can be a personal computer, a server, a network device, etc.) to execute all or some of the steps of the methods of the embodiments of this application. The storage medium can include a flash drive, a portable hard disk drive, a ROM, a RAM, a disk, an optical disc, another medium operable to store program code, or any combination thereof.

[0062] Specific embodiments also provide a system that includes a processor and a non-transitory computer-readable storage medium storing instructions executable by the processor to cause the system to perform operations corresponding to the steps in any of the methods of the above embodiments. Specific embodiments also provide a non-transitory computer-readable storage medium configured with instructions executable by one or more processors to cause the one or more processors to perform operations corresponding to the steps in any of the methods of the above embodiments.

[0063] The embodiments disclosed herein can be implemented through a cloud platform, a server, or a server group (collectively referred to as "service system") that interacts with a client. The client can be a terminal device or a client registered by a user on the platform, where the terminal device can be a mobile terminal, a personal computer (PC), and any device that can install a platform application.

[0064] The various features and processes described above can be used independently of each other or in various combinations. All possible combinations and sub - combinations are intended to fall within the scope of the present invention. Additionally, in some implementations, certain method or process blocks may be omitted. The methods and processes described herein are also not limited to any particular sequence, and the associated blocks or states can be executed in a suitable other sequence. For example, the described blocks or states can be executed in an order different from the particular disclosed order, or multiple blocks or states can be combined into a single block or state. Example blocks or states can be executed serially, in parallel, or in some other manner. Blocks or states can be added to or deleted from the disclosed example embodiments. The exemplary systems and components described herein can be configured differently from those described herein. For example, elements can be added, deleted, or rearranged compared to the disclosed example embodiments.

[0065] The various operations of the exemplary methods described herein can be performed, at least in part, by an algorithm. The algorithm can be included in program code or instructions stored in a memory (e.g., the non - transitory computer - readable storage medium described above). Such an algorithm can include a machine - learning algorithm. In some embodiments, a machine - learning algorithm may not explicitly program a computer to perform a function, but can learn from training samples to build a predictive model for performing the function.

[0066] The various operations of the exemplary methods described herein can be performed, at least in part, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, these processors can constitute a processor - implemented engine for performing one or more of the operations or functions described herein.

[0067] Similarly, the methods described herein can be implemented, at least in part, by a processor, with a particular one or more processors being examples of hardware. For example, at least some operations of a method can be performed by one or more processors or a processor - implemented engine. Additionally, one or more processors can also support the execution of relevant operations in a “cloud computing” environment or as “software as a service” (SaaS). For example, at least some operations can be performed by a group of computers (as an example of a machine including processors), and these operations can be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., application programming interfaces (APIs)).

[0068] The execution of certain operations may be distributed among processors, not only residing within a single machine but also deployed across multiple machines. In some example embodiments, the processor or processor - implemented engine can be located in a single geographical location (e.g., in a home environment, an office environment, or a server farm). In other example embodiments, the processor or processor - implemented engine can be distributed across multiple geographical locations.

[0069] In this specification, multiple instances can implement components, operations, or structures described as a single instance. Although the individual operations of one or more methods are shown and described as separate operations, one or more of the individual operations can be performed simultaneously, and the operations are not required to be performed in the order shown. Structures and functions presented as separate components in an example configuration can be implemented as a combined structure or component. Similarly, structures and functions presented as a single component can be implemented as separate components. These and other variations, modifications, additions, and improvements are within the scope of the subject matter of this disclosure.

[0070] As used herein, "or" is inclusive and not exclusive, unless expressly stated otherwise or the context otherwise indicates. Thus, unless expressly stated otherwise or the context otherwise indicates, "A, B, or C" herein means "A, B, A and B, A and C, B and C, or A, B, and C". Further, unless expressly stated otherwise or the context otherwise indicates, "and" is both conjunctive and disjunctive. Thus, unless expressly stated otherwise or the context otherwise indicates, "A and B" herein means "A and B, jointly or severally". Additionally, multiple instances can be provided for resources, operations, or structures described herein as a single instance. Further, the boundaries between various resources, operations, engines, and data stores are to some extent arbitrary, and particular operations are illustrated in the context of a particular illustrative configuration. Other function allocations are conceivable and can fall within the scope of various embodiments of the present invention. Generally, structures and functions presented as separate resources in an example configuration can be implemented as a combined structure or resource. Similarly, structures and functions presented as a single resource can be implemented as separate resources. These and other variations, modifications, additions, and improvements fall within the scope of the embodiments of the invention as represented by the appended claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

[0071] The term "comprising" or "including" is used to indicate the presence of the subsequently recited features, but does not preclude the addition of other features. Unless otherwise specifically stated, or understood in the context of use, conditional language such as "can", "should", "may", or "could", generally conveys that certain embodiments include, while other embodiments do not include, certain features, elements, and / or steps. Thus, such conditional language generally does not imply that the features, elements, and / or steps are in any way required for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or steps are included or are to be performed in any particular embodiment, regardless of user input or prompting.

[0072] Although the overview of the subject matter has been described with reference to specific example embodiments, various modifications and changes can be made to these embodiments without departing from the broader scope of the embodiments of the present invention. For convenience only, the term "invention" may be used herein, either singly or collectively, to refer to these embodiments of the subject matter, and if multiple embodiments are in fact disclosed, there is no intention to voluntarily limit the scope of the present application to any single disclosure or concept.

[0073] The embodiments shown herein are described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments can be used and derived therefrom, so that structural and logical substitutions and changes can be made without departing from the scope of the present invention. Accordingly, the detailed description should not be construed as limiting, and the scope of the various embodiments is defined only by the appended claims and the full scope of equivalents to which those claims are entitled.

Claims

1. A vector operation accelerator for neural network computing, comprising: A plurality of weight selectors configured to obtain weights; A plurality of activation input interfaces configured to obtain activations; And A plurality of multiplier-accumulator (MAC) circuits organized into a plurality of MAC channels, each of the plurality of MAC channels being configured to: Receive a control signal to indicate whether to perform a convolution operation or a vector operation; Receive one or more weights from at least one of the plurality of weight selectors according to the control signal; Receive one or more activations from at least one of the plurality of activation input interfaces according to the control signal; And Generate output data based on the one or more weights and one or more input activations according to the control signal and feed the output data into an output buffer, Wherein: Each of the plurality of MAC channels includes a plurality of first circuits for performing multiplication operations and a plurality of second circuits for performing addition operations or subtraction operations according to the control signal.

2. The vector operation accelerator according to claim 1, wherein: The plurality of second circuits within the MAC channel are organized as a tree, and the second circuits at the leaf level of the tree are configured to receive data from the plurality of first circuits.

3. The vector operation accelerator according to claim 2, wherein, Each of the plurality of second circuits is configured to: Receive a first input and a second input; Determine whether to perform addition or subtraction based on the control signal; In response to a control signal indicating addition, generate the sum or average of the first input and the second input; And In response to a control signal indicating subtraction, generate the minimum or maximum value between the first input and the second input.

4. The vector operation accelerator according to claim 3, wherein, Each of the first input and the second input includes a vector having the same number of dimensions, and in order to generate the minimum value between the first input and the second input, each of the plurality of second circuits is further configured to: Generate an output vector that contains the minimum value of the vectors on each corresponding dimension.

5. The vector operation accelerator according to claim 1, further comprising: A weight matrix generation circuit configured to generate weights for vector reduction operations, wherein the vector reduction operations include one or more of reduction mean, reduction minimum, reduction maximum, reduction average, reduction addition, or pooling.

6. The vector operation accelerator according to claim 5, wherein, Each of the plurality of weight selectors includes a multiplexer coupled to the weight matrix generation circuit and a weight cache.

7. The vector operation accelerator according to claim 6, wherein Each of the plurality of weight selectors is configured to: Obtain weights from the weight cache in response to a control signal indicating a convolution calculation; and Obtain weights from the weight matrix generation circuit in response to a control signal indicating a vector calculation.

8. The vector operation accelerator according to claim 2 further includes an adder-subtractor circuit outside the tree corresponding to the MAC channel, wherein, The adder-subtractor circuit is configured to receive data from the second circuit at the root level of the MAC channel and write the data into the output buffer.

9. The vector operation accelerator according to claim 8, wherein, The adder-subtractor circuit is further configured to: During a first calculation iteration, write a first set of data received from the second circuit at the root level of the MAC channel into the output buffer; And During a second calculation iteration: Receive a set of temporary data from the second circuit at the root level of the MAC channel, Retrieve the first set of data from the output buffer, Calculate a second set of data based on the set of temporary data, the first set of data, and a control signal indicating whether to perform a convolution calculation or a vector operation, and Write the second set of data to the output buffer.

10. The vector operation accelerator according to claim 5, wherein, The plurality of MAC channels are configured to respectively receive a plurality of weight vectors generated by the weight matrix generation circuit to perform a plurality of vector operations in parallel.

11. The vector operation accelerator according to claim 5, wherein, A first subset of the plurality of MAC channels is configured to receive weights from a weight cache, a second subset of the plurality of MAC channels is configured to receive weights generated by the weight matrix generation circuit, and The first subset of the plurality of MAC channels is further configured to perform convolution calculations, the second subset of the plurality of MAC channels is further configured to perform vector operations, and the convolution calculations and vector operations are performed in parallel.

12. A hybrid convolution-vector operation processing system, comprising: A plurality of weight selectors configured to obtain weights; A plurality of activation input interfaces configured to obtain activations; And A plurality of multiplier-accumulator (MAC) circuits organized into a plurality of MAC channels, each of the plurality of MAC channels being configured to: Receive a control signal to indicate whether to perform a convolution operation or a vector operation; Receive one or more weights from at least one of the plurality of weight selectors according to the control signal; Receive one or more activations from at least one of the plurality of activation input interfaces according to the control signal; And Generate output data based on the one or more weights and one or more input activations according to the control signal and feed the output data into an output buffer, Wherein: Each of the plurality of MAC channels includes a plurality of first circuits for performing multiplication operations and a plurality of second circuits for performing addition or subtraction operations according to the control signal.

13. The hybrid convolution-vector operation processing system according to claim 12, wherein, The plurality of second circuits within the MAC channel are organized as a tree, and the second circuits at the leaf level of the tree are configured to receive data from the plurality of first circuits.

14. The hybrid convolution-vector operation processing system according to claim 13, wherein, Each of the plurality of second circuits is configured to: Receive a first input and a second input; Determine whether to perform addition or subtraction based on the control signal; In response to a control signal indicating addition, generate the sum or average of the first input and the second input; And In response to a control signal indicating subtraction, generate the minimum or maximum value between the first input and the second input.

15. The hybrid convolution-vector operation processing system according to claim 12, further comprising: A weight matrix generation circuit configured to generate weights for vector reduction operations, wherein the vector reduction operations include one or more of reduction mean, reduction minimum, reduction maximum, reduction average, reduction addition, or pooling.

16. The hybrid convolution-vector operation processing system according to claim 15, wherein, Each of the plurality of weight selectors includes a multiplexer coupled to the weight matrix generation circuit and a weight cache.

17. The hybrid convolution-vector operation processing system according to claim 15, wherein the plurality of MAC channels are configured to respectively receive a plurality of weight vectors generated by the weight matrix generation circuit to perform a plurality of vector operations in parallel.

18. The hybrid convolution-vector operation processing system according to claim 15, wherein, A first subset of the plurality of MAC channels is configured to receive weights from a weight cache, and a second subset of the plurality of MAC channels is configured to receive weights generated by the weight matrix generation circuit, and the first subset of the plurality of MAC channels is further configured to perform convolution calculations, and the second subset of the plurality of MAC channels is further configured to perform vector operations, and the convolution calculations and vector operations are performed in parallel.

19. The hybrid convolution-vector operation processing system according to claim 13 further includes an adder-subtractor circuit outside the tree corresponding to the MAC channel, wherein, The adder-subtractor circuit is configured to receive data from the second circuit at the root level of the MAC channel and write the data to the output buffer.

20. The hybrid convolution-vector operation processing system according to claim 19, wherein, The adder-subtractor circuit is further configured to: During a first calculation iteration, write a first set of data received from the second circuit at the root level of the MAC channel to the output buffer; and During a second calculation iteration: Receive a temporary data set from the second circuit at the root level of the MAC channel, Retrieve the first set of data from the output buffer, Calculate a second set of data based on the temporary data set, the first set of data, and a control signal indicating whether to perform a convolution calculation or a vector operation, and Write the second set of data to the output buffer.