Method and device for improving resource utilization rate of on-chip multiplier-adder of FPGA

By using multiplexers and splitters in the matrix multiplier of FPGA, the problem that FPGAs in the prior art is difficult to efficiently utilize multiplier resources, and more efficient resource usage and performance improvement are achieved.

CN113919489BActive Publication Date: 2025-06-06PARADIGM ZHILIAN (SHENZHEN) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010653627.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-08
Publication Date
2025-06-06
Estimated Expiration
2040-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently utilize the on-chip multiplier resources of FPGAs, or it is necessary to introduce a large amount of additional computing overhead at the software level, resulting in limited performance of FPGAs in deep neural network hardware accelerators.

Method used

By introducing multiplexers and splitters into the matrix multiplier, time-sharing multiplexing of the multiplier is realized and the hardware structure is optimized, so that the matrix multiplier can balance more flexible resource usage and performance.

Benefits of technology

It improves the efficiency of FPGA's on-chip multiplier resources, avoids additional software-level overhead, and improves the computing performance of matrix multipliers (such as deep neural network accelerators).

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113919489B_ABST
    Figure CN113919489B_ABST
Patent Text Reader

Abstract

Provided are a method and device for improving the resource utilization rate of an on-chip multiplier of an FPGA. For each multiplication-addition circuit in a matrix multiplier of an FPGA consisting of a predetermined number of multiplication-addition circuits, the input end is connected to a multiplexer, and the output end is connected to a demultiplexer, and the following operations are performed: k inputs are sent to the multiplexer; in k clock cycles, the k inputs are sent to the multiplication-addition circuit through the multiplexer, wherein one input is selected for sending in each clock cycle; in each clock cycle of the k clock cycles, a corresponding output is sent to the demultiplexer through the multiplication-addition circuit; in the k clock cycles, the corresponding k outputs are output through the demultiplexer, wherein the k outputs are sent to the multiplexer to which the subsequent multiplication-addition circuit is connected, wherein k is a multiplexing parameter of the multiplication-addition circuit, and wherein each multiplication-addition circuit includes a multiplier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of hardware structure optimization, and more specifically, to a method and device for improving the utilization rate of on-chip multiplier-adder resources of FPGA by optimizing the hardware structure of a matrix calculator based on the characteristics of the FPGA chip. Background Art

[0002] In the implementation of matrix multiplication calculations that require regularity, for example, in the implementation of deep neural network hardware accelerators, the usage of multiplier-adder resources often directly determines the performance of the accelerator. Since the available multiplier-adder resources on the FPGA chip are fixed and limited, how to improve the efficiency of the use of the on-chip multiplier-adder resources is a very critical optimization hotspot. Due to the limitations of the regular shape of the matrix multiplier, existing solutions often cannot efficiently utilize the available on-chip multiplier-adder resources of the FPGA, or require the introduction of a large amount of additional computing overhead at the software level. Summary of the invention

[0003] Exemplary embodiments of the present invention are intended to overcome the above-mentioned shortcomings of not being able to efficiently utilize the available on-chip multiplier-accumulator resources of the FPGA or having a large amount of additional software calculation overhead.

[0004] According to one aspect of the present invention, a method for improving the resource utilization rate of an on-chip multiplier of an FPGA is provided, wherein, for each multiplication-addition circuit in a matrix multiplier composed of a predetermined number of multiplication-addition circuits of the FPGA, its input end is connected to a multiplexer, and its output end is connected to a demultiplexer, and the method comprises: for each multiplication-addition circuit in a matrix multiplier composed of a predetermined number of multiplication-addition circuits of the FPGA, the following operations are performed: sending k inputs to the multiplexer; in k clock cycles, sending the k inputs to the multiplication-addition circuit through the multiplexer, wherein one input is selected to be sent in each clock cycle; in each clock cycle of the k clock cycles, sending a corresponding output to the demultiplexer through the multiplication-addition circuit; in the k clock cycles, outputting corresponding k outputs through the demultiplexer, wherein the k outputs are sent to the multiplexer to which the subsequent multiplication-addition circuit is connected, wherein k is a multiplexing parameter of the multiplication-addition circuit, and wherein each multiplication-addition circuit comprises a multiplier.

[0005] Optionally, the FPGA can be used to implement a deep neural network hardware accelerator.

[0006] Optionally, the on-chip matrix multiplier of the FPGA may include (2 M ×2 M ) / k multiplication-addition circuits, where M is the bit width of the input channel of the matrix multiplier.

[0007] Optionally, the method may further include: obtaining the number A of available on-chip multipliers and adders of the FPGA; and determining optimal values ​​of the input bit width M and the multiplexing parameter k according to the number A of available multipliers and adders.

[0008] Optionally, based on the number A of available multipliers and adders, the step of determining the optimal values ​​of the input bit width M and the multiplexing parameter k may include: setting initial values ​​of M and k; obtaining performance test results of the FPGA based on the initial values ​​of M and k; obtaining performance test results of the FPGA based on increased M values ​​and / or k values ​​by incrementally setting the M value and / or k value; and determining the M value and k value when the performance of the FPGA was last improved as the optimal values ​​of M and k.

[0009] Optionally, the step of setting initial values ​​of M and k may include: setting the initial value of M to floor(sqrt(A)); and setting the initial value of k to 1.

[0010] Optionally, the step of obtaining the performance test result of the FPGA based on the increased M value and / or k value by incrementally setting the M value and / or k value may include: looping the following operations until the performance of the FPGA does not improve: (1) setting M to M+1; (2) determining 4 M / Is k less than A? (3) When 4 M / When k is not less than A, set k to k+1 and jump to step (2); (4) When 4 M / k is less than A, obtaining a performance test result of the FPGA based on the current M value and k value; (5) determining whether the performance of the FPGA is improved based on a comparison between the currently obtained performance test result of the FPGA and the performance measurement result of the FPGA obtained last time.

[0011] According to another aspect of the present invention, there is provided an apparatus for improving the resource utilization rate of an on-chip multiplier-adder of an FPGA, comprising: a predetermined number of multiplexers; the predetermined number of demultiplexers, wherein for each multiplication-addition circuit in a matrix multiplier of the FPGA composed of the predetermined number of multiplication-addition circuits, its input end is connected to a multiplexer, and its output end is connected to a demultiplexer, wherein the multiplexer is configured to receive k inputs and send the k inputs to the multiplication-addition circuit in k clock cycles, wherein one input is selected for sending in each clock cycle, wherein in each clock cycle of the k clock cycles, the multiplication-addition circuit sends a corresponding one output to the demultiplexer, and the demultiplexer is configured to output corresponding k outputs through the demultiplexer in the k clock cycles, wherein the k outputs are sent to the multiplexer to which the subsequent multiplication-addition circuit is connected, wherein k is a multiplexing parameter of the multiplication-addition circuit, and wherein each multiplication-addition circuit includes a multiplier.

[0012] Optionally, the FPGA can be used to implement a deep neural network hardware accelerator.

[0013] Optionally, the on-chip matrix multiplier of the FPGA may include (2 M ×2 M ) / k multiplication-addition circuits, where M is the bit width of the input channel of the matrix multiplier.

[0014] Optionally, the device may further include: a parameter determiner configured to obtain the number A of available on-chip multipliers and adders of the FPGA, and determine optimal values ​​of the input bit width M and the multiplexing parameter k based on the number A of available multipliers and adders.

[0015] Optionally, the parameter determiner can be configured to: set initial values ​​of M and k; obtain performance test results of the FPGA based on the initial values ​​of M and k; obtain performance test results of the FPGA based on increased M values ​​and / or k values ​​by incrementally setting the M value and / or k value; and determine the M value and k value when the performance of the FPGA was last improved as the optimal values ​​of M and k.

[0016] Optionally, the parameter determiner may be configured to: set an initial value of M to floor(sqrt(A)); and set an initial value of k to 1.

[0017] Optionally, the parameter determiner may be configured to: loop through the following operations until the performance of the FPGA does not improve: (1) set M to M+1; (2) determine 4 M / Is k less than A? (3) When 4 M / When k is not less than A, set k to k+1 and jump to step (2); (4) When 4M / k is less than A, obtaining a performance test result of the FPGA based on the current M value and k value; (5) determining whether the performance of the FPGA is improved based on a comparison between the currently obtained performance test result of the FPGA and the performance measurement result of the FPGA obtained last time.

[0018] According to another aspect of the present invention, a system is provided comprising at least one computing device and at least one storage device storing instructions, wherein the instructions, when executed by the at least one computing device, cause the at least one computing device to perform operations performed by a parameter determiner.

[0019] According to another aspect of the present invention, a computer-readable storage medium storing instructions is provided, wherein when the instructions are executed by at least one computing device, the at least one computing device is caused to perform operations performed by a parameter determiner.

[0020] According to the method and device for improving the resource utilization rate of the on-chip multiplier-adder of the FPGA of the present invention, the hardware structure is optimized by time-sharing multiplexing the on-chip multiplier-adder of the FPGA, so that a balance between resource utilization and performance can be achieved more freely. Without increasing the additional software overhead, the utilization efficiency of the on-chip multiplier-adder of the FPGA is improved, thereby improving the computing performance of the matrix multiplier (for example, a deep neural network accelerator).

[0021] In addition, according to the method and device for improving the resource utilization rate of the on-chip multiplier-adder of the FPGA of the present invention, the parameters that optimize the performance of the matrix multiplier (e.g., a deep neural network accelerator) can be configured through a parameter exploration process, such as the bit width M of the input channel of the matrix multiplier and the multiplexing parameter k of the multiplication-addition circuit, thereby effectively improving the utilization efficiency of the on-chip multiplier-adder of the FPGA. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] These and / or other aspects and advantages of the present invention will become more clear and easier to understand from the following detailed description of the embodiments of the present invention in conjunction with the accompanying drawings, in which:

[0023] Figure 1 is a schematic diagram showing a systolic array of a conventional matrix multiplier.

[0024] Figure 2a , Figure 2b and Figure 2c is a schematic diagram showing an apparatus for improving resource utilization rate of an on-chip multiplier-adder of an FPGA according to an exemplary embodiment of the present invention.

[0025] Figure 3 is a flowchart illustrating a method of determining optimal values ​​of parameters M and k according to an exemplary embodiment of the present invention.

[0026] Figure 4 is a flow chart illustrating a method for improving on-chip multiplier-adder resource utilization of an FPGA according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the present invention, exemplary embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0028] In recent years, artificial intelligence technologies represented by neural networks have developed rapidly and have been widely and deeply applied in image, speech and video processing. However, as deep neural networks quickly become popular, the model depth of various neural networks is increasing, and their model size is also increasing. Obviously, a larger neural network model means that there are more parameters that need to be calculated. The increase in these calculations not only affects the time of model training, but also affects the response speed of the model in actual deployment.

[0029] In the pure software algorithm implementation of deep neural networks, the training and reasoning performance of the model is often directly limited by the computing power of the central processing unit. However, since the central processing unit is a general-purpose hardware, it is often not suitable for making too much improvement for a specific computing power. In contrast, dedicated hardware solutions based on ASIC and FPGA can optimize and improve the computing power of the bottleneck part in the calculation in a targeted manner, and are being increasingly adopted in engineering practice. Since the current popular deep neural networks often contain a considerable number of convolutional layers or fully connected layers, the main operations in these deep neural networks are often disassembled into matrix multiplication operations. Depending on the structure of the neural network, the time occupied by these matrix multiplication operations can often reach more than 90% of the total computing time. Therefore, many hardware acceleration solutions are based on using dedicated hardware to improve the computing performance of matrix multiplication.

[0030] From the mathematical definition of matrix multiplication, we know that the multiplier-adder is the basic computing unit for implementing the matrix multiplier. In existing hardware acceleration solutions, the number of multipliers used directly determines the number of parameters that can be multiplied and added simultaneously in a single clock cycle, and often directly determines the performance of the hardware accelerator. In the FPGA-based implementation, since the number of on-chip multiplier-adder resources of the FPGA is an inherent property of the FPGA device, that is, for a selected FPGA device, the total amount of its multiplier-adder resources is limited and fixed. This means that for FPGA-based hardware optimization solutions, the efficiency of the use of on-chip multiplier-adder resources often directly determines the acceleration performance of the FPGA deep neural network accelerator.

[0031] Since the main operations in common deep neural networks can be decomposed into matrix multiplication operations, and the systolic array is a common engineering method for implementing efficient matrix multiplication accelerators. Therefore, in engineering practice, FPGA-based deep neural network accelerators often implement a 2 M ×2 M A matrix multiplier is implemented with a systolic array of size M, where M is the bit width of the matrix multiplier input channel. Figure 1 is a schematic diagram showing a systolic array of an existing matrix multiplier. Figure 1 As shown, M=2 in the example systolic array, and 4 M The multiplication and addition circuits are arranged in a pulsating array, and each multiplication and addition circuit needs to contain a hardware multiplier. Each group of multiplication and addition circuits is connected to the multiplication and addition circuits adjacent to it above, below, and to the left and right. In each clock cycle, the 4 M Each multiplication-addition circuit will operate simultaneously to perform multiplication-addition operations, and pass the operands and / or operation results to the right or lower multiplication-addition circuit according to the established rules before the end of the clock cycle. M Two 2 are completed once in one clock cycle M ×2 M Matrix multiplication between large and small matrices. Since the matrix multiplier needs to perform 4 M Multiplication and addition operations, it is obvious that to achieve such a 2 M ×2 M A matrix multiplier of size 4 is required M A multiplier-accumulator.

[0032] To improve the performance of the matrix multiplier described above, the most common engineering practice is to increase the size of M. That is, every time M increases by 1, the computing power of the matrix multiplier per clock cycle will reach 4 times the original, but at the same time, the multiplier-adder resources required by the matrix multiplier will also become 4 times the original. In other words, when M = {1, 2, 3, 4, 5, 6, 7, 8...}, to implement the above matrix multiplier on FPGA will require {4, 16, 64, 256, 1024, 4096, 16384, 65536...} multiplier-adder resources respectively.

[0033] Since the number of on-chip multiplier-adder resources of FPGA is fixed and limited, in this usage mode, we often cannot effectively utilize these multiplier-adder resources. For example: The 10GX1150 FPGA device has 3036 on-chip multipliers, but according to the above matrix multiplication implementation scheme, we can only implement a maximum of M = 5 matrix multipliers. In other words, we can only use 1024 / 3036 = 33.7% of the multiplier resources, which greatly limits the performance of the deep neural network accelerator that can be implemented by FPGA.

[0034] There are usually two solutions to solve this problem in the existing solutions, but they both have their own significant defects and disadvantages.

[0035] Existing Solution 1: Implementing irregular-shaped matrix multipliers on FPGA.

[0036] This solution changes the shape of the matrix multiplier to change the number of multiplier resources used. For example, if a 48×60 matrix multiplier is implemented on an FPGA, it can be seen from the above that the number of multipliers used by this multiplier is 2880. It can be seen that this solution can indeed greatly improve the utilization rate of the on-chip multiplier resources of the FPGA. However, since the shape of the matrix is ​​closely related to the network structure such as the filter size / number of channels, the use of irregular matrix multipliers will not only increase the difficulty and complexity of the software layer allocation and scheduling, but also often introduce a large amount of additional computing overhead in the irregular segmentation, thereby further reducing the performance of the deep neural network accelerator.

[0037] Existing Solution 2: Implement multiple smaller and independent matrix multipliers on the FPGA.

[0038] This solution improves the efficiency of the FPGA's on-chip multiplier resources by arranging multiple smaller regular-shaped matrix multipliers in parallel. For example, 5 ×2 5 The number of matrix multipliers required for this type of multiplier is 2048. However, since these matrix multipliers are independent of each other and are subject to data transmission between matrix multipliers, this solution is often used in batch processing applications, that is, scheduling multiple matrix multipliers to calculate multiple independent input groups. In this application mode, although the processing throughput of the deep neural network accelerator is increased, the latency of its input processing cannot be effectively reduced.

[0039] In order to improve the resource utilization rate of the on-chip multiplier-adder of the FPGA and avoid the above-mentioned defects and disadvantages, the present application optimizes the hardware structure of the matrix calculator according to the characteristics of the FPGA device. Specifically, by flexibly applying multiplexers and demultiplexers, time-sharing multiplexing of the multiplier-adder is realized in the matrix multiplier, so that the new matrix multiplier can more freely strike a balance between resource utilization and performance. That is, a new multiplication-addition circuit reuse parameter k is introduced in the accelerator design process, and multiplexers and demultiplexers are applied at the input and output ends of the multiplication-addition circuit, so that a 2 M ×2 M The size of the matrix multiplier can be in 2 M ×k clock cycles to complete two 2 M ×2 M Matrix multiplication between large and small matrices can freely realize k-way time-division multiplexing of multiplier-adder resources in the matrix multiplier. The number of multipliers required for this matrix multiplier is only 1 / k×4 M , making the multiplier resource consumption 1 / k of the original matrix multiplier, thereby greatly improving the multiplier resource utilization rate. In addition, the present application also introduces a parameter exploration process for efficient use of multipliers, by designing and adjusting the bit width M of the input channel of the matrix multiplier and the size of the multiplexing parameter k of the multiplication and addition circuit, so that the set parameters M and k can enable the accelerator to achieve the best performance, thereby efficiently utilizing the inherent multiplier resources on the FPGA chip.

[0040] Next, referring to FIGS. 2 to Figure 4 A device and method for improving the resource utilization rate of an on-chip multiplier-adder of an FPGA according to an exemplary embodiment of the present invention are described in detail.

[0041] Figure 2a , Figure 2b and Figure 2c is a schematic diagram showing an apparatus 200 for improving resource utilization of an on-chip multiplier-adder of an FPGA according to an exemplary embodiment of the present invention.

[0042] Can be targeted at Figure 1 In the systolic array shown, some multiplication-addition circuits are merged, k groups of multiplication-addition circuits in the original circuit are combined into a new group of circuits, and k-1 multiplication-addition circuits are reduced. At the same time, the k inputs in the original k groups of circuits are connected to the multiplexer of the new circuit, and the outputs in the original k groups of circuits are connected to the demultiplexer in the new circuit.

[0043] like Figure 2a As shown, assuming that M is 2 and k is 2, as Figure 1 The original circuit shown requires 16 multiplication-addition circuits, while the new circuit reduces the number of multiplication-addition circuits to only 8 by horizontally compressing the original circuit.

[0044] Another example Figure 2b As shown, assuming that M is 2 and k is 4, Figure 1 The original circuit shown requires 16 multiplication-addition circuits, while the new circuit reduces the number of multiplication-addition circuits to only 4 by compressing the original circuit horizontally and vertically.

[0045] For clarity, in Figure 2a and Figure 2b The multiplexer at the input end and the demultiplexer connected to the output end of each multiplication and addition circuit in the new circuit are not shown. Figure 2c A schematic diagram of a multiply-add circuit connected to a multiplexer and a demultiplexer is shown in FIG. Figure 2a and Figure 2b Each multiplication-addition circuit in the new circuit shown in FIG. Figure 2c The structure shown.

[0046] like Figure 2c As shown, a multiplexer is connected to the input end of each multiplication and addition circuit, and a demultiplexer is connected to the output end of each multiplication and addition circuit. Therefore, in a matrix multiplier including a predetermined number of multiplication and addition circuits, the device 200 may include the predetermined number of multiplexers 201 and the predetermined number of demultiplexers 202. By using multiplexers and demultiplexers, k inputs and outputs can share the original accumulation circuit in different clock cycles. Figure 2c As shown, the solid line input of the multiplexer and the solid line output of the demultiplexer represent the input and output of the current clock cycle, and the dotted line input and the dotted line output represent the input and output of other clock cycles.

[0047] Specifically, the multiplexer 201 can receive k inputs, where k is a multiplexing parameter of the multiplication-addition circuit 222, and k can be determined by the user according to the requirements. An exemplary preferred determination method of the parameter k will be described in detail later.

[0048] The multiplexer 201 sends the k inputs to the multiplication-addition circuit 222 in k clock cycles, wherein one input is selected for transmission in each clock cycle. The multiplication-addition circuit 222 sends a corresponding output to the demultiplexer 202 in each clock cycle in the k clock cycles. The demultiplexer 202 outputs corresponding k outputs through the demultiplexer 202 in k clock cycles, wherein the k outputs can be sent to the multiplexer to which the subsequent multiplication-addition circuit is connected.

[0049] For example, in clock cycle 1, the multiplexer 201 selects the 1st input to be sent to the multiplication and addition circuit 222, the multiplication and addition circuit 222 sends the 1st output corresponding to the 1st input to the demultiplexer 202, and the demultiplexer 202 can select the received output as the 1st output. In clock cycle 2, the multiplexer 201 selects the 2nd input to be sent to the multiplication and addition circuit 222, the multiplication and addition circuit 222 sends the 2nd output corresponding to the 2nd input to the demultiplexer 202, and the demultiplexer 202 can select the received output as the 2nd output. By analogy, the multiplexer 201 selects the kth input to be sent to the multiplication and addition circuit 222, the multiplication and addition circuit 222 sends the kth output corresponding to the kth input to the demultiplexer 202, and the demultiplexer 202 can select the received output as the kth output. The demultiplexer 202 may send the received k outputs to multiplexers connected to subsequent multiplication-addition circuits associated with the multiplication-addition circuits (eg, multiplication-addition circuits arranged on the right side of the multiplication-addition circuit 222 and / or the multiplication-addition circuit below).

[0050] In the deep neural network hardware accelerator, the bit width M of the input channel of the optimal matrix calculator is closely related to the network structure such as the filter size and the number of input and output channels. At the same time, the bit width M of the input channel is also related to the memory bandwidth and the main frequency that the hardware accelerator can achieve. With the help of the reprogrammable characteristics of FPGA, the optimal values ​​of parameters M and k can be determined through the parameter exploration process to improve the utilization rate of on-chip multiplier resources.

[0051] According to an exemplary embodiment of the present invention, the apparatus 200 may further include a parameter determiner (not shown). The parameter determiner may obtain the number A of available multipliers and adders on the FPGA, and determine the optimal values ​​of M and k based on A.

[0052] Specifically, the parameter determiner may set initial values ​​of M and k. For example, the initial value of M may be set to floor(sqrt(A)), and the initial value of k may be set to 1. Subsequently, the performance (e.g., computing speed, etc.) of the FPGA (e.g., a deep neural network hardware accelerator) may be tested based on the initial values ​​of M and k. Therefore, the parameter determiner may obtain performance test results based on the initial values ​​of M and k. Subsequently, the parameter determiner may incrementally set the M value and / or the k value for testing the performance of the FPGA, thereby obtaining performance test results based on the increased M value and / or the k value. Subsequently, the parameter determiner may determine the M value and the k value when the performance of the FPGA was last improved as the optimal values ​​of M and k.

[0053] Below, refer to Figure 3 A method in which a parameter determiner according to an exemplary embodiment of the present invention determines optimal values ​​of parameters M and k is described in detail. Figure 3is a flowchart illustrating a method of determining optimal values ​​of parameters M and k according to an exemplary embodiment of the present invention.

[0054] Reference Figure 3 , in step 301, the parameter determiner sets the initial values ​​of M and k.

[0055] In step 302, the parameter determiner obtains performance test results of the FPGA based on initial values ​​of M and k.

[0056] In step 303 , the parameter determiner may set M to M+1.

[0057] In step 304, the parameter determiner may determine 4 M / Is k less than A?

[0058] If the judgment is no in step 304 , then in step 305 , the parameter determiner may set k to k+1 and jump to step 304 to execute step 304 again.

[0059] If the answer is yes in step 304 , then in step 306 , the parameter determiner may obtain a performance test result of the FPGA based on the current M value and k value.

[0060] In step 307 , based on the comparison between the performance test result of the FPGA obtained in step 306 and the performance measurement result of the FPGA obtained in the last test, it is determined whether the performance of the FPGA is improved.

[0061] If the determination in step 307 is yes, the process jumps to step 303 and executes step 303 again.

[0062] If the determination in step 307 is no, then in step 308, the parameter determiner determines the M value and k value corresponding to the last test (ie, the M value and k value when the performance was last improved) as the optimal value.

[0063] Of course, the above parameter exploration process is only exemplary, and the parameter exploration process may also be performed according to other feasible methods or adjustments of steps.

[0064] Figure 4 is a flow chart illustrating a method for improving on-chip multiplier-adder resource utilization of an FPGA according to an exemplary embodiment of the present invention.

[0065] For each multiplication-addition circuit in the matrix multiplier of the FPGA, which is composed of a predetermined number of multiplication-addition circuits, the input end is connected to a multiplexer, and the output end is connected to a demultiplexer, and the following operations are performed: Figure 4 The operation shown.

[0066] Reference Figure 4, in step 401, k inputs are sent to the multiplexer.

[0067] In each clock cycle of k clock cycles, in step 402, the multiplexer sends one input of the k inputs to the multiplication and addition circuit; in step 403, the multiplication and addition circuit sends one output corresponding to the one input to the demultiplexer, and in step 404, the demultiplexer selects the received output as one output.

[0068] When k clock cycles expire, in step 405, the demultiplexer outputs the corresponding k outputs respectively. For example, the demultiplexer may send the k outputs to the multiplexer connected to the subsequent multiplication and addition circuit (e.g., the right multiplication and addition circuit and / or the lower multiplication and addition circuit). In addition, the demultiplexer may send one output when receiving the output of the multiplication and addition circuit in a single clock cycle, or may temporarily store the output of the multiplication and addition circuit when receiving it in a single clock cycle, and then output it after several clock cycles or when k clock cycles expire. The present invention is not limited to this.

[0069] In addition, the method also includes a method for determining parameters M and k. According to an exemplary embodiment of the present invention, the number A of available multipliers and adders on the FPGA can be obtained; and based on A, the optimal values ​​of M and k can be determined.

[0070] According to an exemplary embodiment of the present invention, the initial values ​​of M and k may be set first. For example, the initial value of M may be set to floor(sqrt(A)), and the initial value of k may be set to 1. Subsequently, the performance (e.g., computing speed, etc.) of the FPGA (e.g., a deep neural network hardware accelerator) may be tested based on the initial values ​​of M and k, thereby obtaining performance test results based on the initial values ​​of M and k. Subsequently, the M value and / or the k value may be set incrementally to test the performance of the FPGA, thereby obtaining performance test results based on the increased M value and / or k value. Subsequently, the M value and the k value when the performance of the FPGA was last improved may be determined as the optimal values ​​of M and k.

[0071] According to an exemplary embodiment of the present invention, the step of obtaining a performance test result based on an increased M value and / or k value may include: looping the following operations until the performance of the FPGA does not improve: (1) setting M to M+1; (2) determining 4 M / Is k less than A? (3) When 4 M / When k is not less than A, set k to k+1 and jump to step (2); (4) When 4 M / k is less than A, obtaining a performance test result of the FPGA based on the current M value and k value; (5) determining whether the performance of the FPGA is improved based on a comparison between the currently obtained performance test result of the FPGA and the performance measurement result of the FPGA obtained last time.

[0072] In addition, the method and device for improving the on-chip multiplier-adder resource utilization rate of FPGA according to the present invention can be applied not only to deep neural network hardware accelerators, but also to the implementation of any other matrix multiplication calculations that require rules.

[0073] According to the method and device for improving the resource utilization rate of the on-chip multiplier-adder of the FPGA of the present invention, the hardware structure is optimized by time-sharing multiplexing the on-chip multiplier-adder of the FPGA, so that a balance between resource utilization and performance can be achieved more freely. Without increasing the additional software overhead, the utilization efficiency of the on-chip multiplier-adder of the FPGA is improved, thereby improving the computing performance of the matrix multiplier (for example, a deep neural network accelerator).

[0074] In addition, according to the method and device for improving the resource utilization rate of the on-chip multiplier-adder of the FPGA of the present invention, the parameters that optimize the performance of the matrix multiplier (e.g., a deep neural network accelerator) can be configured through a parameter exploration process, such as the bit width M of the input channel of the matrix multiplier and the multiplexing parameter k of the multiplication-addition circuit, thereby effectively improving the utilization efficiency of the on-chip multiplier-adder of the FPGA.

[0075] The above has been referred to Figures 2 to Figure 4 A method and apparatus for improving the resource utilization rate of an on-chip multiplier-adder of an FPGA according to an exemplary embodiment of the present invention are described.

[0076] The parameter determiner in the device for improving the resource utilization rate of the on-chip multiplier-adder of the FPGA according to the present invention can be configured as software, hardware, firmware or any combination of the above items that perform a specific function. For example, the parameter determiner can correspond to a dedicated integrated circuit, or to a pure software code, or to a module that combines software and hardware. In addition, one or more functions implemented by the security monitoring module can also be uniformly executed by components in a physical entity device (e.g., a processor, a client or a server, etc.).

[0077] In addition, refer to Figure 3 The operations performed by the described parameter determiner may be implemented by a program (or instruction) recorded on a computer-readable storage medium. For example, according to an exemplary embodiment of the present invention, a computer-readable storage medium storing instructions may be provided, wherein when the instructions are executed by at least one computing device, the at least one computing device is prompted to perform the operations performed by the parameter determiner.

[0078] The computer program in the computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. It should be noted that the computer program can also be used to perform additional steps in addition to the above steps or perform more specific processing when performing the above steps. The contents of these additional steps and further processing have been described in reference to Figure 3 It is mentioned in the description of the related method, so it will not be repeated here to avoid repetition.

[0079] It should be noted that the parameter determiner according to the exemplary embodiment of the present invention can completely rely on the operation of the computer program to realize the corresponding functions, that is, each unit corresponds to each step in the functional architecture of the computer program, so that the entire system is called through a special software package (for example, lib library) to realize the corresponding functions.

[0080] On the other hand, the parameter determiner can be implemented by hardware, software, firmware, middleware, microcode or any combination thereof. When implemented by software, firmware, middleware or microcode, the program code or code segment for performing the corresponding operation can be stored in a computer-readable medium such as a storage medium, so that the processor can perform the corresponding operation by reading and running the corresponding program code or code segment.

[0081] For example, the parameter determiner of the exemplary embodiment of the present invention may also be implemented as a computing device, which includes a storage component and a processor, wherein a set of computer-executable instructions is stored in the storage component, and when the set of computer-executable instructions is executed by the processor, operations performed by the parameter determiner according to the exemplary embodiment of the present invention are performed.

[0082] Specifically, the computing device can be deployed in a server or client, or can be deployed on a node device in a distributed network environment. In addition, the computing device can be a PC, a tablet device, a personal digital assistant, a smart phone, a web application, or other device capable of executing the above instruction set.

[0083] Here, the computing device is not necessarily a single computing device, but may also be any device or circuit collection that can execute the above instructions (or instruction sets) individually or jointly. The computing device may also be part of an integrated control system or system manager, or may be configured as a portable electronic device that is interconnected with a local or remote (e.g., via wireless transmission) interface.

[0084] In a computing device, a processor may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, a processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0085] Some operations described in the operations performed by the parameter determiner according to the exemplary embodiments of the present invention may be implemented by software, some operations may be implemented by hardware, and furthermore, these operations may be implemented by a combination of software and hardware.

[0086] The processor may execute instructions or codes stored in one of the storage components, which may also store data. Instructions and data may also be sent and received over a network via a network interface device, which may employ any known transmission protocol.

[0087] The storage component may be integrated with the processor, for example, RAM or flash memory is arranged within an integrated circuit microprocessor, etc. In addition, the storage component may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The storage component and the processor may be operatively coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor can read files stored in the storage component.

[0088] In addition, the computing device may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the computing device may be connected to each other via a bus and / or a network.

[0089] The operations performed by the parameter determiner according to the exemplary embodiments of the present invention may be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams may be equally integrated into a single logical device or operate according to non-precise boundaries.

[0090] Therefore, refer to Figure 3 The described operations performed by the parameter determiner may be implemented by a system including at least one computing device and at least one storage device storing instructions.

[0091] According to an exemplary embodiment of the present invention, at least one computing device is a computing device for executing an operation performed by a parameter determiner according to an exemplary embodiment of the present invention, and a computer executable instruction set is stored in the storage device. When the computer executable instruction set is executed by the at least one computing device, the execution reference Figure 3 The steps described are those of the operations performed by the parameter determiner.

[0092] The above describes various exemplary embodiments of the present invention, it should be understood that the above description is only exemplary, not exhaustive, and the present invention is not limited to the disclosed exemplary embodiments. Without departing from the scope and spirit of the present invention, many modifications and changes are obvious to those of ordinary skill in the art. Therefore, the scope of protection of the present invention should be based on the scope of the claims.

Claims

1. A method for improving the utilization rate of on-chip multiplier-adder resources of FPGA, in, For each multiplication-addition circuit in a matrix multiplier of the FPGA consisting of a predetermined number of multiplication-addition circuits, the input end of each multiplication-addition circuit is connected to a multiplexer, and the output end is connected to a demultiplexer, the method includes: The following operations are performed for each multiplication-addition circuit in a matrix multiplier composed of a predetermined number of multiplication-addition circuits of the FPGA: Sending k inputs to the multiplexer; In k clock cycles, the k inputs are sent to the multiplication-addition circuit through the multiplexer, wherein one input is selected to be sent in each clock cycle; In each clock cycle of the k clock cycles, a corresponding output is sent to the demultiplexer through a multiplication-addition circuit; In the k clock cycles, the corresponding k outputs are output through the demultiplexer, wherein the k outputs are sent to a multiplexer to which a subsequent multiplication and addition circuit is connected, Wherein, k is the multiplexing parameter of the multiplication and addition circuit, Each multiplication-addition circuit includes a multiplication-addition device.

2. The method according to claim 1, in, The FPGA is used to implement a deep neural network hardware accelerator.

3. The method according to claim 1, in, The FPGA's on-chip matrix multiplier includes A multiplication-addition circuit is provided, where M is the bit width of the input channel of the matrix multiplier.

4. The method according to claim 3, further comprising: include: Obtaining the number A of available on-chip multipliers and adders of the FPGA; According to the number of available multipliers A, the optimal values ​​of the input bit width M and the multiplexing parameter k are determined.

5. The method according to claim 4, in, According to the number A of available multipliers and adders, the step of determining the optimal value of the input bit width M and the multiplexing parameter k comprises: Set the initial values ​​of M and k; Obtaining a performance test result of the FPGA based on initial values ​​of M and k; Acquire a performance test result of the FPGA based on the increased M value and / or k value by incrementally setting the M value and / or k value; The M value and the k value when the performance of the FPGA is last improved are determined as the optimal values ​​of M and k.

6. The method according to claim 5, in, The steps to set the initial values ​​of M and k include: Set the initial value of M to ; Set the initial value of k to 1.

7. The method according to claim 5, in, The step of obtaining the performance test result of the FPGA based on the increased M value and / or k value by incrementally setting the M value and / or k value comprises: The following operations are performed in a loop until the performance of the FPGA does not improve: (1) Set M to M+1; (2) Judgment Is it less than A? (3) When If it is not less than A, set k to k+1 and jump to step (2); (4) When When it is less than A, obtaining the performance test result of the FPGA based on the current M value and k value; (5) Determine whether the performance of the FPGA is improved based on a comparison between the currently acquired performance test result of the FPGA and the last acquired performance measurement result of the FPGA.

8. A device for improving the utilization rate of on-chip multiplier-adder resources of FPGA, include: a predetermined number of multiplexers; the predetermined number of demultiplexers, Wherein, for each multiplication-addition circuit in the matrix multiplier of the FPGA composed of the predetermined number of multiplication-addition circuits, its input end is connected to a multiplexer, and its output end is connected to a demultiplexer, The multiplexer is configured to receive k inputs and send the k inputs to the multiply-add circuit in k clock cycles, wherein one input is selected for transmission in each clock cycle, and wherein the multiply-add circuit sends a corresponding output to the demultiplexer in each clock cycle of the k clock cycles. The demultiplexer is configured to output corresponding k outputs through the demultiplexer in the k clock cycles, wherein the k outputs are sent to a multiplexer to which a subsequent multiplication and addition circuit is connected, Wherein, k is the multiplexing parameter of the multiplication and addition circuit, Each multiplication-addition circuit includes a multiplication-addition device.

9. The device as claimed in claim 8, in, The FPGA is used to implement a deep neural network hardware accelerator.

10. The device according to claim 8, in, The FPGA's on-chip matrix multiplier includes A multiplication-addition circuit is provided, where M is the bit width of the input channel of the matrix multiplier.

11. The device according to claim 10, further comprising: include: The parameter determiner is configured to obtain the number A of available multipliers and adders on the FPGA, and determine the optimal values ​​of the input bit width M and the multiplexing parameter k according to the number A of available multipliers and adders.

12. The device according to claim 11, in, The parameterizer is configured as: Set the initial values ​​of M and k; Obtaining a performance test result of the FPGA based on initial values ​​of M and k; Acquire a performance test result of the FPGA based on the increased M value and / or k value by incrementally setting the M value and / or k value; The M value and the k value when the performance of the FPGA is last improved are determined as the optimal values ​​of M and k.

13. The device according to claim 12, in, The parameterizer is configured as: Set the initial value of M to ; Set the initial value of k to 1.

14. The device according to claim 12, in, The parameterizer is configured as: The following operations are performed in a loop until the performance of the FPGA does not improve: (1) Set M to M+1; (2) Judgment Is it less than A? (3) When If it is not less than A, set k to k+1 and jump to step (2); (4) When When it is less than A, obtaining the performance test result of the FPGA based on the current M value and k value; (5) Determine whether the performance of the FPGA is improved based on a comparison between the currently acquired performance test result of the FPGA and the last acquired performance measurement result of the FPGA.

15. A system comprising at least one computing device and at least one storage device storing instructions, in, When the instructions are executed by the at least one computing device, the at least one computing device is prompted to perform the method for improving the utilization rate of on-chip multiplier-adder resources of FPGA according to any one of claims 1 to 7.

16. A computer-readable storage medium storing instructions, in, When the instructions are executed by at least one computing device, the at least one computing device is prompted to execute the method for improving the utilization rate of on-chip multiplier-adder resources of FPGA as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Neural network unit with memory layout to perform efficient 3-dimensional convolutions

    CN108133262A

  • A method and a system for accelerating circuit optimization

    CN109726413A