A Two-Dimensional Convolution Computation Structure with Variable Step Size and a ZNCC Algorithm Accelerator
By designing a two-dimensional convolutional computing structure with variable step size, the problem that existing ZNCC algorithm accelerators cannot adjust the step size and support non-rectangular templates is solved, efficient and flexible ZNCC calculations are achieved, and memory access bandwidth is reduced.
Patent Information
- Application Number
- CN202111277557.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-10-29
AI Technical Summary
The existing ZNCC algorithm accelerator cannot realize free adjustment of step size and matching templates, and cannot support non-rectangular templates. Parallel algorithm accelerator has high requirements for memory access bandwidth, high structural complexity, and long calculation period.
A two-dimensional convolution calculation structure with variable step size is designed, and the convolution calculation is cascaded by multiple PE units in the two-dimensional direction, and the self-accumulation mode is used to perform convolution calculation, which supports free adjustment of horizontal and vertical step sizes, and supports the calculation of non-rectangular templates.
Effectively reduce the calculation amount of the ZNCC algorithm, improve the computing efficiency, realize flexible and efficient ZNCC calculation, support the calculation of non-rectangular templates, and reduce the access bandwidth requirements for memory.
Smart Images

Figure CN113986193B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital circuits, and particularly relates to a two-dimensional convolution calculation structure with variable step size and a ZNCC algorithm accelerator. Background Art
[0002] ZNCC (Zero-mean Normalized Cross Correlation) is a commonly used image matching algorithm based on gray correlation. The formula of ZNCC is
[0003]
[0004] where A ij is the pixel value at the coordinate (i, j) of the template image with size m×n, and S xy ij is the pixel value at the coordinate (x + i, y + j) in the reference image with size M×N, where 0 ≤ x ≤ M - m and 0 ≤ y ≤ N - n. is the sum of pixels of the template image, is the sum of squares of pixels of the template image, is the sum of pixels of the sub-image of the reference image, is the sum of squares of pixels of the sub-image of the reference image, is the sum of products of pixels of the template image and the corresponding sub-image of the reference image. We call the convolution result of the reference image and the template image the result image. The width of the result image = the width of the reference image - the width of the template image + 1, and the height of the result image = the height of the reference image - the height of the template image + 1.
[0005] In the existing related matching algorithms, ZNCC has the advantages of high matching accuracy and good anti-noise performance, and plays an extremely important role in fields such as aircraft aided navigation, weapon guidance, target search and tracking, intelligent transportation, and environmental monitoring. However, this algorithm needs to translate the template image one by one in the original image, traverse the image to calculate the similarity operator at each position, wastes too much time on invalid positions, has a high algorithm complexity, and a slow matching speed.
[0006] In view of the disadvantages of the high calculation complexity and slow matching speed of the ZNCC matching algorithm, many scholars have proposed improvements to this algorithm, which are basically divided into two aspects. On the one hand, it is the optimization and improvement of the similarity operator. For example, in view of the defect of repeated calculation in the operator itself, the BPC algorithm and the iterative method are respectively used to improve the numerator and denominator of the formula. On the other hand, it is to reduce the matching calculation amount, including the improvement of the search area and search strategy. For the improvement of the search area, some studies use the gray statistics method, use a cross-shaped template or a circular template to replace the original rectangular template for image matching, and reduce the amount of calculation at each search point. For the improvement of the search strategy, some studies use the striding search method (skipping rows and columns) and the pyramid hierarchical search method, etc.
[0007] The problems existing in the current research on reducing the computational complexity of the ZNCC matching algorithm are that most of these studies focus on software implementation, that is, using general-purpose processors such as CPUs or DSPs to implement the ZNCC matching algorithm, and the computational performance of software implementation is weak. The existing ZNCC algorithm accelerators are mainly divided into two types: systolic and parallel. However, the systolic algorithm accelerator cannot achieve free adjustment of the step size and the matching template, and does not support non-rectangular templates. The parallel algorithm accelerator has high requirements for the access bandwidth of the memory, high structural complexity, and long computational cycles. Summary of the Invention
[0008] In order to solve the problems existing in the prior art, the present invention provides a two-dimensional convolution calculation structure with variable step size and a ZNCC algorithm accelerator, which combines the research on reducing the computational complexity of the ZNCC algorithm with the high-efficiency parallel computing ability of FPGAs. The step sizes in the horizontal and vertical directions can be freely adjusted, and non-rectangular templates can be supported, which can effectively reduce the computational complexity of the ZNCC algorithm and improve the computational efficiency to achieve flexible and efficient ZNCC calculations.
[0009] To achieve the above object, the present invention provides the following technical solutions:
[0010] A two-dimensional convolution calculation structure with variable step size, characterized in that it includes a plurality of PE units. Each PE unit includes a computational processor, a plurality of registers, an input port, and an output port. The input port includes a reference graph cascaded input port Ain, a template graph parallel input port Bin, and a calculation result cascaded input port Cin. The output port includes a reference graph cascaded output port Aout and a calculation result cascaded output port Cout;
[0011] A plurality of PE units form a rectangular PE array in the two-dimensional direction. The PE units are interconnected through a cascaded register bank. The input port of the cascaded register bank is connected to the output port of one side of the PE unit, and the output port of the cascaded register bank is connected to the input port of the PE unit on the other side;
[0012] Each row of the PE array is connected through a cascaded FIFO. The data output port of each row of cascaded FIFO is connected to the input port of the last cascaded register in that row, and the data input port of each row of cascaded FIFO is connected to the reference graph cascaded output port Aout of the first PE unit in the previous row.
[0013] Further, when the cascaded FIFO of a row of the PE array is bypassed, the input port of the last cascaded register bank in that row of the PE array is connected to the reference graph cascaded output port Aout of the first PE unit in the previous row of the PE array.
[0014] Further, the parallel input ports Bin of the PE units in the same row of the PE array share a data source.
[0015] Further, the computing processor includes a multiplier and an adder;
[0016] Wherein, the input ends of the multiplier are respectively connected with a first register for storing operand A and a second register for storing operand B, and the output end of the multiplier is connected with the input end of the adder;
[0017] The input end of the adder is provided with a first selector, and the input ends of the first selector are respectively connected with the calculation result cascade input port Cin and the output end of the adder.
[0018] Further, the calculation result cascade output port Cout of the PE unit is provided with a second selector for selecting and outputting the calculation result of the adder or the data stored in the calculation result cascade input port Cin.
[0019] Further, registers are provided at the output ends of the multiplier and the adder for maintaining the results of the previous calculation.
[0020] Further, the computing processor of the PE unit is connected with the calculation enable port EN.
[0021] Further, the cascade register group is formed by cascading multiple levels of registers in sequence, and each level of register is connected with a multiplexer MUX, and the multiplexer MUX is used for selecting and outputting the result of the cascade for the nth time.
[0022] Further, the step size of the convolution calculation of the PE unit is,
[0023] The step size of the convolution calculation in the horizontal direction = the number of levels of the cascade register + 1;
[0024] The step size of the convolution calculation in the vertical direction = ceiling(FIFO depth / width of the reference graph).
[0025] A variable-step ZNCC algorithm accelerator, characterized by comprising,
[0026] A template graph memory MdlMem, a reference graph memory RefMem, a result graph memory RslMem, the two-dimensional convolution calculation structure according to any one of claims 1-9, a reference graph sum / sum of squares calculator Sum / SquSum, and a floating-point pipeline Floatpipeline,
[0027] Among them, the input port of the two-dimensional convolution calculation structure is connected to the output ports of the template graph memory MdlMem, the reference graph memory RefMem, and the result graph memory RslMem, and the output port of the two-dimensional convolution calculation structure Conv is connected to the input port of the result graph memory RslMem.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] The two-dimensional convolution calculation structure provided by the present invention is composed of multiple PE units cascaded in two dimensions. Among them, the PE unit uses the self-accumulation mode to perform convolution calculations, that is, each time the PE unit independently completes all the calculations of the results Figure 1 for all points in the convolution calculation. The PE unit uses a fixed-point number format and consists of a calculation processor and several registers. The calculation processor requires one clock cycle to complete one calculation, and the calculated result is stored in the registers at the internal output ends of each. The PE units form a rectangular PE array through horizontal and vertical expansion to jointly complete the convolution calculation. Among them, the PE unit caches data through a cascaded register group horizontally, and the depth of the cascaded register group is configurable. The PE unit caches data through an inter-row cascaded FIFO vertically. There is a FIFO on the right side of each row of the PE array to cache the reference graph data. The PE array can adjust the number of cascaded register levels between PE units and the depth of the caching FIFO, so as to realize the adjustment of the convolution calculation step size. Thus, ZNCC calculations with any step size can be performed, and non-matrix template calculations are supported, which can effectively reduce the calculation amount of the ZNCC algorithm, improve the calculation efficiency of the ZNCC algorithm, and realize efficient and flexible convolution calculations. Description of the Drawings
[0030] Figure 1 is a schematic diagram of the PE unit structure in an embodiment of the present invention;
[0031] Figure 2 is a schematic diagram of the cascaded register group structure in an embodiment of the present invention;
[0032] Figure 3 is a schematic diagram of the one-dimensional PE array structure in an embodiment of the present invention;
[0033] Figure 4 is a schematic diagram of the PE array structure in an embodiment of the present invention;
[0034] Figure 5 is a schematic diagram of the ZNCC algorithm accelerator structure in an embodiment of the present invention. Detailed Embodiments
[0035] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention; the following embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments, and are not used to limit the scope of the present invention.
[0036] With the continuous development of integrated circuit process technology, FPGA (Field Programmable Gate Array) shows unparalleled advantages over traditional DSP processors and ASICs in the field of parallel computing due to its powerful parallel computing ability and reconfigurable characteristics. The parallel computing ability of FPGA is very suitable for the field of digital signal processing with fixed algorithms and large amounts of computation. In order to make their products more widely used in intensive computing, FPGA suppliers have continuously added powerful computing units inside FPGAs, which can perform complex operations such as double-precision floating-point multiplication, division, addition, subtraction, and square root extraction, solving the problem of low computing accuracy of early FPGAs. Currently, FPGAs are widely used in the field of high-performance intensive computing. The highly parallelized characteristics of the ZNCC algorithm determine that it is very suitable for hardware implementation using FPGA. For typical matching image sizes, implementing the ZNCC algorithm through FPGA is often at least one order of magnitude faster than DSP. In order to combine the research on reducing the computational complexity of the ZNCC algorithm with the efficient parallel computing ability of FPGA, it is necessary to study a hardware implementation architecture of the ZNCC matching algorithm with adjustable computational complexity that can be implemented through FPGA to achieve flexible and efficient ZNCC computing.
[0037] The present invention provides a two-dimensional convolution calculation structure with variable step size, including a plurality of PE units. Each PE unit includes a computing processor, a plurality of registers, an input port, and an output port. The input port includes a reference graph cascaded input port Ain, a template graph parallel input port Bin, and a calculation result cascaded input port Cin. The output port includes a reference graph cascaded output port Aout and a calculation result cascaded output port Cout;
[0038] A plurality of PE units form a rectangular PE array in the two-dimensional direction. The PE units are interconnected through a cascaded register bank. The input port of the cascaded register bank is connected to the output port of one side of the PE unit, and the output port of the cascaded register bank is connected to the input port of the PE unit on the other side;
[0039] Each row of the PE array is connected through a cascaded FIFO. The data output port of each row of cascaded FIFO is connected to the input port of the last cascaded register in that row, and the data input port of each row of cascaded FIFO is connected to the reference graph cascaded output port Aout of the first PE unit in the previous row.
[0040] The two-dimensional convolution calculation structure provided by the present invention is composed of multiple PE units cascaded in two dimensions. Among them, the PE unit performs convolution calculation in a self-accumulating mode, that is, each time the PE unit independently completes all the calculations of Figure 1 the points in the convolution calculation. The PE unit adopts a fixed-point number format and is composed of a calculation processor and several registers. The calculation processor requires one clock cycle to complete one calculation, and the calculation result is stored in the register at its internal output end. The PE units form a rectangular PE array through horizontal and vertical expansion to jointly complete the convolution calculation. Among them, the PE unit caches data through a cascaded register group horizontally, and the depth of the cascaded register group is configurable. The PE unit caches data through an inter-row cascaded FIFO vertically. There is a FIFO on the right side of each row of the PE array to cache the reference map data. The PE array can adjust the number of cascaded register levels between PE units and the depth of the cached FIFO, so as to realize the adjustment of the convolution calculation step size. Thus, ZNCC calculation with any step size can be performed, and non-matrix template calculation is supported, which can effectively reduce the calculation amount of the ZNCC algorithm, improve the calculation efficiency of the ZNCC algorithm, and realize efficient and flexible convolution calculation.
[0041] Further, when the cascaded FIFO of a row of the PE array is bypassed, the input port of the last cascaded register group of this row of the PE array is connected to the reference map cascaded output port Aout of the first PE unit of the previous row of the PE array.
[0042] Further, the template map parallel input ports Bin of the PE units in the same row of the PE array share a data source.
[0043] Further, the calculation processor includes a multiplier and an adder;
[0044] Among them, the input end of the multiplier is respectively connected with a first register for storing operand A and a second register for storing operand B, and the output end of the multiplier is connected to the input end of the adder;
[0045] The input end of the adder is provided with a first selector, and the input ends of the first selector are respectively connected with the calculation result cascaded input port Cin and the output end of the adder.
[0046] Further, the calculation result cascaded output port Cout of the PE unit is provided with a second selector for selecting to output the calculation result of the adder or the data stored in the calculation result cascaded input port Cin.
[0047] Further, registers are provided at the output ends of the multiplier and the adder for maintaining the results of the previous calculation.
[0048] Further, the computing processor of the PE unit is connected to the computing enable port EN. The computation of the PE unit is controlled by the EN signal. Only when EN is valid, the PE unit will perform calculations on the operands. Since the PE unit adopts self-accumulative calculation, using the computing enable to select the data involved in the calculation of the PE unit can enable the PE unit to support the calculation of non-rectangular templates.
[0049] Further, as Figure 2 shown, the cascaded register bank is formed by cascading multiple levels of registers in sequence. Each level of register can cache a PE operand, and each level of register is connected to a multiplexer MUX. The multiplexer MUX is used to select and output the result of the nth cascade.
[0050] Further, the PE array can adjust the number of levels of the cascaded registers between PE units and the depth of the cache FIFO, and set the stride of the convolution calculation;
[0051] Among them, the horizontal stride of the convolution calculation = the number of levels of the cascaded registers + 1;
[0052] The vertical stride of the convolution calculation = ceiling(FIFO depth / width of the reference graph).
[0053] A variable-stride ZNCC algorithm accelerator of the present invention includes
[0054] including a template graph memory MdlMem, a reference graph memory RefMem, a result graph memory RslMem, the above two-dimensional convolution calculation structure Conv, a reference graph sum / square sum calculator Sum / SquSum, and a floating-point pipeline Floatpipeline, which together complete the calculation of the ZNCC algorithm.
[0055] Among them, the input port of the two-dimensional convolution calculation structure Conv is connected to the output ports of the template graph memory MdlMem, the reference graph memory RefMem, and the result graph memory RslMem, and the output port of the two-dimensional convolution calculation structure Conv is connected to the input port of the result graph memory RslMem.
[0056] In this embodiment, it is assumed that there are x PE units in the horizontal direction and y PE units in the vertical direction of the PE array. We call the scale of this PE array x×y, where the first PE unit is called PE(0)(0), and the last PE unit is called PE(y)(x).
[0057] The PE unit caches data horizontally through a cascaded register bank. The depth of the cascaded register bank is configurable. The input port of the cascaded register bank is connected to the cascaded output port of the PE units in the x-th column, and the output port of the cascaded register bank is connected to the cascaded input port of the PE units in the 0-th column. During calculation, the reference graph data is loaded through the AinQ port of the PE units in the x-th column and is loaded sequentially to the right through the reference graph cascaded ports AinQ / AoutQ between the PE units. The template graph data is broadcast to all PE units in sequence through the parallel input ports BinQ / BinEN.
[0058] The PE unit caches data vertically through an inter-row cascaded FIFO. There is a FIFO on the x-th column side of each row of the PE array to cache the reference graph data. The data output port of the FIFO is connected to the cascaded input port Ain of the PE unit in the x-th column of the PE array in this row, and the data input port of the FIFO is connected to the cascaded output port Aout of the PE unit in the 0-th column of the previous row of the PE array.
[0059] Among them, the FIFO of the y-th row of the PE array is connected to the memory.
[0060] As Figure 1 shown, in the embodiment of the present invention, the PE unit adopts a fixed-point number format and is composed of a multiplier, an adder, two selectors, and several registers. The multiplier of the PE unit completes one multiplication calculation per clock cycle. The operand A and operand B of the multiplier are respectively stored in the corresponding registers, which come from the input ports AinQ and BinQ of the PE respectively. Among them, AinQ is a cascaded input interface. The operand passed in through the AinQ interface will be output through the output port AoutQ after being stored in a first-level register. BinQ is a parallel input interface, and the BinQ ports of the PE units in the same row of the PE array share a data source.
[0061] The adder of the PE unit completes one addition calculation per clock cycle. Its operand M comes from the output result of the multiplier, and the operand C comes from a two-way selector. The inputs of this selector are the input port Cin of the PE and the output result of the adder itself respectively.
[0062] In addition to being used as the operand for its own calculation, the calculation result of the adder of the PE unit will also be output through the Cout port. In addition to being used as the operand of the adder, the Cin input port of the PE unit will also be output through the Cout port after being stored in a first-level register. There is a two-way selector on the Cout port of the PE unit to select whether to output the calculation result of the adder or the data stored for one clock cycle from the Cin input port.
[0063] The PE unit performs convolution calculations in a self-accumulating manner, that is, the calculation result of each multiply-accumulate operation will be used as the input of the adder in the next step. In this way, a single PE unit can complete all the calculations for one pixel in the result image. Assume the reference image matrix The template image matrix is The first point of the result image matrix is A00B00 + A01B01 + A10B10 + A11B11. The process of using the PE unit for calculation is as follows: In cycle0 (the 0th cycle), A00 is input into the register A of the PE unit through the cascaded input port Ain, and B00 is input into the register B of the PE unit through the parallel input port Bin. The two are multiplied to obtain A00B00. In cycle1 (the 1st cycle), A01 is cascaded into the register A of the PE unit, and B01 is input into the register B of the PE unit. The two are multiplied to obtain A01B01, and it is accumulated with the result generated in the previous cycle in the adder to obtain A00B00 + A01B01. In cycle2 (the 2nd cycle), A02 is cascaded into the register A of the PE unit. At this time, the PE unit does not perform calculations because for the first result of the result image matrix, A02 is invalid data. In cycle3, A10 is input into the register A of the PE unit, and B10 is input into the register B of the PE unit. The two are multiplied to obtain A10B10, and it is accumulated with the result calculated in cycle1 in the adder to obtain A00B00 + A01B01 + A10B10. In cycle4, A11 is input into the register A of the PE unit, and B11 is input into the register B of the PE unit. The two are multiplied to obtain A11B11, and it is accumulated with the result generated in the previous cycle in the adder to finally obtain the calculation result of the first point of the result image A00B00 + A01B01 + A10B10 + A11B11. After the calculation is completed, the PE unit outputs the calculation result through the Cout port and can then load new data for the next round of calculation.
[0064] Since the PE unit performs convolution calculations in a self-accumulating mode, calculating the multiply-accumulate of one template image pixel and the corresponding reference image pixel per clock cycle, it is possible to control the calculation enable port EN of the PE unit to only allow valid pixels to participate in the calculation, enabling the PE unit to support calculations on non-matrix templates. Assume the reference image matrix is The template image matrix is It can be seen that the template image is not a standard rectangle but a cross. When convolving the reference image and the template image, the first point of the resulting image matrix is A01B01 + A10B10 + A11B11 + A12B12 + A21B21. The process of using the PE unit for calculation is as follows: At cycle0 (the 0th cycle), A01 is input into the register A of the PE unit through the cascaded input port Ain, and B01 is input into the register B of the PE unit through the parallel input port Bin, and the two are multiplied to obtain A01B01. At cycle1 (the 1st cycle), A02 is input into the register A of the PE unit. At this time, the PE unit does not perform calculations because for the first result of the resulting image matrix, A02 is invalid data. At cycle2 (the 2nd cycle), A03 is input into the register A of the PE unit. At this time, the PE unit does not perform calculations. At cycle3 (the 3rd cycle), A10 is input into the register A of the PE unit, and B10 is input into the register B of the PE unit. The two are multiplied to obtain A10B10, which is added to the calculation result of cycle0 to obtain A01B01 + A10B10. At cycle4 (the 4th cycle), A11 is input into the register A of the PE unit, and B11 is input into the register B of the PE unit. The two are multiplied to obtain A11B11, which is added to the calculation result of cycle3 to obtain A01B01 + A10B10 + A11B11. At cycle5 (the 5th cycle), A12 is input into the register A of the PE unit, and B12 is input into the register B of the PE unit. The two are multiplied to obtain A12B12, which is added to the calculation result of cycle11 (the 11th cycle) to obtain A01B01 + A10B10 + A11B11 + A12B12. At cycle6 (the 6th cycle), A13 is input into the register A of the PE unit. At this time, the PE unit does not perform calculations. At cycle7 (the 7th cycle), A20 is input into the register A of the PE unit. At this time, the PE unit does not perform calculations. At cycle8 (the 8th cycle), A21 is input into the register A of the PE unit, and B21 is input into the register B of the PE unit. The two are multiplied to obtain A21B21, which is added to the calculation result of cycle5 to obtain the final calculation result A01B01 + A10B10 + A11B11 + A12B12 + A21B21 of the first point of the resulting image matrix.
[0065] Since the method of using the PE unit to calculate the convolution has strong regularity, when calculating the convolution, the values of the EN signal in each cycle in the calculation process of the PE unit for a pixel point of the resulting image can be stored in advance in the order from first to last, and then taken out in sequence during the calculation and sent to the PE unit to control the multiplier-accumulator. Using this method, the PE unit can perform convolution calculations for templates of any shape.
[0066] The PE units are expanded into a one-dimensional PE array by cascading registers horizontally. When performing horizontal expansion, each PE is interconnected through a cascaded register bank. The input port of the cascaded register bank is connected to the AoutQ port of the right PE unit, and the output port is connected to the AinQ port of the left PE unit, which is used for cascaded caching of the PE operand Ain. The maximum number of cascaded levels in the cascaded register bank can be set according to actual needs. The value of this setting plus one is the maximum step number that the PE array can support horizontally during convolution calculation. For example, assuming that the cascaded register bank can cache the PE operand A for up to 8 levels, then the step numbers that the PE array can support horizontally during convolution calculation are 1 to 9.
[0067] Specifically, the structure of the one-dimensional PE array is as Figure 3 shown, which includes the structure diagram of the PE unit and the structure diagram of the cascaded register bank. The PE array shown in the figure consists of n PEs and n cascaded register banks, and the working mode of any one PE / cascaded register bank can be independently configured. The PEs in the one-dimensional PE array are cascaded through the Ain / Aout ports and Cin / Cout ports. The Bin port of the PE unit is a parallel input interface, that is, the Bin input interfaces of all PE units in the one-dimensional PE array share the same data source.
[0068] Assume the reference diagram is the matrix The template diagram is the matrix To calculate the convolution with a step size of 2 between the two, the number of cascaded register stages between PE units needs to be set to 1. Suppose there are 4 PEs for calculation, and the PEs are PE(0), PE(1), PE(2), and PE(3) from left to right. After the loading is completed, the reference graph data in PE(0) is A00, the reference graph data in PE(1) is A02, the reference graph data in PE(2) is A04, and the reference graph data in PE(3) is A06. A01, A03, and A05 are cached in the cascaded registers between PEs, A07 is cached in the FIFO, and A10 - A17 are cached in the memory. After the loading is completed, the PE array starts to calculate. At cycle0, the template graph data B00 is broadcast to all PEs to start the calculation. After the calculation is completed, the result of PE(0) is A00B00, the result of PE(1) is A02B00, the result of PE(2) is A04B00, and the result of PE(3) is A06B00. At cycle1, the reference graph data is shifted left through the cascaded interface AinQ / AoutQ. After the shift, the reference graph data in PE(0) is A01, the reference graph data in PE(1) is A03, the reference graph data in PE(2) is A05, and the reference graph data in PE(3) is A07. While shifting, the template graph data B01 is broadcast to all PEs for calculation. After the calculation is completed, the result of PE(0) is A00B00 + A01B01, the result of PE(1) is A02B00 + A03B01, the result of PE(2) is A04B00 + A05B01, and the result of PE(3) is A06B0 + A07B01. At cycle2 - cycle7, the PEs do not calculate, and the reference graph data is shifted left through the cascaded interface in turn. At cycle8, the reference graph data in PE(0) is A10, the reference graph data in PE(1) is A12, the reference graph data in PE(2) is A14, and the reference graph data in PE(3) is A16. At this time, the template graph data B10 is broadcast to all PEs to start the calculation. After the calculation is completed, the result of PE(0) is A00B00 + A01B01 + A10B10, the result of PE(1) is A02B00 + A03B01 + A12B10, the result of PE(2) is A04B00 + A05B01 + A14B10, and the result of PE(3) is A06B0 + A07B01 + A16B10. At cycle9, the reference graph data is shifted left by one bit through the cascaded interface. The reference graph data in PE(0) is A11, the reference graph data in PE(1) is A14, the reference graph data in PE(2) is A16, and the reference graph data in PE(3) is A17. The PE array broadcasts the template graph data B11 to all PEs to start the calculation.The result after the calculation is the final result. Among them, the result of PE(0) is A00B00 + A01B01 + A10B10 + A11B11, the result of PE(1) is A02B00 + A03B01 + A12B10 + A12B11, the result of PE(2) is A04B00 + A05B01 + A14B10 + A14B11, and the result of PE(3) is A06B0 + A07B01 + A16B10 + A16B11. After the calculation is completed, the calculation result is cascaded and output to the left through the Cin / Cout ports of the PE unit.
[0069] The one-dimensional PE array can be extended into a two-dimensional PE array longitudinally through the Ain / Aout ports and Cin / Cout ports of the PE units at both left and right ends. When performing longitudinal expansion, the Ain / Aout ports between adjacent rows of one-dimensional PE arrays are cascaded through a cascaded FIFO. The input port of the cascaded FIFO is connected to the Aout port of the leftmost PE unit of the lower PE array, and the output port of the cascaded FIFO is connected to the Ain port of the rightmost PE unit of the upper PE array. The Cin / Cout ports between adjacent rows of one-dimensional PE arrays are cascaded in a direct connection manner, that is, the Cin port of the rightmost PE of the upper PE array is directly connected to the Cout port of the leftmost PE of the lower PE array. The depth of the cascaded FIFO can accommodate the number of rows of the reference graph plus 1, which is the maximum step number supported by the PE array longitudinally during convolution calculation. For example, assuming that the depth of the cascaded FIFO can accommodate at most 4 rows of reference graph data, then when the PE array performs convolution calculation, the step numbers supported longitudinally are 1 to 4.
[0070] The depth of the cascaded FIFO can be adjusted as needed. The maximum value = (the step number supported by the PE array longitudinally - 1) × the number of rows of the reference graph. The minimum value = 0, that is, the FIFO is bypassed, and the Ain / Aout ports of the PE units at both ends of the FIFO are directly connected.
[0071] The structure of the PE array is as Figure 4 shown. We define the number of one-dimensional PE arrays as the number of rows y of the PE array, and the number of PE units in a one-dimensional PE array as the number of columns x of the PE array. The PE array shown in the figure consists of y one-dimensional PE arrays and y cascaded FIFOs, and the depth of any one of the FIFOs can be independently adjusted. The number of rows y and the number of columns x of the PE array can be any values, and can be set according to the multiplier resource layout of the FPGA device in practical applications. However, considering avoiding repeated calculations, the value of x × y should be greater than or equal to the result Figure 1 size of the row. From Figure 4It can be seen that the PE array can be expanded into a one-dimensional PE array. The expanded one-dimensional PE array is composed of y PE sub-arrays, each of which includes x PE units and a FIFO on the right.
[0072] The PE array uses the reference image data to be loaded through the AinQ / AoutQ port cascade, the template image data to be loaded through the BinQ port broadcast, and the result is output through the Cout / Cin port cascade to perform convolution calculation. Before performing the convolution calculation, the PE array first needs to group the PE units according to the image size. Assuming that the number of columns of the result image (that is, the number of pixels in the horizontal direction of the result image) is ResultX, the number of rows (that is, the number of pixels in the vertical direction of the result image) is ResultY, the result of ResultX / m is rounded up to the integer value W, which is the number of rows of the PE array required to load a row of reference image data. W=1 indicates that one row of PE can complete the loading of a row of reference image data. W>1 indicates that multiple rows of PE are required to complete the loading of a row of reference image data. If this happens, the FIFOs between the PE arrays responsible for jointly loading a row of reference image data will be bypassed. For example, assuming W=3, it means that three rows of PE are required to realize the loading of a row of reference image data. At this time, the FIFOs between these three rows of PE (two in total) will be bypassed, and the three rows of PE are jointly responsible for loading a row of reference image data.
[0073] Due to the limited resources of the PE array, it is sometimes impossible to completely load all the reference image data. Therefore, the PE array adopts a batch loading method when performing convolution calculations. The reference image is loaded in order from top to bottom and from left to right. At least one row of reference image data is loaded in each batch. When a row of reference image data in the group does not fill the PE array allocated to it, the excess PE units do not participate in the calculation and are only used as cascade storage of data. When a row of reference image data in the group is more than the PE array allocated to it, the excess data is stored in the rightmost FIFO of the PE array allocated to it. For example, suppose 3 rows of PE are required to load a row of reference image data, but the number of PE units in the 3 rows of PE is less than the reference image data. Figure 1 If the number of rows of data is equal, then after the PE array loads the reference image into the PE unit, it will cache the remaining data in the rightmost FIFO of these three rows of PE.
[0074] After completing the PE unit grouping, it is necessary to preload data for the PE units involved in the calculation. Preloading is divided into preloading of reference images and preloading of template images. When preloading the reference image, the PE array takes the reference image data from the memory and loads it in the order from top to bottom and from left to right through the PE unit at the lower right corner of the PE array ( Figure 4The Ain ports of PE(y)(x) in Figure 4 are sequentially loaded into the PE units. Referring to the figure, the data will be sequentially passed forward through the Ain / Aout cascading interfaces between the PE units until the first data of the reference figure is loaded into the PE unit (
[0075] PE(0)(0)) in the upper left corner of the PE array.
[0076] When preloading the template figure, the PE array fetches the first data of the template figure from the memory and broadcasts it to all PE units through the parallel input port Bin of the PE array.
[0077] After completing the data preloading, the PE array can start the convolution calculation. When performing the calculation, the PE array needs to update the operands of the PE units every cycle. The update method is the same as the preloading method. When updating the reference figure data, the PE array needs to fetch the new data from the memory and load the data into the PE array through the cascading input port Ain(y)(x) of the PE(y)(x) unit in the lower right corner of the PE array by cycle. When updating the template figure data, the PE array needs to fetch the new data from the memory and broadcast the data to all PE units through the parallel input port Bin of the PE array by cycle. When performing the convolution calculation, data update is not required every cycle, so dedicated control logic is needed to control the timing of data update. The final result of the convolution calculation is cached in the output register of the adder of the PE unit. The PE array outputs the results sequentially through the Cin / Cout cascading interfaces between the PE units.
[0078] The following is an example to illustrate the configuration method of the PE array. Assume that the scale of the PE array is 16×64, that is, the PE array has 64 rows and each row contains 16 PE units.
[0079] Example 1: Calculate the convolution of a 128×128 reference figure and a 68×68 template figure. The stride is 1.
[0080] Configuration method: The PE array is divided in the way of 64×16, that is, 4 rows of PE units form a one-dimensional PE array, and the FIFO in the middle of these 4 rows of PE units is set as bypass. The depth of the rightmost FIFO is set to 64 to ensure that the 4 rows of PE units and the FIFO can just store the reference Figure 1 row data. The cascading level of the cascading registers between all PE units is set to 0. When loading, a one-dimensional PE array (including the FIFO) completes the loading of one row of the reference graph. When calculating, every 4 rows of PE units can complete the calculation of one row of the result graph. The 64×16 PE array can calculate 16 rows of the result graph simultaneously. Since the number of 4 rows of PE units (64) is greater than the result Figure 1 row size (61), the rightmost 3 PE units of the one-dimensional PE array do not participate in the calculation.
[0081] Example 2: Calculate the convolution of a 128×128 reference graph and a 68×68 template graph. The stride is 2.
[0082] Configuration method: The PE array is divided in the way of 32×32, that is, 2 rows of PE units form a one-dimensional PE array, and the FIFO in the middle of these 2 rows of PE units is set as bypass. The depth of the rightmost FIFO is set to 192. Ensure that 2 rows of PE units (including the cascading registers between them) and the FIFO can just store the data of two rows of the reference graph. The cascading level of the cascading registers of all PE units is set to 1. When loading, a one-dimensional PE array (including the FIFO) completes the loading of two rows of the reference graph. When calculating, every 2 rows of PE units can complete the calculation of one row of the result graph. The 32×32 PE array can calculate 32 rows of the result graph simultaneously. Since the number of 2 rows of PE units (32) is greater than the result Figure 1 row size (31), the rightmost 1 PE unit of the one-dimensional PE array does not participate in the calculation.
[0083] Based on the above two-dimensional convolution calculation structure, the present invention proposes a structure of a ZNCC algorithm accelerator with variable stride and supporting non-matrix templates. The structure of the ZNCC algorithm accelerator is as Figure 5 shown.
[0084] The ZNCC algorithm accelerator includes a template graph memory MdlMem, a reference graph memory RefMem, a result graph memory RslMem, the above two-dimensional convolution calculation structure Conv, a reference graph sum / square sum calculator Sum / SquSum, and a floating-point pipeline Floatpipeline.
[0085] Among them, MdlMem is used to store the template graph data, and during the process of loading the template graph data, it automatically completes the summation of the template graph in the ZNCC formula ( Figure 5The sum in summdl) and the square difference of the template graph ( Figure 5 The calculation of stdmdl) in.
[0086] Among them, the accumulated sum of the template graph adopts the fixed-point number format, and the calculation of the square difference of the template graph adopts the 32-bit floating-point number format.
[0087] Among them, RefMem is used to store the reference graph data. Since RefMem needs to provide the reference graph data to the convolution accelerator and the reference graph accumulated sum / square accumulated sum calculator at the same time, two independent FIFOs are set inside the reference graph memory to load the data for convolution calculation and accumulated sum / square accumulated sum calculation respectively.
[0088] To ensure continuous calculation, the bandwidth of data written into the FIFO is set to be more than twice the read bandwidth of the FIFO. RslMem is used to store the result graph data, and the result graph is stored in the floating-point number format. The two-dimensional convolution calculation structure Conv is used to complete the calculation of convdataout) in the ZNCC formula ( Figure 5 using the fixed-point number format. The accumulated sum / square accumulated sum module Sum / SquSum is used to complete the calculation of sumref) in the ZNCC formula ( Figure 5 and ( Figure 5 the calculation of squsumref) in using the fixed-point number format. The floating-point pipeline FloatPipeline is used to complete the final ZNCC formula calculation. Before calculation, the floating-point pipeline will convert each sub-item in the ZNCC formula calculated by other modules into 32-bit floating-point numbers and then perform the final calculation. Since the square root / division of floating-point numbers in the FPGA takes a long time, the floating-point pipeline requires a long pipeline establishment time. After the pipeline is established, the floating-point pipeline outputs a ZNCC final calculation result every clock cycle, and the calculation result is stored in the result graph memory ( Figure 5 in RslData).
[0089] A first-level MUX is set at the input port of the result graph memory to select whether the data source of the memory is the calculation result of the convolution module or the calculation result of the floating-point pipeline. The result of the floating-point pipeline is the final result of the ZNCC calculation. If the convolution module cannot complete all the result graph calculations at one time, it needs to calculate in multiple loops, store the results in batches in the result graph memory, and then uniformly retrieve the convolution calculation results from the memory and send them to the floating-point pipeline module after all calculations are completed.
[0090] Compared with the pulsating accelerator, the ZNCC algorithm accelerator proposed by the present invention can support the free adjustment of the step size and the matching template, with strong flexibility; compared with the parallel accelerator, the accelerator proposed by the present invention has the advantage of a small data access bandwidth, because the accelerator proposed by the present invention only needs to load one reference graph data and one template graph data into the PE array in one clock cycle to realize the pipelining of convolution calculation, which greatly reduces the access bandwidth to the memory, reduces the design complexity, and helps to improve the calculation frequency.
[0091] On the other hand, the ZNCC algorithm accelerator designed by the present invention can well support the existing research on improving the efficiency of the ZNCC algorithm by reducing the matching calculation amount. These studies reduce the ZNCC calculation amount by improving the matching template (cross-shaped, circular, etc.), search strategy (interlaced row and column matching, etc.). These algorithms can be easily implemented on the algorithm accelerator proposed in this patent; compared with software implementation, the hardware accelerator has stronger computing performance.
[0092] Compared with the limitation of the fixed step size of 1 and only supporting rectangular templates in its previous generation product, the horizontal / vertical step size of the present invention can be arbitrarily adjusted, and non-rectangular convolution templates can be supported, which can well support the optimization algorithm of reducing the ZNCC matching calculation amount by adjusting the step size and template, effectively improving the computing performance.
Claims
1. A two-dimensional convolution calculation structure with variable step size, Characterized in that, It includes multiple PE units. Each PE unit includes a calculation processor, multiple registers, an input port and an output port. The input port includes a reference graph cascaded input port Ain, a template graph parallel input port Bin and a calculation result cascaded input port Cin. The output port includes a reference graph cascaded output port Aout and a calculation result cascaded output port Cout; Multiple PE units form a rectangular PE array in two dimensions. The PE units are interconnected through a cascaded register bank. The input port of the cascaded register bank is connected to the output port of one side of the PE unit, and the output port of the cascaded register bank is connected to the input port of the PE unit on the other side; Each row of the PE array is connected through a cascaded FIFO. The data output port of each row of cascaded FIFO is connected to the input port of the last cascaded register in that row, and the data input port of each row of cascaded FIFO is connected to the reference graph cascaded output port Aout of the first PE unit in the previous row.
2. The two-dimensional convolution calculation structure with variable step size according to claim 1, Characterized in that, When the cascaded FIFO of a row of the PE array is bypassed, the input port of the last cascaded register bank in that row of the PE array is connected to the reference graph cascaded output port Aout of the first PE unit in the previous row of the PE array.
3. The two-dimensional convolution calculation structure with variable step size according to claim 1, Characterized in that, The template graph parallel input ports Bin of the PE units in the same row of the PE array share a data source.
4. The two-dimensional convolution calculation structure with variable step size according to claim 1, Characterized in that, The calculation processor includes a multiplier and an adder; Among them, the input end of the multiplier is respectively connected with a first register for storing operand A and a second register for storing operand B. The output end of the multiplier is connected to the input end of the adder; A first selector is arranged at the input end of the adder. The input ends of the first selector are respectively connected with the calculation result cascaded input port Cin and the output end of the adder.
5. The two-dimensional convolution calculation structure with variable step size according to claim 4, Characterized in that, A second selector is arranged at the calculation result cascaded output port Cout of the PE unit for selecting to output the calculation result of the adder or the data stored in the calculation result cascaded input port Cin.
6. The two-dimensional convolution calculation structure with variable step size according to claim 4, Characterized in that, Registers are arranged at the output ends of both the multiplier and the adder for maintaining the result of the previous calculation.
7. The two-dimensional convolution calculation structure with variable step size according to claim 1, Characterized in that, The calculation processor of the PE unit is connected to the calculation enable port EN.
8. The two-dimensional convolution calculation structure with variable step size according to claim 1, Characterized in that, The cascaded register bank is formed by cascading multiple levels of registers in sequence. Each level of register is connected to a multiplexer MUX, and the multiplexer MUX is used to select and output the result of the cascading for the nth time.
9. A two-dimensional convolution calculation structure with variable step size according to claim 8, characterized in that, the step size of the convolution calculation of the PE unit is, the step size in the horizontal direction of the convolution calculation = the number of stages of the cascaded register + 1; the step size in the vertical direction of the convolution calculation = ceiling(FIFO depth / width of the reference graph).
10. A variable step size ZNCC algorithm accelerator, characterized in that, it includes, a template graph memory MdlMem, a reference graph memory RefMem, a result graph memory RslMem, the two-dimensional convolution calculation structure according to any one of claims 1-9, a reference graph sum / square sum calculator Sum / SquSum, and a floating-point pipeline Floatpipeline, wherein, the input port of the two-dimensional convolution calculation structure is connected to the output ports of the template graph memory MdlMem, the reference graph memory RefMem, and the result graph memory RslMem, and the output port of the two-dimensional convolution calculation structure Conv is connected to the input port of the result graph memory RslMem.
Citation Information
Patent Citations
Configurable multiply accumulation cell and multiply accumulation array consisting of same
CN103677739A
An FPGA parallel system of convolution neural network algorithm
CN109032781A