A super-resolution system and method for high-speed image acquisition

CN115965528BActive Publication Date: 2026-09-25XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211651686.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-21
Publication Date
2026-09-25
Estimated Expiration
2042-12-21

AI Technical Summary

Technical Problem

这种外部存储器增加了系统的整体成本,降低了性能,并增加了功耗

Benefits of technology

[0038]本发明超分辨率系统,基于16点-卷积双三次插值算法,Bicubic顶层计算模块通过纵向窗先按列计算一维插值中间结果,将算出的一维插值中间结果暂存在移位寄存器中,每周期向右滑动一次;4个周期后,横向窗按行计算一维插值中间结果的一维插值,采用移位寄存器进行缓存并实现4点一维插值中间结果值并行输出,能对运算过程中的中间结果进行高效复用,可以在纵向窗并行度减少到1的同时保持输出数据不变,从而节省硬件资源,在一定程度上降低了资源消耗。本发明为高频率、低功耗的超分辨率系统,此系统实现将地面站接收到的低分辨率图像转换为高分辨率图像的功能,可满足无人机系统在高速移动下的高精度监测需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application provides a kind of high-speed image acquisition-oriented super-resolution system and method, comprising: Padding module, Linebuffer module, Bicubic top layer calculation module and shift register;Padding module supplements original image, and outputs supplementary image to Linebuffer module;Linebuffer module carries out line buffering to supplementary image, and synchronously outputs four rows of data in sequence;Bicubic top layer calculation module receives four rows of data output by Linebuffer module, calculates one-dimensional interpolation intermediate result by longitudinal window by column and temporarily stores in shift register, after four columns are calculated, one-dimensional interpolation of one-dimensional interpolation intermediate result is calculated by horizontal window by row, and interpolation point pixel value is obtained and output.The application improves processing efficiency and reduces resource consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of FPGA platform algorithm hardware acceleration, specifically relating to a super-resolution system and method for high-speed image acquisition. Background Technology

[0002] When high-speed mobile drones monitor the ground, they often use high-frame-rate airborne cameras to take pictures. However, due to the limitations of transmission bandwidth and onboard computer computing power, the resolution of the cameras they carry is usually low. As a result, the images they transmit back to the ground station are blurry and lack details, making it difficult to meet the needs of high-precision monitoring.

[0003] Super-resolution, also known as upsampling, is a class of algorithms that improve the resolution and quality of videos or images. It is widely used in video editing and image processing fields such as surveillance video processing, high-resolution cameras, and film restoration, and can convert low-resolution images received by ground stations into high-resolution images.

[0004] Traditional interpolation algorithms in super-resolution still hold significant research value and hardware implementation significance due to their advantages such as regular formulas and ease of hardware implementation. Traditional interpolation algorithms mainly include nearest-neighbor interpolation, bilinear interpolation, and bicubic interpolation. Nearest-neighbor interpolation is simple and fast, but the quality of the interpolated image is not high. Bilinear interpolation has the properties of a low-pass filter, weakening high-frequency components in the image information, resulting in blurred edges in the interpolated image. Bicubic interpolation uses a 4×4 matrix (16 points in total) near the interpolation point as a reference to obtain the pixel value of the interpolation point. It overcomes the stepped boundary problem of nearest-neighbor interpolation, and its interpolation effect on image edges is better than that of bilinear interpolation.

[0005] There are two main implementations of bicubic interpolation algorithms: 16-point ordinary bicubic interpolation and 16-point convolutional bicubic interpolation. The main difference lies in how the interpolation kernel is calculated. The former uses the bicubic term formula of the convolution kernel to obtain the interpolation coefficients; the latter simplifies the calculation of the two-dimensional planar interpolation kernel to the x and y directions, calculating the interpolation coefficients at each one-dimensional direction sequentially. The 16-point convolutional bicubic interpolation method has a simpler formula, and the hardware implementation circuits for the x and y directions can be reused.

[0006] However, the existing method does not separate the horizontal and vertical windows, requiring the caching of 16 raw pixels, which increases hardware consumption and waiting time.

[0007] Furthermore, the 16-point convolutional bicubic interpolation method first calculates coefficients in the x and y directions. Existing methods typically select α = -0.5 or α = -0.75 as the convolution kernel coefficients for the interpolation algorithm. Since general-purpose computer platforms use floating-point ALUs (Arithmetic Logic Units) for operations, choosing coefficients with decimal places does not incur additional computational overhead. Moreover, due to the extremely large dynamic range of floating-point numbers, the extra decimal places can be accurately represented without introducing errors. However, for FPGA platforms, floating-point operations are extremely slow and incur significant resource overhead. Fixed-point arithmetic units are highly sensitive to data bit width. Existing solutions using α = -0.5 or α = -0.75 cause the data bit width involved in the operation to increase exponentially, leading to additional resource consumption. If low-order data is discarded to maintain the data bit width in order to save this consumption, the quality of the interpolated image will deteriorate. Meanwhile, the current 16-point convolutional bicubic interpolation method does not consider the special optimizations for 4x upsampling. For 4x upsampling, if interpolation is performed directly, the number of points inserted between two adjacent reference points is not fixed. This requires counting the number of points to be inserted between two reference points, introducing additional computational complexity. In addition, due to the irregular relative distances, finite decimals must be used to approximate infinite decimals in fixed-point calculations, which introduces rounding errors. Furthermore, existing schemes use different relative distances for calculation when calculating interpolation coefficients. While this simplifies the formula and reduces the computational load of addition and multiplication, it does not consider the cumulative errors caused by multi-stage pipelines.

[0008] The optimal architecture represents the best trade-off between accuracy and hardware cost. Therefore, careful consideration must be given to output image quality, hardware resource consumption, and throughput performance. The bicubic interpolation architecture proposed by G. Mahale et al. generates high-quality upsampled images but requires a large amount of output resources and is extremely energy-intensive. Bicubic interpolation implemented by A. Nuno-Magand, K. Gribbon, F. Sabbetzadeh, and others stores the entire image pixel in external memory, thus requiring a considerable amount of external memory. This external memory increases the overall system cost, reduces performance, and increases power consumption. Furthermore, most current hardware implementations of bicubic image interpolation using FPGAs employ floating-point units. Floating-point units incur significant area overhead, generate high power consumption, and negatively impact overall system performance.

[0009] In summary, although there are various solutions for super-resolution hardware system design based on 16-point convolutional bicubic interpolation, a method that balances hardware resource consumption, output image quality, and throughput performance is still lacking. Summary of the Invention

[0010] To address the problems existing in the prior art, this invention provides a super-resolution system and method for high-speed image acquisition, which improves processing efficiency and reduces resource consumption.

[0011] This invention is achieved through the following technical solution:

[0012] A super-resolution system for high-speed image acquisition includes: a padding module, a linebuffer module, a Bicubic top-level calculation module, and a shift register;

[0013] The Padding module is used to pad the original image and output the padding image to the Linebuffer module;

[0014] The Linebuffer module is used to buffer the supplementary image lines and output four lines of data synchronously in sequence.

[0015] The Bicubic top-level calculation module receives four rows of data output from the Linebuffer module. It calculates the intermediate one-dimensional interpolation results column by column through a vertical window and temporarily stores them in a shift register. After the four columns are calculated, it calculates the one-dimensional interpolation of the intermediate one-dimensional interpolation results row by row through a horizontal window, obtains the pixel values ​​of the interpolation points, and outputs them.

[0016] Preferably, the Bicubic top-level computing module includes three single-channel modules: an R-channel module, a G-channel module, and a B-channel module; each single-channel module includes a Bicubic module.

[0017] The top-level Bicubic calculation module receives four rows of data output from the Linebuffer module and sends the R, G, and B pixel values ​​of the four rows of data to the R channel module, G channel module, and B channel module, respectively. The Bicubic module of each single channel module calculates the intermediate one-dimensional interpolation results column by column through a vertical window and temporarily stores them in a shift register. After the four columns are calculated, the one-dimensional interpolation of the intermediate one-dimensional interpolation results is calculated row by row through a horizontal window to obtain the R, G, and B pixel values ​​of the interpolation points. The R, G, and B pixel values ​​of the interpolation points are then concatenated to obtain the interpolation point pixel value and output.

[0018] Furthermore, the Linebuffer module provides an enable signal to enable the Bicubic module to perform calculations.

[0019] Furthermore, the Bicubic module includes the cal_B module, cal_q module, shift_reg module, cal_B_post module, and cal_q_post module;

[0020] The cal_B module receives four lines of data, performs multiplication by two, subtraction, and addition operations in sequence, and outputs four intermediate coefficients to the cal_q module.

[0021] The cal_q module receives four intermediate coefficients from the cal_B module, passes them sequentially through multiplication operations and a two-level addition tree, and outputs the one-dimensional interpolation intermediate result to the shift_reg module.

[0022] The shift_reg module is used to receive the one-dimensional interpolation intermediate results output by the cal_q module and buffer them through a shift register, so as to output the four-point one-dimensional interpolation intermediate results to the cal_B_post module in parallel.

[0023] The cal_B_post module receives the four-point one-dimensional interpolation intermediate results output by the shift_reg module, performs multiplication by two, subtraction, and addition operations in sequence to obtain four horizontal intermediate coefficients, and outputs them to the cal_q_post module.

[0024] The cal_q_post module receives the four horizontal intermediate coefficients output by the cal_B_post module, and sequentially performs multiplication operations and a two-level addition tree to obtain the two-dimensional interpolation result, namely the R, G, or B pixel values ​​of the interpolation point.

[0025] Furthermore, the cal_q module performs operations on the four intermediate coefficients of the received input with the pre-stored parameters. The first level completes the multiplication operation, and then the two-level addition tree is used to obtain the one-dimensional interpolation intermediate result.

[0026] Furthermore, the cal_q_post module performs operations on the four received horizontal intermediate coefficients and the pre-stored parameters. The first level completes the multiplication operation, and then the two-dimensional interpolation result is obtained through a two-level addition tree.

[0027] Four cal_q_post modules are configured, each with different pre-stored parameters. The four cal_q_post modules perform parallel calculations and obtain the two-dimensional interpolation result of four adjacent points in one clock cycle.

[0028] A super-resolution method for high-speed image acquisition includes:

[0029] S1, supplement the original image to obtain the supplemented image;

[0030] S2, perform row buffering on the supplementary image to obtain four rows of data;

[0031] S3: Perform one-dimensional interpolation on the four rows of data column by column through the vertical window to obtain the intermediate one-dimensional interpolation result and temporarily store it. After the four columns are calculated, perform one-dimensional interpolation on the intermediate one-dimensional interpolation result row by row through the horizontal window to obtain the interpolation point pixel value.

[0032] Preferably, S1 specifically involves: assigning the pixel values ​​of the edges of the original image to the pixels that need to be supplemented in the original image to obtain the supplemented image.

[0033] Preferably, in S3, when performing one-dimensional interpolation, the interpolation points are inserted according to the rule of fixed intervals.

[0034] Preferably, in S3, the convolution kernel used for one-dimensional interpolation is as follows:

[0035]

[0036] Where |d| is the distance from the interpolation point to the reference point.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] This invention presents a super-resolution system based on a 16-point convolutional bicubic interpolation algorithm. The Bicubic top-level computation module first calculates intermediate one-dimensional interpolation results column-by-column using a vertical window, temporarily storing these results in a shift register, which slides right once per cycle. After four cycles, the horizontal window calculates the one-dimensional interpolation results row-by-row, using a shift register for caching and parallel output of the four-point intermediate interpolation results. This allows for efficient reuse of intermediate results during computation, maintaining consistent output data while reducing the parallelism of the vertical window to 1, thus saving hardware resources and reducing resource consumption to some extent. This invention is a high-frequency, low-power super-resolution system that converts low-resolution images received by a ground station into high-resolution images, meeting the high-precision monitoring requirements of unmanned aerial vehicle (UAV) systems operating at high speeds.

[0039] Furthermore, the constants used in the calculation (such as the distance d between two interpolation points) are pre-stored in the hardware, thereby avoiding a large number of redundant calculations.

[0040] The super-resolution method of this invention is based on a 16-point convolutional bicubic interpolation algorithm. It first calculates the intermediate one-dimensional interpolation results column by column through a vertical window, and temporarily stores the calculated intermediate one-dimensional interpolation results. It slides to the right once every cycle. After 4 cycles, the horizontal window calculates the one-dimensional interpolation results row by row. The intermediate one-dimensional interpolation results of the 4 points are output in parallel through caching, so as to efficiently reuse the intermediate results in the calculation process. It can keep the output data unchanged while reducing the parallelism of the vertical window to 1, thereby saving hardware resources and improving resource utilization to a certain extent.

[0041] Furthermore, in this invention, the pixel values ​​of the outermost edge of the original image are assigned to the values ​​to be filled, thereby maximizing the preservation of the interpolation points without distortion.

[0042] Furthermore, considering that direct 4x upsampling would cause additional resource consumption and accuracy loss due to the insertion of different numbers of points between reference points, this invention uses padding to evenly distribute the interpolation points between the reference points, and the relative distance between them and the reference points is a fixed number, i.e., |d| is fixed, which can reduce hardware consumption.

[0043] Furthermore, in the parameter selection process, this invention takes into account the errors and computational load that may be introduced by existing solutions, and selects α = -1 as the convolution kernel parameter to avoid the presence of decimals in the convolution kernel parameter, while avoiding pipeline delay deterioration caused by the introduction of a divider execution unit, thereby improving the pipeline throughput of convolution kernel calculation. Attached Figure Description

[0044] Figure 1 This is a schematic diagram of bicubic interpolation;

[0045] Figure 2 A diagram comparing different padding schemes;

[0046] Figure 3 This is a schematic diagram showing the distance between the interpolation point and the reference point;

[0047] Figure 4 This is a schematic diagram illustrating the simplification scheme for the interpolation formula;

[0048] Figure 5 A diagram showing the pixels that need to be added to the original image;

[0049] Figure 6 This is a schematic diagram of the padding process of the present invention;

[0050] Figure 7 This is a schematic diagram of the algorithm flow of the present invention;

[0051] Figure 8 A schematic diagram of the Bicubic module;

[0052] Figure 9 This is a schematic diagram of the top-level computing module structure of Bicubic;

[0053] Figure 10 This is a schematic diagram of a single-channel module structure;

[0054] Figure 11 Here is a schematic diagram of the cal_B module structure;

[0055] Figure 12 Here is a schematic diagram of the cal_q module structure;

[0056] Figure 13 Schematic diagram of horizontal and vertical windows;

[0057] Figure 14 This is a schematic diagram of the shift_reg module structure;

[0058] Figure 15 This is a schematic diagram of the cal_B_post module structure;

[0059] Figure 16 Here is a schematic diagram of the cal_q_post module structure;

[0060] Figure 17 This is a block diagram of the overall system of the present invention;

[0061] Figure 18 This is a schematic diagram of the test vector process;

[0062] Figure 19 The comparison results are shown for the C99 at the beginning of the first row of the upsampled test image and the simulation results of the post-implementation.

[0063] Figure 20 The comparison results are shown for the C99 at the end of the first row of the upsampled test image and the post-implementation simulation result.

[0064] Figure 21 Comparison of the C code at the beginning of the fifth line of the upsampled test image with the simulation results of post-implementation;

[0065] Figure 22 The comparison results are shown below: the C code at the end of the fifth line of the upsampled test image and the simulation results of the post-implementation.

[0066] Figure 23 The result is a display of the 4K image obtained by this invention.

[0067] Figure 24 This invention provides the resource usage and resource utilization rate.

[0068] Figure 25 For the system's clock constraints;

[0069] Figure 26 Analysis of setup and hold time margin after Bicubic IP placement and routing. Detailed Implementation

[0070] To further understand the present invention, the present invention will be described below with reference to embodiments. These descriptions are only for further explaining the features and advantages of the present invention and are not intended to limit the claims of the present invention.

[0071] This invention comprises three parts: an algorithm, a hardware, and a system.

[0072] Algorithm section: The one-dimensional bicubic interpolation algorithm is improved. In the final improved algorithm, the multiplication operation is 75% and 47% of that of the original one-dimensional and two-dimensional bicubic interpolation algorithms, respectively.

[0073] Hardware aspects: Constants used in the calculations (such as the distance d between two interpolation points) are pre-stored in the hardware, thus avoiding a large amount of redundant computation. Furthermore, shift registers are used to efficiently reuse intermediate results during the calculation process, thereby reducing resource utilization to some extent. Finally, multi-stage pipelining is used in the bicubic interpolation implementation to improve computational efficiency.

[0074] System component: When building the system on the Zynq 7020 development platform, VDMA is used to interconnect the PL end where the computing module is located with the PS end where the DDR is located. The high bandwidth provided by VDMA provides the foundation for the subsequent implementation of video streaming.

[0075] The three parts will be introduced separately below.

[0076] I. The following technical solutions are adopted for the algorithm part:

[0077] This invention employs a 16-point convolutional bicubic interpolation algorithm for subsequent hardware implementation.

[0078] The bicubic interpolation algorithm can be viewed as a two-dimensional extension of the one-dimensional interpolation function. The one-dimensional interpolation function is shown in equation (1):

[0079]

[0080] Where k takes the values ​​0, 1, 2, 3, and A k x represents the pixel values ​​of four reference points adjacent to the point to be determined. k Here are the x-coordinates of four reference points, x is the x-coordinate of the point to be determined, and d... k β is the distance between two points, β is the convolution kernel, and g(x) is the interpolation output.

[0081] In this invention, the convolution kernel provided by R.Keys is selected, and the interpolation point pixel value is calculated using four adjacent points as references. The convolution kernel is as follows:

[0082]

[0083] Where d is the distance between the reference point and the point to be calculated. The parameter α is generally set to -0.5 or -0.75, but in this invention, the parameter α is set to -1. That is, the convolution kernel is a piecewise cubic multinomial.

[0084] The bicubic interpolation algorithm performs further interpolation based on the one-dimensional interpolation result. For example... Figure 1 As shown, the bicubic interpolation algorithm selects 16 adjacent reference points. The position of the selected reference points depends on the coordinates of the interpolation points, such that the interpolation point g(x,y) falls on the reference point A. 11 A 21 A 12 A 22 Within the enclosed square area, interpolation pixel values ​​are calculated sequentially in both the vertical and horizontal directions. The formula for calculating the interpolation pixel values ​​is given below:

[0085]

[0086] Where i and j take values ​​of 0, 1, 2, 3, and A j The pixel values ​​of four adjacent reference points in the vertical direction, y k and x k y and x are the ordinates and abscissas of the four reference points, respectively, and y and x are the ordinates and abscissas of the point to be determined.

[0087] That is, interpolation is performed sequentially in the y and x directions. Since formula (3) is a linear operation, changing the order of calculation in the two-dimensional direction will not affect the final result. Therefore, in practical applications, the calculation order can be determined by referring to indicators such as the space complexity and time complexity of the algorithm implementation.

[0088] In the four arithmetic operations of FPGA (Field Programmable Gate Array), division has a greater overhead in terms of both time and space than the other three. Therefore, the introduction of division operations should be reduced in the hardware design of the algorithm. From the perspective of the interpolation processing of the entire image, the 16 reference points required to solve the interpolation points need to be substituted into formula (3) in sequence. In this invention, the parameter α is set to -1 to avoid the presence of decimals in the convolution kernel parameters, and at the same time to avoid pipeline delay deterioration caused by the introduction of the divider execution unit, thereby improving the pipeline throughput of β(d) calculation. The convolution kernel β(d) with parameter α = -1 is selected as shown in formula (4):

[0089]

[0090] Where |d| is the distance from the interpolation point to the reference point.

[0091] Hardware implementations of bicubic interpolation algorithms can be divided into two categories: non-uniform structures and uniform structures. Non-uniform structures are used in fields such as image rotation and infinite scaling. Uniform structures are essentially a special case of non-uniform structures, applicable only when the interpolation point positions are fixed. Since this invention uses a fixed magnification factor, and considering reducing hardware consumption, a uniform structure is chosen. Taking 8-point one-dimensional quadruple interpolation as an example, the specific implementation is as follows... Figure 2 As shown in the figure. Without padding, to perform one-dimensional four-fold uniform interpolation, the number of points inserted between two reference points is 6 or 7, and the relative distances between the interpolation points and the reference points are different. This requires calculating the distance between each interpolation point and the reference point in the hardware. Furthermore, the different number of interpolations between adjacent reference points introduces an additional counter to control the calculation count of the next-level module, which also introduces additional resource consumption. The case with padding of 1 is similar. In this invention, padding of 2 is used. As shown in the figure, if this method is used, the interpolation points will be evenly distributed between the reference points, that is, 4 points are inserted at a fixed interval between two reference points. Assuming the distance between two reference points is 1, and |d| is the distance between the interpolation point and the adjacent reference point on the left, it is easy to see that |d| is equal to 1 / 8, 3 / 8, 5 / 8, and 7 / 8 respectively. Since the value of |d| is fixed, to avoid repeated calculations, |d| and |d| are pre-calculated. 2 , |d| 3 The result is calculated and assigned as a parameter in the Verilog code, i.e., a pre-stored parameter. Historically, compared to fixed-point arithmetic, the stricter timing constraints and slower processing speed of floating-point arithmetic in FPGAs have led FPGA designers to use fixed-point arithmetic whenever possible. Similarly, in this invention, fixed-point arithmetic is used. To ensure that the loss of accuracy in fixed-point calculations is within a tolerable range, this invention uses |d|, |d| 2 , |d| 3 The value is shifted 18 bits to the left and substituted into the calculation. Finally, the pixel value is shifted 18 bits to the right to restore the correct data.

[0092] This invention simplifies the interpolation algorithm and minimizes the use of DSP (hardware multiplier).

[0093] like Figure 3 As shown, assuming the distance between reference points is 1, the interpolation point g(x,y) and the reference point A... 11 The horizontal distance |d| is u. Then the interpolation point g(x,y) and the reference point A... 10 A 12 A 13 The distances |d| are (1+u), (1-u), and (2-u) respectively. According to equations (4) and (3), we have:

[0094]

[0095] I c0 I c1 I c2 I c3 Let g(x0,y0) and reference point A be the coordinates of the coordinates. 10 A 11 A 12 A 13 Substituting the horizontal distance between them into the inner terms of equation (3), we obtain the result of each component, sum them up, and simplify them to get:

[0096] g(u)=u×(u×(u×B4+B3)+B2)+B1 (6)

[0097] Further merging yields:

[0098] g(u)=u 3 ×B4+u 2 ×B3+u×B2+B1 (7)

[0099] in:

[0100]

[0101] For the original one-dimensional bicubic interpolation algorithm, the formula for calculating the pixel value of the interpolation point is given by equation (7), and the hardware structure diagram is as follows. Figure 4 As shown in (a), it is easy to see that this structure diagram requires 5 fixed-point multipliers and 3 adders. Factoring out the common factor of formula (7) yields formula (6). Figure 4 (b) shows the optimized structure diagram, where the multiplication operation is reduced to 75% of the original. Furthermore, since the padding method of this invention allows only four fixed values ​​for the distance between the interpolation point and the reference point, this characteristic can be utilized to... 2 u 3 Pre-store, and then reuse formula (7) to obtain Figure 4 The hardware structure shown in (c) has the advantage of avoiding the accumulated error introduced by using multi-stage fixed-point multipliers. Meanwhile, compared to... Figure 4 The hardware shown in (b) utilizes a two-level addition tree to reduce the original pipeline stages from six to three, and compared to... Figure 4 As shown in (a), this optimization scheme also reduces the amount of multiplication operations to 75%.

[0102] The two-dimensional bicubic interpolation algorithm is an extension of one-dimensional interpolation in terms of spatial dimension, and its final formula is given by equation (9):

[0103] g(v)=v 3 ×C4+v 2 ×C3+v 2×C2+C1 (9)

[0104] in:

[0105]

[0106] g i (u)=u 3 ×B i4 +u 2 ×B i3 +u 1 ×B i2 +B i1 (11)

[0107]

[0108] The values ​​of u and v are 1 / 8, 3 / 8, 5 / 8, and 7 / 8, respectively.

[0109] When performing interpolation calculations, the starting position of the interpolation is as follows: Figure 5 As shown, the circular dots represent pixels in the original image. When calculating pix_d(0,0), its reference point in the original image and the 4×4 pixel window used for interpolation are pix_re(0,0) and pix_s 4×4, respectively. Therefore, it can be seen that some additional pixels need to be added at this point. Figure 5 Interpolation can only be performed using square points. The process of padding is also called padding.

[0110] In this invention, the padding process is as follows: Figure 6 As shown, this invention does not choose to simply fill with 0, but instead assigns the pixel values ​​of the outermost edge of the original image to the values ​​that need to be filled, thereby maximizing the preservation of the interpolation points without distortion.

[0111] The algorithm flowchart of this invention is shown below. Figure 7 As shown.

[0112] The algorithm includes the following steps:

[0113] S1, Padding the original image, specifically: adding pixels around the original image, preferably assigning the pixel values ​​of the original image edges to the pixels that need to be added to the original image, to obtain the supplemented image;

[0114] S2, the supplementary image after padding is buffered by a line buffer and outputs four lines of data synchronously in sequence;

[0115] S3: The four rows of data output synchronously are fed into the Bicubic top-level calculation module in parallel. The Bicubic top-level calculation module is obtained by instantiating the Bicubic module three times, and processes the pixel values ​​of the R, G, and B channels in parallel. The Bicubic module is divided into a vertical window and a horizontal window; the vertical window receives the four rows of data sent in parallel by the Linebuffer, calculates the intermediate one-dimensional interpolation result column by column, and temporarily stores the calculated intermediate one-dimensional interpolation result in a shift register, sliding to the right once per cycle; after 4 cycles, the horizontal window calculates the one-dimensional interpolation value of the intermediate one-dimensional interpolation result row by row and outputs it. This value is the final single-channel pixel value of the interpolation point, and the horizontal window slides to the right once per cycle. The Bicubic top-level calculation module stitches the pixel values ​​of the R, G, and B channels of the interpolation point bit by bit and outputs them.

[0116] S4: Determine if all image interpolation points have been calculated. If the calculation is complete, proceed to the next step; otherwise, wait for completion.

[0117] S5 performs operations such as image data integration and adding BMP file headers.

[0118] I. The hardware component adopts the following technical solution:

[0119] The hardware of the bicubic interpolation algorithm of this invention includes a Linebuffer module and a Bicubic top-level calculation module. The Bicubic top-level calculation module includes three single-channel modules, namely an R-channel module, a G-channel module, and a B-channel module. Each single-channel module includes a Bicubic module.

[0120] The Linebuffer module is used to buffer the padding supplementary image and output four lines of data sequentially and synchronously to the Bicubic top-level calculation module.

[0121] like Figure 8 As shown, the Bicubic module consists of a vertical window and a horizontal window. Essentially, they are structurally similar hardware modules used in different dimensions (horizontal and vertical). Square dots represent pixels that the Linebuffer module is currently outputting or will output; circular dots represent points output by the Linebuffer module in past clock cycles; triangular dots represent intermediate one-dimensional interpolation results from the vertical window; and diamond dots represent the final calculated interpolated pixel values.

[0122] It is worth mentioning that, because the horizontal and vertical windows are separated, the data output by the Linebuffer module can be directly processed in the vertical window, eliminating the need to cache 16 raw pixels and reducing hardware consumption and waiting time.

[0123] The Bicubic top-level calculation module of the present invention, as follows: Figure 9As shown, the start / stop and calculation mode selection of this Bicubic top-level calculation module are controlled by the upper-level module. The input is the RGB pixel values ​​of four lines of data, which are sent to the corresponding single-channel modules (R channel module, G channel module, B channel module) for calculation. After the calculation is completed, the values ​​are combined into four RGB pixel values ​​and output to the next level.

[0124] The clock and reset signals provide synchronization and reset for the Bicubic top-level calculation module. The enable signal is provided by the previous-level Linebuffer module to start the Bicubic module for computation. Line0_pixel to line4_pixel represent the four lines of RGB data of the image to be interpolated. To reduce the computational load required for the processor in the FPGA (Field-Programmable Gate Array) development board to recover data from DDR3 memory, this design intends to access the original image in DDR3 four times. The calculation results are output by traversing the rows, outputting four interpolation points adjacent to each row of data in the order of 1, 5, ..., 2157, 2, 6, ..., 2158, 3, 7, ..., 2159, 4, 8, ..., 2160. The sel signal is used to select the internal parameter data, choosing which row of data to calculate and output. vld_out is used to enable the next-level module to correctly receive valid data.

[0125] The single-channel module consists of eight sub-modules, which calculate the pixel values ​​of four connected points in each row. For example... Figure 10 As shown, it includes the cal_B module, cal_q module, shift_reg module, cal_B_post module, and four cal_q_post modules.

[0126] cal_B module description:

[0127] The structural diagram of this module is as follows: Figure 11 As shown, the input is the pixel values ​​of the original image. First, a shift operation is used to perform a multiplication by two. Then, subtraction and addition operations are performed in the second and third pipeline stages to finally obtain intermediate coefficients. The specific definitions of each signal in the cal_B module are given in Table 1.

[0128] Table 1. Specific definitions of each signal in the cal_B module.

[0129] clk 1 unsigned clock signal en 1 unsigned Input enable signal rst 1 unsigned Reset signal line0_pixel 8 unsigned The first line of raw data provided by linebuffer line1_pixel 8 unsigned The second line of raw data provided by linebuffer line2_pixel 8 unsigned The third line of raw data provided by linebuffer line3_pixel 8 unsigned The fourth line of raw data provided by linebuffer vld_out 1 unsigned Output enable signal b1 8 unsigned Output intermediate coefficient b1 b2 8 unsigned Output intermediate coefficient b2 b3 8 unsigned Output intermediate coefficient b3 b4 8 unsigned Output intermediate coefficient b4

[0130] cal_q module description:

[0131] The structural diagram of this cal_q module is shown below. Figure 12As shown, the input consists of four intermediate coefficients. A data selector selects pre-stored parameters for computation. The first level performs multiplication, and then a two-level addition tree yields the one-dimensional interpolation intermediate result q. The specific definitions of each signal in the cal_q module are given in Table 2.

[0132] Table 2 defines the specific signals in the cal_q module.

[0133] clk 1 unsigned clock signal en 1 unsigned Input enable signal rst 1 unsigned Reset signal sel 2 unsigned The selection signal is used to select the parameters to be used in the calculation. b1 8 unsigned intermediate coefficient b1 b2 8 unsigned intermediate coefficient b2 b3 8 unsigned intermediate coefficient b3 b4 8 unsigned intermediate coefficient b4 vld_out 1 unsigned Output enable signal q 23 signed Intermediate result q, 13 integer digits, 9 decimal digits

[0134] Description of the shift_reg module:

[0135] This `shift_reg` module implements the horizontal sliding window functionality. For example... Figure 13 As shown, x0-x7 are 8 points upsampled in a certain row, and q0-q4 are the calculation results of the vertical window. Specifically, q0-q3 in horizontal window 1 are used to calculate the upsampled points x0-x3; and q1-q4 in horizontal window 2 are used to calculate the upsampled points x4-x7. Therefore, it can be concluded that the vertical window calculation result q can be reused when calculating pixel values ​​at different positions (as shown in horizontal window 1 and horizontal window 2). Thus, a shift register is used for caching, and the intermediate q values ​​of the 4-point one-dimensional interpolation are output in parallel to provide input for the next-level cal_B_post module. In this way, the parallelism of the vertical window can be reduced to 1 while maintaining the output data unchanged, thereby saving hardware resources.

[0136] The schematic diagram and signal definitions of the shift_reg module are provided by Figure 14 As shown in Table 3.

[0137] Table 3 defines the specific signals in the shift_reg module.

[0138] clk 1 unsigned clock signal en 1 unsigned Input enable signal rst 1 unsigned Reset signal q 23 signed Input intermediate result q, a 13-digit integer and a 9-digit decimal. q0 8 signed Output the intermediate result q0, a 13-bit integer and a 9-bit decimal. q1 8 signed Output the intermediate result q1, a 13-digit integer and a 9-digit decimal. q2 8 signed Output the intermediate result q2, a 13-digit integer and a 9-digit decimal. q3 8 signed Output the intermediate result q3, a 13-digit integer and a 9-digit decimal. vld_out 1 signed Output enable signal

[0139] Description of the cal_B_post module:

[0140] This `cal_B_post` module is used to calculate the lateral intermediate coefficients. The input is the 4-point one-dimensional interpolation intermediate result `q` output by the `shift_reg` module. First, a multiplication by two operation is performed through a shift operation. Then, subtraction and addition operations are implemented in the second and third pipeline stages, ultimately yielding four lateral intermediate coefficients as output. The module structure diagram and signal definitions are provided by [the relevant documentation / component]. Figure 15 As shown in Table 4.

[0141] Table 4 Signal Definitions for the cal_B_post Module

[0142]

[0143] Description of the cal_q_post module:

[0144] The structural diagram of this module is as follows: Figure 16 As shown, the input consists of four horizontal intermediate coefficients. The calculation is performed using pre-stored parameters. The first level performs the multiplication operation, and then the two-dimensional interpolation result q is obtained through a two-level addition tree.

[0145] It's worth noting that there are four similar modules, `cal_q_post`, each with different pre-stored parameters. Through parallel computation, the interpolation result of four adjacent points is obtained within one clock cycle. Since the software algorithm may produce negative results or values ​​greater than 255 (the software assigns 255 to values ​​greater than 255 and 0 to values ​​less than 0), hardware also needs to handle this judgment and rounding. First, the sign bit is used to determine if the number is negative; if so, the output result is set to zero. Then, the bit above the highest integer bit is used to check for overflow; if overflow occurs, the output result is set to 255. For general output results, since the calculation involves decimal places, hardware rounding is required to obtain the final pixel value. When the integer part is less than 255 and the decimal part is greater than 0.5, the integer part is incremented by one, and the output is the final result. In other cases, the integer part is used directly as the final result.

[0146] The signal definitions for the cal_q_post module are given in Table 5.

[0147] Table 5 Signal Definitions for the cal_q_post Module

[0148]

[0149] III. The system adopts the following technical solutions:

[0150] The overall system block diagram of this invention is as follows: Figure 17 As shown, it is mainly divided into data flow and control flow.

[0151] The data stream can be divided into an upsampling processing section and an HDMI output section:

[0152] Upsampling Processing: The PS end reads the BMP image, discards the file header and performs other preprocessing, and stores it in a designated address in DDR3 memory. VDMA0 enables the read channel, reads the raw image data from the aforementioned address in DDR3, and sends the data stream to the Linebuffer IP core via the AXI4-stream interface. The Linebuffer IP core buffers 3 lines and synchronizes with the downlink data, outputting a total of 4 lines of data. The Bicubic IP core processes the data received from the Linebuffer IP core and outputs the 1st, 5th, 9th... lines of the upsampled image sequentially through the AXI4-stream interface. VDMA1 enables the write channel, receives the data stream processed by the Bicubic IP core, and writes it back to another designated address in DDR. The above steps are repeated 4 times. The difference from the first time is that in the second time, the Bicubic IP core outputs the 2nd, 6th, 10th... lines of the upsampled image, and so on. The PS end reads and tunes the upsampled output data, performs BMP post-processing such as adding a file header, names the file, and writes it to the SD card.

[0153] HDMI output section: VDMA2 enables the read channel, reads the image data after PS tuning, and sends it to the AXI-Stream to Video Out IP core via the AXI4-stream interface. Under the control of the VideoTiming Controller IP core, the AXI-Stream to Video Out IP core converts the AXI4-Stream format data into RGB888 format and sends it to the DVI Transmitter to drive the HDMI interface.

[0154] The control flow involves the host computer interacting with the PS terminal via UART. This includes sending initial coordinate information from the host computer to the PS terminal for image display, and sending a completion signal from the PS terminal to the host computer for writing the BMP image to the SD card. The GP interface connects to the peripheral configuration interface via AXIInterconnect, enabling the PS terminal to control peripherals on the PL terminal. The PS terminal configures the frame buffer space address and size of the VDMI IP core, as well as the read / write channels, and initializes and configures the output timing parameters of the Video Timing Controller IP core to control the start or stop of display.

[0155] Specific implementation examples

[0156] Hardware simulation verification:

[0157] The purpose of the simulation verification is to verify whether the hardware module in this valve can correctly implement the super-resolution algorithm designed in this invention and whether the results are completely consistent with the C code execution results. Therefore, in this test vector, the main implementation is to transmit the pixels in the 1K original image as stimuli to the hardware module and check whether there is a difference between the output pixels and the results of the C code execution.

[0158] In the control signal section of the test vector, the clk signal provides the clock signal for the entire hardware, so it needs to be assigned a value cyclically; the rst signal provides a reset, so it is 0 during initialization and then pulled high after 21ns to reduce the impact of system startup on the calculation results. en provides the enable signal for the entire system, so it is also 0 during initialization and then pulled high after 21ns.

[0159] In the data signal section of the test vector, firstly, MATLAB is used to convert the pixel values ​​in the 1K image to hexadecimal format and store them in a txt file. Then, the readmemh function is used to store the pixel values ​​from the txt file into pixel0_mem, pixel0_mem, pixel0_mem, and pixel0_mem. After the en signal goes high, the values ​​from these four memories are sequentially written to the hardware data input interface, as detailed below. Figure 18 As shown.

[0160] The upsampled image size is 3840×2160. Comparing the entire image output by the hardware with the C code would be quite complex. Therefore, we randomly select a few rows of data from the upsampled result and compare the data at the beginning and end of each row with the results in the C code. If the comparison results at the beginning and end of each row are the same, it proves that the calculation of the entire row is correct. Comparing the data from different rows can basically confirm that the calculation result of the entire image is correct.

[0161] For the test image, MATLAB reads the upsampled image generated by the C code and converts it to hexadecimal format. Simultaneously, the original test image is written to the testbench, followed by a post-implementation timing simulation of the hardware. A comparison of the C code at the beginning and end of the first line of the upsampled test image with the post-implementation simulation results is shown below. Figure 19 and Figure 20 As shown. Figure 19In the diagram, `vld_out_all` represents the valid output data signal. As can be seen from the graph, after `vld_out_all` is raised, there are still three cycles of invalid data, and in the fourth cycle, the first two data points are also invalid. Only then does the correct upsampling result appear. Based on the results marked in the boxes in the graph, it can be concluded that the data at the beginning of the first line of the simulation results for the super-resolution image obtained by the C code in the post-implementation is consistent. And... Figure 20 In the output data of the last cycle of the first line, there are still two invalid data points. Apart from this, the results of the C code execution are consistent with the hardware simulation results.

[0162] A comparison of the C code at the beginning and end of the fifth line of the upsampled test image with the post-implementation simulation results. Figure 21 and Figure 22 As shown in the figure. From the results in the figure, it can be seen that, excluding... Figure 19 and Figure 20 Invalid data in a similar format. The results obtained from the C code and hardware simulation are also basically consistent.

[0163] The simulation results above demonstrate that the hardware implemented in this invention can achieve the upsampling function. Furthermore, the output results are consistent with the results of running the C code.

[0164] System verification:

[0165] During system verification, firstly, the 1K image requiring super-resolution is copied to an SD card, and then the SD card is inserted into the development board. After powering on the development board, the program is downloaded to the board, and the hardware implementation of super-resolution begins. The calculated pixel values ​​of the 4K image are transferred from the PL terminal (where the calculation module is located) to the PS terminal via VDMA for writing to the SD card and HDMI display. The displayed result is as follows... Figure 23 As shown, the image displayed on the monitor on the left is the super-resolution 4K image. Due to the limitations of the monitor's resolution, a 4K image is displayed four times.

[0166] To comprehensively evaluate the algorithm design and hardware implementation, this invention divides the performance evaluation into two processes: image quality evaluation and hardware evaluation. Image quality evaluation primarily assesses the quality of the super-resolution image, using three metrics: PSNR, SSIM, and L obtained from the LPIPS model. 2 Distance. Hardware evaluation is divided into three parts: resource usage, operating frequency, and system latency.

[0167] Image quality assessment:

[0168] The image quality evaluation metrics selected in this invention are: PSNR, SSIM, and L obtained from the LPIPS model. 2 Distance is one of the metrics used in image similarity assessment. PSNR (Peak Signal Noise Ratio) is used to evaluate the mean square error between the original and processed images. SSIM (Structural Similarity) is a metric that measures the similarity between two images. LPIPS (Learned Perceptual Image Patch Similarity) uses deep learning to extract features and determine the similarity between two images. All three metrics are widely used in image similarity evaluation.

[0169] To facilitate a more intuitive evaluation of the quality of the super-resolution image, a certain number of 4K images were selected and compressed to 1K using an average method. Then, super-resolution processing was performed on them, and the processed 4K image was compared with the original image to determine the quality of the image processed by the algorithm.

[0170] Therefore, comparing the 4K image with the upsampling results implemented in C code, and considering PSNR, SSIM, and L... 2 The average performance of the algorithm implemented in this invention for distance is 31.76, 0.867 and 0.268, respectively.

[0171] Hardware evaluation:

[0172] For the hardware evaluation process, the implemented hardware system undergoes synthesis and implementation to obtain the logic resource usage. Next, master clock constraints are applied, and the highest operating frequency of the system is obtained under the conditions of establishing and maintaining time margins, and ensuring VDMA read / write bandwidth meets requirements. The clock constraint and time margin results are derived from... Figure 24 , Figure 25 The following is given. Finally, based on the system's operating frequency and the cycle time of the complete image, the delay for the system to upsample a single image can be obtained.

[0173] After implementing the hardware, the resource usage of the complete system, including IPs such as VDMA and Video Timing Controller, is obtained as follows: Figure 26 As shown.

[0174] Taking a single image as an example, the total delay required for upsampling can be obtained from formula (8). Wherein, Latency single N is the upsampling delay for a single image. processwith f max These represent the number of cycles required for a single image and the maximum operating frequency of the system, respectively.

[0175]

[0176] Based on the hardware module design process, the processing cycle required for the hardware system of this invention to process a single image is:

[0177] N process =N row ×Num row +t linebuffer =964×2160+12×960=2,093,760 (9)

[0178] Where, N row The number of processing steps required for upsampling each row; Num row t is the number of rows in the upsampled image; buffer This is the time required for row caching.

[0179] Therefore, the total delay for upsampling a single image can be derived as:

[0180]

[0181] Therefore, the total system latency is 0.013s, and the theoretical frame rate is 76FPS, which can be applied to high-speed image acquisition by UAVs.

[0182] A comprehensive comparison of this invention (Bicubic IP) with other mainstream hardware implementations yields the following table:

[0183] Table 6. Comparison of this invention with other mainstream hardware.

[0184]

[0185] It is easy to see that, while having a higher clock frequency, the output image quality and hardware logic resource usage of this invention have certain advantages compared with other implementation methods.

[0186] The function achieved by this invention is as follows: based on the Xilinx Zynq7020 development platform, a single 1K image with a lower resolution is processed by a super-resolution algorithm to obtain a 4K image with a higher resolution and displayed using HDMI, while the obtained 4K image output result is written to an SD card.

[0187] This invention first considers the potential errors and computational overhead introduced by existing schemes during parameter selection, choosing α = -1 as the convolution kernel parameter. Secondly, considering that direct 4x upsampling would lead to additional resource consumption and accuracy loss due to the insertion of different numbers of points between reference points, this invention performs a padding operation on the image. Padding ensures that the interpolation points are evenly distributed among the reference points, with the relative distance between them being four fixed finite decimals. Furthermore, since there are only four distance parameters, the formula can be further simplified to reduce pipeline stages and decrease accumulated error or data bit width. Considering the problems of increased area overhead, higher power consumption, and reduced computation speed caused by floating-point arithmetic in existing schemes, this invention uses fixed-point units for computation. Through parameter selection, padding image processing, and formula simplification, this invention can obtain interpolation results very close to floating-point results using fixed-point arithmetic with a narrower bit width. Similarly, addressing the issue of existing solutions' over-reliance on external storage, this invention employs a pipelined parallel architecture. The Bicubic IP contains no additional internal memory; the input is provided by a 4×1 vertical window generated by the Linebuffer, receiving four pixel values ​​per clock cycle. Furthermore, this invention explores the possibility of data multiplexing within the IP. Since the distance parameter is fixed, by using a shift register to generate a horizontal window, this invention optimizes the calculation of intermediate results from four times per clock cycle to once, reducing the negative impact of redundant computation. Additionally, by rationally arranging the connection order of multipliers and adders, this invention reduces the number of pipeline stages and the required data bit width, achieving a balance between computational accuracy and resource consumption.

Claims

1. A super-resolution system for high-speed image acquisition, characterized in that, include: Padding module, Linebuffer module, Bicubic top-level calculation module and shift register; The Padding module is used to pad the original image and output the padding image to the Linebuffer module; The Linebuffer module is used to buffer the supplementary image lines and output four lines of data synchronously in sequence. The Bicubic top-level calculation module receives four rows of data output from the Linebuffer module. It calculates the intermediate one-dimensional interpolation results column by column through a vertical window according to the rule of inserting interpolation points at fixed intervals and temporarily stores them in a shift register. After the four columns are calculated, it calculates the one-dimensional interpolation of the intermediate one-dimensional interpolation results row by row through a horizontal window according to the rule of inserting interpolation points at fixed intervals, obtains the pixel value of the interpolation point, and outputs it. The Bicubic top-level computation module consists of three single-channel modules: the R-channel module, the G-channel module, and the B-channel module; each single-channel module includes a Bicubic module. The top-level Bicubic calculation module receives four lines of data output from the Linebuffer module and sends the R, G, and B pixel values ​​of the four lines of data to the R channel module, G channel module, and B channel module, respectively. The Bicubic module of each single channel module calculates the intermediate one-dimensional interpolation results column by column through a vertical window according to the rule of inserting interpolation points at fixed intervals and temporarily stores them in a shift register. After the four columns are calculated, the one-dimensional interpolation of the intermediate one-dimensional interpolation results is calculated row by row through a horizontal window according to the rule of inserting interpolation points at fixed intervals, and the R, G, and B pixel values ​​of the interpolation points are obtained respectively. The R, G, and B pixel values ​​of the interpolation points are concatenated to obtain the interpolation point pixel value and output. The Bicubic module includes the cal_B module, cal_q module, shift_reg module, cal_B_post module, and cal_q_post module; The cal_B module receives four lines of data, performs multiplication by two, subtraction, and addition operations in sequence, and outputs four intermediate coefficients to the cal_q module. The cal_q module receives four intermediate coefficients output by the cal_B module, inserts interpolation points at fixed intervals, and then passes them through multiplication operations and a two-level addition tree to obtain a one-dimensional interpolation intermediate result, which is then output to the shift_reg module. The shift_reg module is used to receive the one-dimensional interpolation intermediate results output by the cal_q module and buffer them through a shift register, so as to output the four-point one-dimensional interpolation intermediate results to the cal_B_post module in parallel. The cal_B_post module receives the four-point one-dimensional interpolation intermediate results output by the shift_reg module, performs multiplication by two, subtraction, and addition operations in sequence to obtain four horizontal intermediate coefficients, and outputs them to the cal_q_post module. The cal_q_post module receives the four horizontal intermediate coefficients output by the cal_B_post module, and sequentially performs multiplication operations and a two-level addition tree according to the rule of inserting interpolation points at fixed intervals to obtain the two-dimensional interpolation result, namely the R, G or B pixel value of the interpolation point; The cal_q_post module performs operations on the four received horizontal intermediate coefficients and pre-stored parameters. The first level completes the multiplication operation, and then the two-dimensional interpolation result is obtained through a two-level addition tree. There are four cal_q_post modules, each with different pre-stored parameters. The four cal_q_post modules are calculated in parallel, and the two-dimensional interpolation result of four adjacent points is obtained in one clock cycle.

2. The super-resolution system for high-speed image acquisition according to claim 1, characterized in that, The Linebuffer module provides an enable signal to start the Bicubic module for computation.

3. The super-resolution system for high-speed image acquisition according to claim 1, characterized in that, The cal_q module performs operations on the four intermediate coefficients of the received input with the pre-stored parameters. The first level completes the multiplication operation, and then the two-level addition tree obtains the one-dimensional interpolation intermediate result.

4. A super-resolution method for high-speed image acquisition, characterized in that, The super-resolution system for high-speed image acquisition based on claim 1 includes: S1, supplement the original image to obtain the supplemented image; S2, perform row buffering on the supplementary image to obtain four rows of data; S3: The four rows of data output synchronously enter the Bicubic top-level calculation module in parallel. The Bicubic top-level calculation module is obtained by instantiating the Bicubic module three times and processes the pixel values ​​of the R, G, and B channels in parallel. The Bicubic module is divided into a vertical window and a horizontal window. The vertical window receives the four rows of data sent in parallel by the Linebuffer module, performs one-dimensional interpolation column by column to obtain the intermediate one-dimensional interpolation result, and temporarily stores the obtained intermediate one-dimensional interpolation result in a shift register, sliding to the right once per cycle. After 4 cycles, the horizontal window performs one-dimensional interpolation row by row on the intermediate one-dimensional interpolation result to obtain the final interpolation point single-channel pixel value, and the horizontal window slides to the right once per cycle. The Bicubic top-level calculation module stitches the pixel values ​​of the R, G, and B channels of the interpolation point bit by bit and outputs them. When performing one-dimensional interpolation, the interpolation points are inserted according to the rule of fixed intervals. S4, determine whether all image interpolation points have been calculated. If the calculation is complete, proceed to the next step; otherwise, wait for completion. S5 performs image data integration and adds BMP file headers.

5. The super-resolution method for high-speed image acquisition according to claim 4, characterized in that, S1 specifically involves assigning the pixel values ​​of the edges of the original image to the pixels that need to be added to the original image, thus obtaining the supplemented image.

6. The super-resolution method for high-speed image acquisition according to claim 4, characterized in that, In S3, the convolution kernel used for one-dimensional interpolation is as follows: in, This is the distance from the interpolation point to the reference point.