Data acceleration processing method based on affine inverse transformation of FPGA

By dividing the result image into odd column blocks and even column blocks, using ping-pong operation and affine transform inverse matrix to optimize FPGA data processing, the problems of excessive FPGA data processing and insufficient BRAM space are solved, and fast and accurate data processing and transformation are achieved.

CN116132606BActive Publication Date: 2025-07-11ZHEJIANG DALI TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211640901.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2025-07-11
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

In the prior art, FPGAs take too long to perform data processing, and insufficient BRAM space leads to failure of affine reduction transformation.

Method used

The result image is divided into odd column blocks and even column blocks, and the preloading steps and addressing reading steps are performed through ping-pong operations. The affine transformation inverse matrix is used to determine the preloaded sub-blocks, and whether it is necessary to shrink according to the BRAM space size is determined, and data storage is optimized through the pre-reducing magnification parameters.

Benefits of technology

The rapidity of FPGA data processing is achieved, the reduction transformation in affine transformation is ensured successfully, the calculation accuracy is improved and data loss is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116132606B_ABST
    Figure CN116132606B_ABST
Patent Text Reader

Abstract

The present invention relates to a data acceleration processing method based on FPGA affine inverse transformation, belonging to the technical field of image processing, and solves the problems of excessive time consumption in data processing by FPGA and the inability to implement affine reduction transformation due to insufficient BRAM space in the prior art. The data acceleration processing method includes: dividing the result image into multiple sub-image blocks as filling sub-image blocks, and the filling sub-image blocks are divided into odd-column blocks and even-column blocks; performing a preloading step and an addressing and reading step on the odd-column blocks and the even-column blocks based on ping-pong operation; the preloading step includes: determining corresponding second preloading sub-image blocks in the original image according to the odd-column block / even-column block and the affine transformation inverse matrix, and storing the second preloading sub-image blocks into the corresponding BRAM space; the addressing and reading step includes: sequentially traversing each pixel point in the odd-column block / even-column block, and filling the pixel value of each pixel point. It realizes the fast processing of data by FPGA and the normal execution of affine reduction transformation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a data acceleration processing method based on affine inverse transformation of FPGA. Background Art

[0002] With the development of technology, people's demand for high-definition videos is becoming increasingly strong. Scaling low-resolution videos into high-definition videos has become a major issue. With the increasingly wide application of FPGA (Field Programmable Gate Array), implementing high-definition videos based on FPGA has gradually become the mainstream.

[0003] The affine transformation algorithm can achieve functions such as image rotation, translation, and scaling, and is widely used in many video image processing fields.

[0004] However, in the prior art, data reading and writing can only be independently performed for the random access BRAM (Block Random Access Memory) space, resulting in excessive time consumption when the FPGA processes data; at the same time, when facing image reduction transformation, a large amount of original image data needs to be loaded, and the BRAM space is not sufficient to place the required original image data, resulting in data execution failure. Summary of the Invention

[0005] In view of the above analysis, an embodiment of the present invention aims to provide a data acceleration processing method based on affine inverse transformation of FPGA, so as to solve the problems of excessive time consumption when the FPGA processes data in the prior art and the inability to implement affine reduction transformation due to insufficient BRAM space.

[0006] An embodiment of the present invention provides a data acceleration processing method based on affine inverse transformation of FPGA, and the data acceleration processing method includes:

[0007] Dividing the result image into multiple sub-blocks as filling sub-blocks, and the filling sub-blocks are divided into odd-column blocks and even-column blocks;

[0008] Performing a preloading step and an addressing reading step on the odd-column blocks and even-column blocks based on ping-pong operation;

[0009] The preloading step includes: determining corresponding second preloading sub-blocks in the original image according to the odd-column blocks / even-column blocks and the affine transformation inverse matrix, storing the second preloading sub-blocks corresponding to the odd-column blocks in the first BRAM space, and storing the second preloading sub-blocks corresponding to the even-column blocks in the second BRAM space;

[0010] The addressing and reading step includes: sequentially traversing each pixel point in the odd-column block / even-column block, determining the pixel value of each pixel point according to the inverse affine transformation matrix and the second pre-loaded sub-block in the first BRAM space / second BRAM space corresponding to the odd-column block / even-column block, and filling the pixel value of each pixel point.

[0011] Based on a further improvement of the above method, the determining of the corresponding second pre-loaded sub-block in the original image according to the odd-column block / even-column block and the inverse affine transformation matrix includes:

[0012] Determining a matte sub-block according to the odd-column block / even-column block and the inverse affine transformation matrix, and determining a first pre-loaded sub-block according to the matte sub-block; judging whether the size of the first pre-loaded sub-block is larger than the size of the first BRAM space / second BRAM space; if the judgment is yes, determining a pre-shrinking magnification parameter according to the first pre-loaded sub-block and the size of the first BRAM space / second BRAM space, and shrinking the first pre-loaded sub-block based on the pre-shrinking magnification parameter to obtain a second pre-loaded sub-block; if the judgment is no, taking the first pre-loaded sub-block as the second pre-loaded sub-block.

[0013] Based on a further improvement of the above method, the determining of the pixel value of each pixel point according to the inverse affine transformation matrix and the second pre-loaded sub-block in the first BRAM space or second BRAM space corresponding to the odd-column block / even-column block includes:

[0014] Combining the inverse affine transformation matrix and the pre-shrinking magnification parameter to determine the first coordinate corresponding to the coordinate of each pixel point in the odd-column block / even-column block in the first BRAM space / second BRAM space;

[0015] Determining the pixel value of each corresponding pixel point according to the first coordinate.

[0016] Based on a further improvement of the above method, the combining the inverse affine transformation matrix and the pre-shrinking magnification parameter to determine the first coordinate corresponding to the coordinate of each pixel point in the odd-column block / even-column block in the first BRAM space / second BRAM space includes:

[0017] Obtaining the coordinate of the corresponding pixel point in the original image according to the coordinate of each pixel point and the inverse affine transformation matrix, and multiplying the coordinate of the pixel point in the original image minus the starting coordinate of the first pre-loaded sub-block by the pre-shrinking magnification parameter to obtain the first coordinate.

[0018] Further improvement based on the above method, the preloading step and the addressing and reading step are performed on the odd-column blocks and the even-column blocks based on the ping-pong operation, including:

[0019] All the filling sub-blocks in the result image are arranged in the order from left to right and from top to bottom, and the preloading step and the addressing and reading step are performed on the odd-column blocks / even-column blocks based on the ping-pong operation;

[0020] When performing the preloading step on the odd-column blocks / even-column blocks, the addressing and reading step is simultaneously performed on the even-column blocks / odd-column blocks.

[0021] Further improvement based on the above method, the determination of the matte sub-block according to the odd-column blocks / even-column blocks and the inverse affine transformation matrix includes:

[0022] Determine the filling coordinates corresponding to the odd-column blocks / even-column blocks; the filling coordinates include the four filling vertex coordinates corresponding to the odd-column blocks / even-column blocks;

[0023] Based on the inverse affine transformation, determine the matte coordinates of the matte sub-block according to the four filling vertex coordinates and the inverse affine transformation matrix, and the matte coordinates include four matte vertex coordinates;

[0024] Determine the matte sub-block according to the four matte vertex coordinates.

[0025] Further improvement based on the above method, the determination of the first preloading sub-block according to the matte sub-block includes:

[0026] According to the four matte vertex coordinates of the matte sub-block, determine the maximum and minimum values of the matte sub-block on two coordinate axes;

[0027] According to the maximum and minimum values on the two coordinate axes, determine the four vertex coordinates of the first preloading sub-block:

[0028]

[0029] Wherein, A, B, C and D represent the four vertex coordinates of the first preloading sub-block, Xmax represents the maximum value of the matte sub-block on the X-axis, Xmin represents the minimum value of the matte sub-block on the X-axis, Ymax represents the maximum value of the matte sub-block on the Y-axis, and Ymin represents the minimum value of the matte sub-block on the Y-axis.

[0030] Further improvement based on the above method, the judgment of whether the size of the first preloading sub-block is greater than the size of the first BRAM space / second BRAM space includes:

[0031] Determine the length and width of the first pre-loaded sub-block and determine the length and width of the first BRAM space / second BRAM space;

[0032] Calculate the length ratio of the length of the first pre-loaded sub-block and the length of the first BRAM space / second BRAM space, and calculate the width ratio of the width of the first pre-loaded sub-block and the width of the first BRAM space / second BRAM space; the sizes of the first BRAM space and the second BRAM space are the same;

[0033] If any one of the length ratio and the width ratio is greater than 1, it is judged as yes, otherwise, it is judged as no.

[0034] Based on a further improvement of the above method, the determining the pre-shrinking magnification parameter according to the sizes of the first pre-loaded sub-block and the first BRAM space / second BRAM space includes:

[0035] Judge whether the length ratio is greater than the width ratio; if the judgment is yes, determine the pre-shrinking magnification parameter according to the length ratio; otherwise, determine the pre-shrinking magnification parameter according to the width ratio; the pre-shrinking magnification parameter is N / M, where N and M are both integers in [1, 6], and M > N.

[0036] Based on a further improvement of the above method, the shrinking the first pre-loaded sub-block based on the pre-shrinking magnification parameter to obtain a second pre-loaded sub-block includes:

[0037] Insert N - 1 pixel points between every two adjacent pixels in each row of the first pre-loaded sub-block, and then insert N - 1 pixel points between every two adjacent pixels in each column of the first pre-loaded sub-block, so as to expand the first pre-loaded sub-block by N times to obtain an intermediate sub-block;

[0038] In the intermediate sub-block, select every M - 1th row as a row to be sampled, and in each row to be sampled, select every M - 1th point as a sampling point, so as to shrink the intermediate sub-block by M times to obtain a second pre-loaded sub-block.

[0039] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:

[0040] 1. By determining whether the size of the original image data (the first pre-loaded sub-block) to be loaded is greater than the size of the BRAM space; if the determination is yes, the original image data is reduced according to the original image data and the size of the BRAM space, and the reduced original image data is saved to the BRAM space, realizing the normal saving of the original image data to the BRAM space and ensuring the successful execution of the reduction transformation in the affine transformation.

[0041] 2. By filling in the filling coordinates of the filling sub-blocks, the original image data to be loaded is determined based on the affine inverse transformation, improving the accuracy of calculating the original image data to be loaded.

[0042] 3. According to the length ratio and width ratio of the original image data to be loaded and the BRAM space, the pre-reduction magnification parameter is further determined, ensuring the reduction range when reducing the original image data and reducing the loss of the original image data.

[0043] 4. By performing the pre-loading step and the addressing and reading step on the odd-column blocks and even-column blocks into which the result image is divided through ping-pong operation, the FPGA realizes the fast processing of data and reduces the processing delay.

[0044] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can be made obvious from the description, or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the content specifically pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference numerals represent the same components.

[0046] Figure 1 It is a schematic flowchart of the data acceleration processing method based on the FPGA affine inverse transformation provided by the embodiment of the present invention;

[0047] Figure 2 It is a schematic structural diagram of dividing the result image into filling sub-blocks provided by the embodiment of the present invention;

[0048] Figure 3 It is one of the schematic structural diagrams of the cropping sub-block corresponding to the filling sub-block provided by the embodiment of the present invention;

[0049] Figure 4 It is the other schematic structural diagram of the cropping sub-block corresponding to the filling sub-block provided by the embodiment of the present invention;

[0050] Figure 5Schematic diagram for interpolating and enlarging the first pre-loaded sub-block in the embodiments of the present invention;

[0051] Figure 6 Schematic diagram for sampling and shrinking intermediate sub-blocks at every other point in the embodiments of the present invention;

[0052] Figure 7 One of the schematic diagrams of the first coordinate in the embodiments of the present invention;

[0053] Figure 8 Another schematic diagram of the first coordinate in the embodiments of the present invention. Detailed implementation manners

[0054] The preferred embodiments of the present invention will be specifically described below with reference to the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principle of the present invention, rather than to limit the scope of the present invention.

[0055] A specific embodiment of the steps of the present invention discloses a data acceleration processing method based on FPGA affine inverse transformation, as Figure 1 shown.

[0056] Step S1: Divide the result image into multiple sub-blocks as filling sub-blocks. The filling sub-blocks are divided into odd-column blocks and even-column blocks;

[0057] Step S2: Perform a pre-loading step and an addressing and reading step on the odd-column blocks and even-column blocks based on ping-pong operation; the pre-loading step includes: determining corresponding second pre-loaded sub-blocks in the original image according to the odd-column block / even-column block and the affine transformation inverse matrix, storing the second pre-loaded sub-blocks corresponding to the odd-column blocks in the first BRAM space, and storing the second pre-loaded sub-blocks corresponding to the even-column blocks in the second BRAM space; the addressing and reading step includes: sequentially traversing each pixel point in the odd-column block / even-column block, determining the pixel value of each pixel point according to the affine transformation inverse matrix and the second pre-loaded sub-blocks in the first BRAM space / second BRAM space corresponding to the odd-column block / even-column block, and filling the pixel value of each pixel point.

[0058] Specifically, when using FPGA for video image processing, it is necessary to store the original image pixel data from the DDR (Double Data Rate, double-speed synchronous dynamic random access memory) memory externally configured to the FPGA into the on-chip BRAM space of the FPGA, and then map the original image pixel data to the result image according to the affine transformation matrix. The result image is also saved in the DDR.

[0059] Specifically, in step S1, the result image is divided into multiple sub - image blocks, which are used as filling sub - image blocks. According to the rows and columns where the filling sub - image blocks are located, they can be divided into odd - column blocks and even - column blocks. After the size of each filling sub - image block is designed, it will not be changed anymore. Exemplarily, as Figure 2 shown, the size of the result image is 1280*1024. The length of the result image is 1280 pixels, and the width is 1024 pixels; the size of the filling sub - image block is 128*128. The length of the filling sub - image block is 128 pixels, and the width is 128 pixels; the result image is divided into 10 filling sub - image blocks in length and 8 filling sub - image blocks in width, and a total of 80 filling sub - image blocks are obtained by dividing the result image.

[0060] It can be understood that the two - dimensional coordinate system is used to locate all pixel points of the result image and the original image.

[0061] The inverse matrix of the affine transformation in the two - dimensional coordinate system includes:

[0062] 1. The 3*3 control matrix for two - dimensional translation is:

[0063] In this matrix, tx and ty are the pixel distances of translation, corresponding to the arrow directions of the X and Y coordinate axes. For the inverse transformation, its forward translation is exactly opposite to the arrow direction of the coordinate axis, so it is represented by a negative sign.

[0064] 2. The 3*3 control matrix for two - dimensional rotation is:

[0065] In this matrix, θ represents the rotation angle. In the inverse transformation, the sine parameter sign of this matrix is exactly opposite to that in the forward transformation.

[0066] 3. The 3*3 control matrix for two - dimensional scaling is:

[0067] In this matrix, Sx and Sy are the magnification factors. Since the magnification factor in the forward transformation is the reciprocal of the reduction factor in the inverse transformation, the magnification factors here exist in reciprocal form.

[0068] According to the inverse matrix of the affine transformation, the image coordinates of the result image and the original image are re - arranged. It should be noted that the pixel point coordinates of the result image are mapped to the coordinates of the original image, and then the pixel values on the coordinates of the original image are transferred to the pixel point coordinates of the result image. Its matrix operation expression is:

[0069]

[0070] Among them, the coordinates (u, v) are the coordinates of any pixel point in the result image, and the coordinates (x, y) are the corresponding coordinates in the original image obtained according to the inverse matrix of the affine transformation. is the inverse matrix of the affine transformation.

[0071] It should be noted that the image transformation is achieved by changing the parameters in the control matrix. Of course, multiple deformation parameters can also be mixed to achieve the purpose of performing multiple deformations simultaneously in one affine transformation process.

[0072] Specifically, in step S2, a preloading step and an addressing and reading step are performed on the odd-column blocks and even-column blocks based on the ping-pong operation; the preloading step is to store the pixel value data of the pixel points of the sub-blocks of the original image corresponding to the filling sub-blocks (odd-column blocks / even-column blocks) divided in the result image into the corresponding BRAM space; the addressing and reading step is to fill the pixel value data stored in the corresponding BRAM space into each pixel point in the filling sub-blocks (odd-column blocks / even-column blocks) divided in the result image, so that each pixel point in the filling sub-block is filled with the corresponding pixel value data.

[0073] Preferably, performing the preloading step and the addressing and reading step on the odd-column blocks and even-column blocks based on the ping-pong operation includes:

[0074] All the filling sub-blocks in the result image are arranged in the order from left to right and from top to bottom, and the preloading step and the addressing and reading step are performed on the odd-column blocks / even-column blocks based on the ping-pong operation;

[0075] When performing the preloading step on the odd-column blocks / even-column blocks, the addressing and reading step is simultaneously performed on the even-column blocks / odd-column blocks.

[0076] Specifically, as Figure 2 shown, the result image is divided into 10 filling sub-blocks in length and 8 filling sub-blocks in width. All the filling sub-blocks divided in the result image are performed in the order from left to right and from top to bottom. There are a total of 40 odd-column blocks and 40 even-column blocks in the result image. In the filling sub-blocks of the first row, the 1st, 3rd, 5th, 7th, and 9th filling sub-blocks are odd-column blocks, and the 2nd, 4th, 6th, 8th, and 10th filling sub-blocks are even-column blocks; similarly, in the filling sub-blocks of the second row, the 11th filling sub-block is an odd-column block, and the 12th filling sub-block is an even-column block, and so on. It should be noted that the execution order of all the filling sub-blocks in the result of the present invention can also be in other orders, as long as all the filling sub-blocks are executed in a certain order.

[0077] Exemplary:

[0078] In the first execution cycle, a preloading step is performed on the first tiling sub-block;

[0079] In the second execution cycle, a preloading step is performed on the second tiling sub-block, and an addressing and reading step is performed on the first tiling sub-block at the same time;

[0080] In the third execution cycle, a preloading step is performed on the third tiling sub-block, and an addressing and reading step is performed on the second tiling sub-block at the same time;

[0081] In the 80th execution cycle, a preloading step is performed on the 80th tiling sub-block, and an addressing and reading step is performed on the 79th tiling sub-block at the same time;

[0082] In the 81st execution cycle, an addressing and reading step is performed on the 80th tiling sub-block.

[0083] It should be noted that in the embodiment of the present invention, the preloading step and the addressing and reading step are performed on the tiling sub-blocks divided from the result image through ping-pong operation, which speeds up the data processing speed of the FPGA, realizes the fast processing of data by the FPGA, and reduces the processing delay.

[0084] Specifically, in the preloading step, according to the odd-column block / even-column block and the inverse affine transformation matrix, the corresponding second preloading sub-block of the odd-column block / even-column block is determined in the original image, the second preloading sub-block corresponding to the odd-column block is stored in the first BRAM space, and the second preloading sub-block corresponding to the even-column block is stored in the second BRAM space. Exemplarily, as Figure 2 shown, the second preloading sub-blocks corresponding to the 1st, 3rd, 5th, 7th... 79th tiling sub-blocks are stored in the first BRAM space, and the second preloading sub-blocks corresponding to the 2nd, 4th, 6th, 8th... 80th tiling sub-blocks are stored in the second BRAM space.

[0085] It should be noted that the first BRAM space and the second BRAM space are designed with the same size, and the first BRAM space and the second BRAM space are two independent BRAM spaces, both belonging to the on-chip BRAM space of the FPGA.

[0086] Preferably, the determining the corresponding second preloading sub-block in the original image according to the odd-column block / even-column block and the inverse affine transformation matrix includes:

[0087] Determine the matte sub - patch according to the odd - column block / even - column block and the inverse affine transformation matrix, and determine the first pre - loaded sub - patch according to the matte sub - patch; determine whether the size of the first pre - loaded sub - patch is greater than the size of the first BRAM space / second BRAM space; if the determination result is yes, determine the pre - scaling ratio parameter according to the size of the first pre - loaded sub - patch and the size of the first BRAM space / second BRAM space, and scale down the first pre - loaded sub - patch based on the pre - scaling ratio parameter to obtain the second pre - loaded sub - patch; if the determination result is no, use the first pre - loaded sub - patch as the second pre - loaded sub - patch.

[0088] Specifically, calculate the matte sub - patch according to the odd - column block / even - column block and the inverse affine transformation matrix. The matte sub - patch is a sub - patch in the original image and is the sub - patch in the original image corresponding to the filling sub - patch.

[0089] Preferably, the determining the matte sub - patch according to the odd - column block / even - column block and the inverse affine transformation matrix includes:

[0090] Determine the filling coordinates corresponding to the odd - column block / even - column block; the filling coordinates include the four filling vertex coordinates corresponding to the odd - column block / even - column block;

[0091] Based on the inverse affine transformation, determine the matte coordinates of the matte sub - patch according to the four filling vertex coordinates and the inverse affine transformation matrix. The matte coordinates include four matte vertex coordinates;

[0092] Determine the matte sub - patch according to the four matte vertex coordinates.

[0093] Specifically, as Figure 3 shown, the filling sub - patch is any sub - patch in the result image. The filling sub - patch can be an odd - column block or an even - column block. Combining with the two - dimensional coordinate system of the result image, four corresponding filling vertex coordinates Q1, Q2, Q3, and Q4 can be determined for the filling sub - patch.

[0094] Based on the inverse affine transformation, according to the four filling vertex coordinates Q1, Q2, Q3, and Q4, calculate and determine the corresponding four matte vertex coordinates Q1’, Q2’, Q3’, and Q4’ in the original image respectively through the inverse affine transformation matrix. As Figure 3 shown, the matte sub - patch is the sub - patch in the original image corresponding to the filling sub - patch.

[0095] It should be noted that for the burst length reading in the DDR externally configured by the FPGA, if the pixel data corresponding to the cropped sub-blocks is directly extracted from the DDR, a large amount of timing waste will be generated. Therefore, it is necessary to determine the first pre-loaded sub-block according to the cropped sub-blocks of the original image, and the first pre-loaded sub-block is the pixel value data that needs to be stored from the DDR to the BRAM space.

[0096] Preferably, the determining the first pre-loaded sub-block according to the cropped sub-blocks includes:

[0097] Determine the maximum and minimum values of the cropped sub-block on the two coordinate axes according to the four cropping vertex coordinates of the cropped sub-block;

[0098] Determine the four vertex coordinates of the first pre-loaded sub-block according to the maximum and minimum values on the two coordinate axes:

[0099]

[0100] Wherein, A, B, C, and D represent the four vertex coordinates of the first pre-loaded sub-block, Xmax represents the maximum value of the cropped sub-block on the X-axis, Xmin represents the minimum value of the cropped sub-block on the X-axis, Ymax represents the maximum value of the cropped sub-block on the Y-axis, and Ymin represents the minimum value of the cropped sub-block on the Y-axis.

[0101] Specifically, as Figure 3 shown, the four cropping vertex coordinates Q1’, Q2’, Q3’, and Q4’ have a maximum value of Xmax and a minimum value of Xmin on the X-axis; a maximum value of Ymax and a minimum value of Ymin on the Y-axis. The four vertex coordinates of the first pre-loaded sub-block are It should be noted that when reading the pixel value of a pixel point from the DDR, the coordinate of each pixel point is an integer.

[0102] It can be understood that the rectangle enclosed by the first pre-loaded sub-block includes the cropping area formed by the cropped sub-blocks, and the four vertex coordinates of the first pre-loaded sub-block are all integers.

[0103] Specifically, determine whether the size of the first pre-loaded sub-block is larger than the size of the first BRAM space / the second BRAM space; if the determination result is yes, determine a pre-shrinking magnification parameter according to the sizes of the first pre-loaded sub-block and the first BRAM space / the second BRAM space, and shrink the first pre-loaded sub-block based on the pre-shrinking magnification parameter to obtain a second pre-loaded sub-block; if the determination result is no, use the first pre-loaded sub-block as the second pre-loaded sub-block.

[0104] Specifically, the size of the BRAM space is generally designed to be 1.5 - 2 times the size of the tiling sub-block. Exemplarily, when the size of the tiling sub-block is 128 * 128, the size of the BRAM space is designed to be 256 * 256. After the size of the BRAM space is determined, it will not be changed anymore.

[0105] It should be noted that when shrinking the original image, when loading pixel data from the original image according to the tiling sub-block of the result image, it may be necessary to load a larger area of the original image, and the pre-determined BRAM space may not be sufficient to place the required pixel data of the original image, which may cause the data execution to fail. Therefore, before storing the first pre-loaded sub-block into the BRAM, the present invention compares the size of the first pre-loaded sub-block with the size of the BRAM space.

[0106] As Figure 3 shown, when the size of the first pre-loaded sub-block is larger than the size of the corresponding first BRAM space / the second BRAM space, the first pre-loaded sub-block is shrunk according to the sizes of the first pre-loaded sub-block and the first BRAM space / the second BRAM space to obtain a second pre-loaded sub-block, so that the first BRAM space / the second BARM space is larger than the second pre-loaded sub-block, enabling the second pre-loaded sub-block to be normally placed in the first BRAM space / the second BRAM space and ensuring the successful execution of the data.

[0107] As Figure 4 shown, when the size of the first pre-loaded sub-block is less than or equal to the size of the first BRAM space / the second BRAM space, the first pre-loaded sub-block is directly used as the second pre-loaded sub-block.

[0108] Preferably, the determination of whether the size of the first pre-loaded sub-block is larger than the size of the first BRAM space / the second BRAM space includes:

[0109] Determine the length and width of the first pre-loaded sub-block and determine the length and width of the first BRAM space / the second BRAM space;

[0110] Calculate the length ratio of the length of the first pre-loaded sub-block to the length of the first BRAM space / second BRAM space and calculate the width ratio of the width of the first pre-loaded sub-block to the width of the first BRAM space / second BRAM space; the sizes of the first BRAM space and the second BRAM space are the same;

[0111] If any one of the length ratio and the width ratio is greater than 1, it is judged as yes, otherwise, it is judged as no.

[0112] Specifically, the designed sizes of the first BRAM space and the second BRAM space are the same, and the sizes of the first BRAM space and the second BRAM space are fixed and unchanged, such as both being L*W; L is the length, usually 256 pixels, and W is the width, usually 256 pixels. The size of the first pre-loaded sub-block is calculated and determined according to the four vertex coordinates A, B, C, and D of the first pre-loaded sub-block. The length of the first pre-loaded sub-block is pixels, and the width is pixels.

[0113] Specifically, the length ratio of the length of the first pre-loaded sub-block to the length of the first BRAM space / second BRAM space is / L, and the width ratio of the width of the first pre-loaded sub-block to the width of the first BRAM space / second BRAM space is

[0114] Specifically, compare the length ratio / L and 1, compare the width ratio and 1. If any one of the length ratio and the width ratio is greater than 1, it is judged as yes, otherwise, it is judged as no.

[0115] Specifically, determine the pre-shrinking magnification parameter according to the sizes of the first pre-loaded sub-block and the first BRAM space / second BRAM space.

[0116] Preferably, the determining the pre-shrinking magnification parameter according to the sizes of the first pre-loaded sub-block and the first BRAM space / second BRAM space includes:

[0117] Judge whether the length ratio is greater than the width ratio; if the judgment is yes, determine the pre-shrinking magnification parameter according to the length ratio; otherwise, determine the pre-shrinking magnification parameter according to the width ratio; the pre-shrinking magnification parameter is N / M, where N and M are both integers in [1, 6], and M > N.

[0118] Specifically, judge the length ratio and the width ratio The size. If the length ratio is greater than the width ratio then it is judged as yes, and the pre-shrinking magnification parameter is determined according to the length ratio If the length ratio / L is less than or equal to the width ratio then it is judged as no, and the pre-shrinking magnification parameter is determined according to the width ratio to determine the pre-shrinking magnification parameter.

[0119] Preferably, the pre-shrinking magnification parameter is N / M, where both N and M are integers in [1, 6], and M > N.

[0120] It should be noted that when the length ratio is greater than the width ratio at this time, by taking the reciprocal of the length ratio, and then simplifying the reciprocal value into a fraction form, and making the fraction value as close as possible to the reciprocal value on the basis of not being greater than the reciprocal value, this fraction value is the pre-shrinking magnification parameter N / M, so that both N and M are integers in [1, 6], and M > N. Similarly, when the length ratio is less than or equal to the width ratio at this time, take the reciprocal of the width ratio, and then simplify the reciprocal value into a fraction form, and make the fraction value as close as possible to the reciprocal value on the basis of not being greater than the reciprocal value, this fraction value is the pre-shrinking magnification parameter N / M, so that both N and M are integers in [1, 6], and M > N.

[0121] Exemplarily, when the calculated length ratio is 1.6 and the width ratio is 1.2, the pre-shrinking magnification parameter determined according to the length ratio can be 3 / 5. Specifically, first take the reciprocal of the length ratio as 1 / 1.6 = 0.625, then select integers N and M such that N / M is not greater than 0.625 and as close as possible to 0.625. Therefore, N is selected as 3 and M is selected as 5, that is, the reciprocal value is simplified into the fraction 3 / 5, and the pre-shrinking magnification parameter is 3 / 5.

[0122] Preferably, the step of shrinking the first pre-loaded sub-block according to the pre-shrinking magnification parameter to obtain a second pre-loaded sub-block includes:

[0123] Insert N - 1 pixel points between two adjacent pixels in each row of the first pre-loaded sub-block. After that, insert N - 1 pixel points between two adjacent pixels in each column of the first pre-loaded sub-block, so as to expand the first pre-loaded sub-block by N times to obtain an intermediate sub-block;

[0124] In the middle sub-block, one row is selected as the row to be sampled every M - 1 rows. In each row to be sampled, one pixel is selected as the sampling point every M - 1 points, so as to reduce the middle sub-block by M times and obtain the second pre-loading sub-block.

[0125] Specifically, as Figure 5 shown, for example: when N is 2, one pixel is inserted between two adjacent pixels in each row of the first pre-loading sub-block. Then, one pixel is inserted between two adjacent pixels in each column of the first pre-loading sub-block, so as to expand the first pre-loading sub-block by 2 times and obtain the middle sub-block; when N is 3, two pixels are inserted between two adjacent pixels in each row of the first pre-loading sub-block. Then, two pixels are inserted between two adjacent pixels in each column of the first pre-loading sub-block, so as to expand the first pre-loading sub-block by 3 times and obtain the middle sub-block. The original pixels represent the pixels in the first pre-loading sub-block, and the interpolated pixels represent the inserted pixels.

[0126] Preferably, when inserting pixels in the first pre-loading sub-block, the expression of the pixel value of the interpolated pixel is:

[0127] Dix_j = DixJ - j[(DixJ - DixK) / N];

[0128] where DixJ and DixK respectively represent the pixel values of two adjacent pixels J and K in the first pre-loading sub-block, and Dix_j represents the pixel value of the j-th pixel among the N - 1 pixels inserted in sequence from pixel J to pixel K, 1 ≤ j ≤ (N - 1).

[0129] Specifically, as Figure 5 shown, for example: when N is 2, in the first row of the first pre-loading sub-block, one pixel is inserted between two adjacent original pixels. From left to right, the pixel values of two adjacent original pixels are 20 and 30 respectively, then the pixel value of the inserted pixel is 20 - 1*[(20 - 30) / 2], that is, the pixel value of the inserted one pixel is 25; when N is 3, in the first row, two pixels are inserted between two adjacent original pixels. From left to right, the pixel values of two adjacent original pixels are 20 and 30 respectively, then the pixel values of the inserted pixels are 20 - 1*[(20 - 30) / 3] and 20 - 2*[(20 - 30) / 3], that is, the pixel values of the inserted two pixels are 23 and 26.

[0130] Specifically, as Figure 5As shown, for example: when N is 2, in the first column of the first pre-loading sub-block, 1 pixel point is inserted between two adjacent original pixels. From top to bottom, the pixel values of two adjacent original pixels are 30 and 40 respectively. Then the pixel value of the inserted pixel point is 30 - 1 * [(30 - 40) / 2], that is, the pixel value of the inserted 1 pixel point is 35; when N is 3, in the first column of the first pre-loading sub-block, 2 pixel points are inserted between two adjacent original pixels. From top to bottom, the pixel values of two adjacent original pixels are 30 and 40 respectively. Then the pixel values of the inserted pixel points are 30 - 1 * [(30 - 40) / 3] and 30 - 2 * [(30 - 40) / 3], that is, the pixel values of the inserted 2 pixel points are 33 and 36.

[0131] Specifically, as Figure 6 shown, for example: when M is 2, in the middle sub-block, every other row is selected as the row to be sampled. In each row to be sampled, every other point is selected as a sampling point, so as to reduce the middle sub-block by 2 times to obtain the second pre-loading sub-block; when M is 3, in the middle sub-block, every two rows are selected as the row to be sampled. In each row to be sampled, every two points are selected as a sampling point, so as to reduce the middle sub-block by 3 times to obtain the second pre-loading sub-block. The adopted pixels represent the pixels of the points, and the discarded pixels represent the pixel points that are not sampled.

[0132] Compared with the prior art, the data acceleration processing method based on FPGA affine inverse transformation provided by the embodiment of the present invention performs a pre-loading step and an addressing and reading step on the filling sub-blocks divided from the result image through ping-pong operation, which speeds up the data processing speed of the FPGA, realizes the fast processing of data by the FPGA, and reduces the processing delay; by judging whether the size of the original image data to be loaded is greater than the size of the BRAM space; if the judgment result is yes, the original image data is reduced according to the size of the original image data and the BRAM space, and the reduced original image data is saved to the BRAM space, which solves the technical problem that the affine reduction transformation cannot be realized due to the too large original image and the too small BRAM space during the affine transformation process; based on the filling coordinates of the filling sub-blocks, the original image data to be loaded is determined based on the affine inverse transformation, which improves the accuracy of calculating the original image data to be loaded; according to the length ratio and width ratio of the original image data to be loaded and the BRAM space, the pre-reduction magnification parameter is further determined, so that when the original image data is reduced, the reduction range is ensured and the loss of the original image data is reduced.

[0133] In step S2, the addressing and reading step includes: sequentially traversing each pixel point in the odd column block / even column block, determining the pixel value of each pixel point according to the inverse affine transformation matrix and the second pre-loaded sub-block in the first BRAM space / second BRAM space corresponding to the odd column block / even column block, and filling the pixel value of each pixel point.

[0134] Preferably, the determining the pixel value of each pixel point according to the inverse affine transformation matrix and the second pre-loaded sub-block in the first BRAM space or the second BRAM space corresponding to the odd column block / even column block includes:

[0135] Combining the inverse affine transformation matrix and the pre-shrinking magnification parameter to determine the first coordinate corresponding to the coordinate of each pixel point in the odd column block / even column block in the first BRAM space / second BRAM space;

[0136] Determining the pixel value of each corresponding pixel point according to the first coordinate.

[0137] Preferably, the combining the inverse affine transformation matrix and the pre-shrinking magnification parameter to determine the first coordinate corresponding to the coordinate of each pixel point in the odd column block / even column block in the first BRAM space / second BRAM space includes:

[0138] Obtaining the coordinate of the corresponding pixel point in the original image according to the coordinate of each pixel point and the inverse affine transformation matrix, and multiplying the coordinate of the pixel point in the original image minus the starting coordinate of the first pre-loaded sub-block by the pre-shrinking magnification parameter to obtain the first coordinate.

[0139] Specifically, for any pixel point coordinate (u, v) in the filling sub-block, the coordinate (u', v') of the corresponding original image is calculated according to the inverse radiation transformation matrix. It should be noted that both (u, v) and (u', v') are coordinate points in the same coordinate system. It can be understood that the coordinate system in the BRAM space is different from the above coordinate system. The corresponding coordinate (u'v') needs to be subtracted by the starting coordinate of the first pre-loaded sub-block corresponding to the filling sub-block so that the corresponding coordinate point (u', v') is converted into a coordinate point (u", v") in the BRAM space coordinate system; according to the pre-shrinking magnification parameter N / M, the first coordinate corresponding to the pixel point (u, v) in the BRAM space is (u"*N / M, v"*N / M).

[0140] It can be understood that the coordinate systems of the first BRAM space and the second BRAM space are the same.

[0141] Specifically, the determining the pixel value of each corresponding pixel point according to the first coordinate includes:

[0142] Let P(x, y) represent the first coordinate;

[0143] When both x and y are integers, take the pixel value of point P as the pixel value of each corresponding pixel point;

[0144] When both x and y are not integers, determine four second coordinates around point P according to point P, and address from the BRAM space to determine the pixel values of the four second coordinates; based on the bilinear interpolation algorithm, calculate the pixel value corresponding to the first coordinate through the pixel values of the four second coordinates, and take the pixel value corresponding to the first coordinate as the pixel value of each corresponding pixel point;

[0145] When only one of x and y is an integer, determine two third coordinates around point P according to point P, and address from the BRAM space to determine the pixel values of the two third coordinates; based on the bilinear interpolation algorithm, calculate the pixel value corresponding to the first coordinate through the pixel values of the two third coordinates, and take the pixel value corresponding to the first coordinate as the pixel value of each corresponding pixel point.

[0146] Specifically, when both x and y are integers, it indicates that the first coordinate P(x, y) is a pixel point. At this time, directly take out the pixel value corresponding to point P in the BRAM space as the pixel value of the pixel point in the filling sub - tile.

[0147] When only one of x and y is an integer, for example, Figure 7 As shown, y is an integer and x is not an integer, indicating that the first coordinate P(x, y) is between the integer coordinates of two adjacent pixel points. At this time, determine two third coordinates P1 and P2 according to point P, take out the pixel values DixP1 and DixP2 corresponding to the third coordinates P1 and P2 in the BRAM space, then the pixel value corresponding to the first coordinate P point is DixP = DixP1 * k1 + Dix * k2, where Take the pixel value corresponding to point P as the pixel value of the pixel point in the filling sub - tile.

[0148] When both x and y are not integers, such as Figure 8As shown in the figure, it shows that the first coordinate P(x, y) is between 4 adjacent pixels. At this time, four second coordinates P1, P2, P3, and P4 are determined according to point P, and the pixel values of the second coordinates P1, P2, P3, and P4 in the BRAM space are taken out as DixP1, DixP2, DixP3, and DixP4. Then the pixel value corresponding to the first coordinate P point is DixP = (DixP1 * k1 + DixP2 * k3 + DixP3 * k1 + DixP4 * k3 + DixP1 * k4 + DixP3 * k2 + DixP2 * k4 + DixP4 * k2) / 4, where The pixel value corresponding to point P is used as the pixel value of the pixel point in the filling sub-block. Among them, is rounding down, is rounding up.

[0149] After obtaining the pixel values in each filling sub-block in turn according to the above method, all the filling sub-blocks form the result image.

[0150] Compared with the prior art, the data acceleration processing method based on FPGA affine inverse transformation provided by the embodiment of the present invention determines the first coordinate corresponding in the BRAM space through any pixel point in the filling sub-block, determines the pixel value corresponding to the first coordinate according to different situations of the first coordinate, and uses the pixel value of the first coordinate as the pixel value of the corresponding pixel point in the filling sub-block.

[0151] Those skilled in the art can understand that all or part of the processes of implementing the method of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.

[0152] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any change or replacement that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A data acceleration processing method based on FPGA affine inverse transformation, characterized in that The data acceleration processing method includes: Dividing the result image into multiple sub - image blocks as filling sub - image blocks, and the filling sub - image blocks are divided into odd - column blocks and even - column blocks; Performing a pre - loading step and an addressing and reading step on the odd - column blocks and even - column blocks based on ping - pong operation; The pre - loading step includes: determining corresponding second pre - loading sub - image blocks in the original image according to the odd - column block / even - column block and the inverse affine transformation matrix, storing the second pre - loading sub - image blocks corresponding to the odd - column blocks into a first BRAM space, and storing the second pre - loading sub - image blocks corresponding to the even - column blocks into a second BRAM space; The addressing and reading step includes: sequentially traversing each pixel point in the odd - column block / even - column block, determining the pixel value of each pixel point according to the inverse affine transformation matrix and the second pre - loading sub - image blocks in the first BRAM space / second BRAM space corresponding to the odd - column block / even - column block, and filling the pixel value of each pixel point; The determining of the corresponding second pre - loading sub - image blocks in the original image according to the odd - column block / even - column block and the inverse affine transformation matrix includes: Determining a cropping sub - image block according to the odd - column block / even - column block and the inverse affine transformation matrix, and determining a first pre - loading sub - image block according to the cropping sub - image block; judging whether the size of the first pre - loading sub - image block is greater than the size of the first BRAM space / second BRAM space; if the judgment result is yes, determining a pre - reduction magnification parameter according to the first pre - loading sub - image block and the size of the first BRAM space / second BRAM space, and reducing the first pre - loading sub - image block based on the pre - reduction magnification parameter to obtain a second pre - loading sub - image block; if the judgment result is no, taking the first pre - loading sub - image block as the second pre - loading sub - image block.

2. The data acceleration processing method according to claim 1, wherein The determining of the pixel value of each pixel point according to the inverse affine transformation matrix and the second pre - loading sub - image blocks in the first BRAM space or second BRAM space corresponding to the odd - column block / even - column block includes: Combining the inverse affine transformation matrix and the pre - reduction magnification parameter to determine a first coordinate corresponding to the coordinate of each pixel point in the odd - column block / even - column block in the first BRAM space / second BRAM space; Determining the pixel value of each corresponding pixel point according to the first coordinate.

3. The data acceleration processing method according to claim 2, wherein The combining of the inverse affine transformation matrix and the pre - reduction magnification parameter to determine a first coordinate corresponding to the coordinate of each pixel point in the odd - column block / even - column block in the first BRAM space / second BRAM space includes: Obtaining the coordinate of the corresponding pixel point in the original image according to the coordinate of each pixel point and the inverse affine transformation matrix, and multiplying the coordinate of the pixel point in the original image minus the starting coordinate of the first pre - loading sub - image block by the pre - reduction magnification parameter to obtain the first coordinate.

4. The data acceleration processing method according to claim 1, wherein The performing of the pre - loading step and the addressing and reading step on the odd - column blocks and even - column blocks based on ping - pong operation includes: Arrange all the filled sub - tiles in the result image in the order from left to right and from top to bottom, and perform the pre - loading step and the addressing and reading step on the odd - column blocks / even - column blocks based on the ping - pong operation; When performing the pre - loading step on the odd - column blocks / even - column blocks, simultaneously perform the addressing and reading step on the even - column blocks / odd - column blocks.

5. The data acceleration processing method according to claim 1, wherein The determining of the cropped sub - tiles according to the odd - column blocks / even - column blocks and the inverse affine transformation matrix includes: Determine the filling coordinates corresponding to the odd - column blocks / even - column blocks; the filling coordinates include the four filling vertex coordinates corresponding to the odd - column blocks / even - column blocks; Based on the inverse affine transformation, determine the cropping coordinates of the cropped sub - tiles according to the four filling vertex coordinates and the inverse affine transformation matrix, and the cropping coordinates include four cropping vertex coordinates; Determine the cropped sub - tiles according to the four cropping vertex coordinates.

6. The data acceleration processing method according to claim 5, wherein The determining of the first pre - loaded sub - tiles according to the cropped sub - tiles includes: According to the four cropping vertex coordinates of the cropped sub - tiles, determine the maximum and minimum values of the cropped sub - tiles on the two coordinate axes; According to the maximum and minimum values on the two coordinate axes, determine the four vertex coordinates of the first pre - loaded sub - tiles: ; ; ; ; Wherein, A, B, C, and D represent the four vertex coordinates of the first pre - loaded sub - tiles, Xmax represents the maximum value of the cropped sub - tiles on the X - axis, Xmin represents the minimum value of the cropped sub - tiles on the X - axis, Ymax represents the maximum value of the cropped sub - tiles on the Y - axis, and Ymin represents the minimum value of the cropped sub - tiles on the Y - axis.

7. The data acceleration processing method according to claim 1, wherein The judging whether the size of the first pre - loaded sub - tiles is larger than the size of the first BRAM space / the second BRAM space includes: Determine the length and width of the first pre - loaded sub - tiles and determine the length and width of the first BRAM space / the second BRAM space; Calculate the length ratio of the length of the first pre - loaded sub - tiles and the length of the first BRAM space / the second BRAM space and calculate the width ratio of the width of the first pre - loaded sub - tiles and the width of the first BRAM space / the second BRAM space; the sizes of the first BRAM space and the second BRAM space are the same; If any one of the length ratio and the width ratio is greater than 1, judge as yes, otherwise, judge as no.

8. The data acceleration processing method according to claim 7, wherein The determining of the pre - reduction magnification parameter according to the size of the first pre - loaded sub - tiles and the first BRAM space / the second BRAM space includes: Judge whether the length ratio is greater than the width ratio; if the judgment is yes, determine the pre - reduction magnification parameter according to the length ratio; otherwise, determine the pre - reduction magnification parameter according to the width ratio; the pre - reduction magnification parameter is N / M, where N and M are both integers in [1, 6], and M > N.

9. The data acceleration processing method according to claim 8, wherein The shrinking of the first pre - loaded sub - tiles based on the pre - reduction magnification parameter to obtain the second pre - loaded sub - tiles includes: Insert N - 1 pixel points between every two adjacent pixels in each row of the first pre - loaded sub - block. Then, insert N - 1 pixel points between every two adjacent pixels in each column of the first pre - loaded sub - block, thereby expanding the first pre - loaded sub - block by N times to obtain an intermediate sub - block; In the intermediate sub - block, select every M - 1th row as a row to be sampled. In each row to be sampled, select every M - 1th pixel point as a sampling point, thereby shrinking the intermediate sub - block by M times to obtain a second pre - loaded sub - block.

Citation Information

Patent Citations

  • Image remapping method and device based on programmable logic device

    CN105528758A

  • Image projection transformation method and device and electronic equipment

    CN110503602A