High-performance real-time image scaling method based on FPGA
Through geometric center method and appropriate magnification selection, the image scaling algorithm on the FPGA platform is optimized, which solves the problems of large resource consumption and slow computing speed in the existing technology, and achieves efficient and real-time image scaling.
Patent Information
- Application Number
- CN202510158951.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to save logical resources and improve computing speed while ensuring the scaling effect during image scaling, especially on FPGA platforms.
The weight coefficients of the original image and the target image are determined by the geometric center method and magnified by 2N times. Combined with the appropriate magnification selection, the use of BRAM and multiplier is optimized, and the ping-pong cache technology is used to save resources and improve calculation speed.
It realizes that while ensuring the scaling effect, save logical resources, improve calculation speed, and improve the real-time performance of the image scaling algorithm.
Smart Images

Figure CN120070165A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image scaling method, specifically a high-performance real-time image scaling method based on FPGA, belonging to the field of digital image processing technology. Background Art
[0002] With the development of digital image processing algorithms and video high-speed transmission technologies, people's requirements for image quality are constantly increasing. Digital video images have characteristics such as large data volume and high real-time performance. The processing of video images is a key issue in the field of image processing, and image scaling algorithms play a particularly important role in improving image quality. Due to the continuous development of FPGA, it has been widely used in various fields, especially in the 4K video image field, and using FPGA for image scaling is also one of the current mainstream methods. The bilinear interpolation algorithm has gradually become one of the current mainstream scaling algorithms due to its advantages such as simplicity, high efficiency, image smoothing, and low computational complexity.
[0003] In the prior art, a method and device for image scaling processing based on FPGA disclosed in a patent with publication number CN106910162A obtain the original image data, put the original image data into the internal cache of the FPGA at a preset input speed, read the original image data from the internal cache at a reading speed corresponding to the preset input speed, and perform interpolation calculation according to the interpolation algorithm and the read original image data to obtain image interpolation data, and obtain the scaled image data according to the image interpolation data. Caching the original image data in the internal cache of the FPGA eliminates the need to store it using an external memory, reducing costs, and it can be read from the internal cache of the FPGA without having to read from the external memory again, accelerating the data reading speed and improving the overall image scaling efficiency. In addition, reading the original image data from the internal cache at a reading speed corresponding to the preset input speed, that is, reading the data at a reasonable reading speed, can ensure that the reading speed avoids data loss due to too slow a reading speed and improves the scaling processing efficiency. In order to ensure the scaling effect, most current processing methods consume a large amount of BRAM resources or multiplier resources. Therefore, how to implement an image scaling algorithm that can save logic resources and improve the calculation speed while ensuring the scaling effect is a technical problem that needs to be solved in this field. Summary of the Invention
[0004] The purpose of the present invention is to provide a high-performance real-time image scaling method based on FPGA to solve at least one of the above technical problems. This method, while ensuring the scaling effect, balances the BRAM resources and the number of multipliers, selects an appropriate scaling accuracy, can save logic resources, improve the calculation speed, and has high real-time performance.
[0005] The present invention achieves the above object through the following technical solutions: A method for high-performance real-time image scaling based on FPGA, the method comprising the following steps: S1. Determine the weight coefficients of the original image and the target image in the row and field directions using the geometric center method, and magnify by 2 N times; S2. By selecting an appropriate magnification factor, balance the calculation accuracy and calculation speed; S3. Sequentially ping-pong cache two consecutive rows of original image data into two BRAMs. After the second row is completely written into the BRAM, read these two BRAMs simultaneously, and send them together with the row weight coefficients to the calculation unit to obtain the calculation result; S4. Concatenate the calculation results and send them to BRAM2 for caching. At the same time, read out the cached data and send it together with the field weight coefficients to the calculation unit, thereby calculating the complete data of one row of the target image.
[0006] As a further aspect of the present invention: The determination of the weight coefficients of the original image and the target image specifically includes: Select the geometric center of the image data as the reference point, obtain the relationship between the original image coordinates and the target image coordinates, add 0.5 to both the original image coordinates and the target image coordinates, and then optimize them to obtain the weight coefficients and magnify the weight coefficients. The optimized formula is as follows: srcX = ((2 * destX + 1) * srcWidth / destWidth - 1) / 2; srcY = ((2 * destY + 1) * srcHeight / destHeight - 1) / 2; Wherein, srcWidth is the width of the source image, destWidth is the width of the destination image, srcHeight is the height of the source image, destHeight is the width of the source image, srcX is the position of the source image row pixel, destX is the position of the destination image row pixel, srcY is the position of the source image field pixel, and destY is the position of the destination image field pixel.
[0007] As a further aspect of the present invention: The magnification of the weight coefficients specifically includes: Magnify the weight coefficients of the original image coordinates and the target image coordinates by 2 N times, and at the same time let para_l = 2 N * srcWidth / destWidth; para_h = 2 N * srcHeight / destHeight; Among them, para_l is the row scaling parameter, para_h is the field scaling parameter, and the parameters are all integers; by calculating the values of para_l and para_h as described above, the division operation can be converted into a multiplication operation; then the above formula can be further simplified to: srcX = ((2 * destX + 1) * para_l - 2 N )>>(N + 1); srcY = ((2 * destY + 1) * para_h - 2 N )>>(N + 1).
[0008] As a further solution of the present invention: The cache calculation of the original data specifically includes: S31. Write the first row of the original image data into BRAM0, and write the second row into BRAM1. Reasonably set the depth of the BRAM according to the number of pixels in the row direction of the original image to improve resource utilization; S32. After writing two consecutive rows of the original image data, use f0 and f0 + 1 as addresses to simultaneously read the data in BRAM0 and BRAM1. The data read from BRAM0 is dout00 and dout01, and the data read from BRAM1 is dout10 and dout11; S33. Send dout00, dout01, and coe0 to the calculation unit at the same time. Amplify dout00 by 2 N+1 times, then use coe0 * (dout01 - dout00), and finally add the two with an adder. Discard the lower 13 bits of the calculation result to obtain the first scaled pixel q0(x, y0). As the address continuously accumulates, the intermediate data q1(x, y0), q2(x, y0),..., qw(x, y0) of all pixels in the first row can be obtained; similarly, the intermediate data q0(x, y1), q1(x, y1),..., qw(x, y1) of all pixels in the second row can also be obtained; S34. Repeat the above steps. Through continuous ping-pong caching, the scaling operation of the original image to the target image column can be completed, and the scaled data of two consecutive rows of the target image can be obtained each time.
[0009] As a further solution of the present invention: The calculation of the complete data specifically includes: S41. Concatenate the intermediate data of the two rows of pixels obtained by calculation and send them to BRAM2 for caching. While caching, read the data according to the coordinates of the destination pixels and send it to the calculation unit together with the field weight coefficient coe1. Reasonably set the depth of the BRAM according to the total number of columns of the target image data to improve resource utilization; S42. Send dout0, dout1, and coe1 into the calculation unit simultaneously; first, amplify dout0 by 2 N+1 times, then use coe1 * (dout1 - dout0), and finally add these two results together with an adder. Discard the lower 13 bits of the calculation result to obtain the first scaled pixel q(x, y). As the address accumulates continuously, the intermediate data q1(x, y), q2(x, y),..., qw(x, y) of all pixels in the first row can be obtained; S43. Repeat the above steps to complete the scaling operation from the original image to the columns of the target image and obtain all the target image data.
[0010] The beneficial effects of the present invention are as follows: 1) The present invention uses the geometric center method to determine the weight coefficients of the original image and the target image in the row and field directions and amplifies them by 2 N times. The geometric center method is used to eliminate the influence of errors caused by coordinate selection on the scaling algorithm, and the row and column scaling parameters are obtained through calculation. The division operation is removed and converted into a multiplication operation, optimizing the use of resources; 2) The present invention selects an appropriate magnification factor to balance the calculation accuracy and calculation speed, and can improve the calculation speed while meeting the accuracy requirements; 3) The present invention sequentially pings and caches two consecutive rows of original image data into two BRAMs. After the second row is completely written into the BRAM, these two BRAMs are read simultaneously and sent into the calculation unit together with the row weight coefficients. The dual-port BRAM is used, saving the number of BRAMs. Only two multipliers and four adders and subtracters are needed to obtain two consecutive rows of image data, saving the number of multipliers used, and reducing the calculation amount of the algorithm. Only two BRAMs can complete the scaling calculation of all rows of the entire frame through the ping-pong caching method, and two consecutive rows of data can be extracted; 4) The present invention splices the calculation results and sends them into BRAM2 for caching, and at the same time reads out the cached data and sends it into the calculation unit together with the field weight coefficients, so that a complete row of data can be calculated, saving the number of multipliers used. Only one multiplier and two adders and subtracters are needed to obtain the finally scaled target image data; 5) The present invention can scale the input source video image signal with any resolution to the target video image signal with the required resolution. On the premise of ensuring the scaling effect, a small amount of BRAM resources inside the FPGA are used, reducing the number of multiplier resources, saving logic resources within a reasonable accuracy range, and improving the calculation speed of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a block diagram for implementing the scaling algorithm of the present invention; Figure 2 Schematic diagram of the process of mapping the source image to the target image according to the present invention; Figure 3 Schematic diagram of the principle of the scaling algorithm according to the present invention; Figure 4 Schematic diagram of the implementation process of the scaling algorithm according to the present invention. Specific embodiments
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0013] Embodiment 1, as Figures 1 to 4 shown, this embodiment provides a method for high-performance real-time image scaling based on FPGA, and the method includes the following steps: First: Use the geometric center method to determine the weight coefficients of the original image and the target image in the row and field directions, and magnify by 2 N times.
[0014] The determination of the weight coefficients of the original image and the target image specifically includes: Select the geometric center of the image data (original image data and target image data) as the reference point, obtain the relationship between the original image coordinates and the target image coordinates, add 0.5 to both the original image coordinates and the target image coordinates, and then optimize them to obtain the weight coefficients, and magnify the weight coefficients. The optimized formula is as follows: srcX = ((2 * destX + 1) * srcWidth / destWidth - 1) / 2; srcY = ((2 * destY + 1) * srcHeight / destHeight - 1) / 2; where srcWidth is the width of the source image, destWidth is the width of the destination image, srcHeight is the height of the source image, destHeight is the width of the source image, srcX is the position of the source image row pixel, destX is the position of the destination image row pixel, srcY is the position of the source image field pixel, and destY is the position of the destination image field pixel, as Figure 2 shown.
[0015] Use the geometric center method to map the target image coordinates to the source image coordinates (original image coordinates), and eliminate the influence of errors caused by different coordinate selections on the scaling effect.
[0016] The magnification of the weight coefficient specifically includes: Magnify the weight coefficients of the original image coordinates and the target image coordinates by 2 N times simultaneously, and at the same time let para_l = 2 N *srcWidth / destWidth; para_h = 2 N *srcHeight / destHeight; Among them, para_l is the row scaling parameter, para_h is the field scaling parameter, and the parameters are all integers; by calculating the values of para_l and para_h through the above formula, the division operation can be converted into a multiplication operation, and the above formula can be further simplified as: srcX = ((2 * destX + 1) * para_l - 2 N )>>(N + 1); srcY = ((2 * destY + 1) * para_h - 2 N )>>(N + 1); By calculating the row and column scaling parameters, removing the division operation and converting it into a multiplication operation, the use of resources is optimized.
[0017] Second: By selecting an appropriate magnification factor, the calculation accuracy and calculation speed are balanced.
[0018] It can be calculated by the geometric center method that as the exponent N increases, the accuracy of mapping the destination pixel coordinates to the source pixels is higher. When N = 12, the error of mapping the destination pixel coordinates to the source pixels is already very small and tends to be stable. Therefore, select the magnification factor to be 2 12 , that is, 4096 times. Since the geometric center formula has a built-in magnification factor of 2, the total magnification factor is 8192 times.
[0019] The high 13 bits of the above srcX calculation result are used as the integer f0, and the low 13 bits are used as the decimal coefficient coe0; the high 13 bits of the srcY calculation result are used as the integer f1, and the low 13 bits are used as the decimal coefficient coe1.
[0020] Third: Sequentially ping-pong cache two consecutive rows of the original image data into two BRAMs. After the second row is completely written into the BRAM, read these two BRAMs simultaneously and send them to the calculation unit together with the row weight coefficients.
[0021] The caching calculation of the original image data specifically includes: 1) Write the first row of the original image into BRAM0, and write the second row into BRAM1. Reasonably set the depth of the BRAM according to the number of pixels in the row direction of the original image to improve the resource utilization rate; 2) After writing two consecutive rows of original image data, use f0 and f0 + 1 as addresses to read out the data in BRAM0 and BRAM1 simultaneously. The data read from BRAM0 are dout00 and dout01, and the data read from BRAM1 are dout10 and dout11. The dual-port BRAM is used, saving the number of BRAMs; 3) Send dout00, dout01 and coe0 into the calculation unit at the same time. First, magnify dout00 by 2 N+1 times, then use coe0 * (dout01 - dout00), and finally add these two with an adder. Discard the lower 13 bits of the calculation result to obtain the first scaled pixel q0(x, y0). As the address accumulates continuously, the intermediate data q1(x, y0), q2(x, y0), …, qw(x, y0) of all pixels in the first row can be obtained. Similarly, the intermediate data q0(x, y1), q1(x, y1), …, qw(x, y1) of all pixels in the second row can also be obtained, saving the number of multipliers used. Only two multipliers and four adders and subtracters are needed to obtain two consecutive rows of image data; 4) Repeat the above steps. Through continuous ping-pong buffering, the scaling operation from the original image to the target image column can be completed. Two consecutive rows of scaled data of the target image can be obtained each time. Only two BRAMs can complete the scaling calculation of all rows of the entire frame through the ping-pong buffering method, and two consecutive rows of data can be extracted. It is neither necessary to use BRAM to cache a complete frame nor an external storage device for caching, greatly simplifying the utilization of BRAM.
[0022] Fourth: During the above calculation, splice the calculation results and send them into BRAM2 for caching. At the same time, read out the cached data and send it into the calculation unit together with the field weight coefficient; in this way, the complete data of one row of the target image can be calculated.
[0023] The calculation of the complete data specifically includes: 1) Splice the intermediate data of the two rows of pixels obtained by calculation and send them into BRAM2 for caching. While caching, read out the data according to the coordinates of the destination pixels and send it into the calculation unit together with the field weight coefficient coe1. Reasonably set the depth of BRAM according to the total number of columns of the target image data to improve resource utilization; 2) Send dout0, dout1 and coe1 into the calculation unit at the same time. First, magnify dout0 by 2 N+1times, then use coe1*(dout1 - dout0), and finally use an adder to add these two together. Discard the lower 13 bits of the calculation result to obtain the first scaled pixel q(x, y). As the address accumulates continuously, the intermediate data q1(x, y), q2(x, y), …, qw(x, y) of all pixels in the first row can be obtained, saving the number of multipliers used. Only one multiplier and two adder / subtracters are needed to obtain the final scaled target image data; 3) Repeat the above steps to complete the scaling operation from the original image to the columns of the target image and obtain all the target image data.
[0024] Working principle: To scale the original image to the target image, first, calculate the scaling parameters of the row and field through the geometric center method. The calculated scaling parameters are magnified by 2 N times; Secondly, select the exponent N of the magnification factor as 12, which can well balance the calculation accuracy and calculation speed; Then, write the data of the first row of the original image into BRAM0, and then write the data of the second row into BRAM1. After the data of the second row is completely written into BRAM1, start reading the data of BRAM0 and BRAM1, and send the data and their respective row weight wash data into the calculation unit to obtain the intermediate data of these two rows. Concatenate the data of these two rows into a row of pixel data and send it into BRAM2 for caching. While caching, read the data in BRAM2 and send it and the corresponding column weight coefficients into the calculation unit to obtain the data of the first row of the target image. Finally, repeat the above steps and write the original image data in a ping-pong manner to obtain the data of the first row, the second row, and so on until the last row of data.
[0025] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0026] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A high-performance real-time image scaling method based on FPGA, characterized in that: The method comprises the following steps: S1. Use the geometric center method to determine the weight coefficients of the original image and the target image in the row field direction, and enlarge it by 2 N times; S2, by selecting the magnification, to balance the calculation accuracy and speed; S3, ping-pong buffering two consecutive rows of original image data into two BRAMs in sequence, and after the second row is completely written into the BRAM, reading the two BRAMs at the same time and sending them together with the row weight coefficient into the calculation unit to obtain the calculation result; S4, the calculation results are spliced and sent to BRAM2 for caching, and the cached data is read out at the same time, and sent to the calculation unit together with the field weight coefficient, so as to calculate the complete data of a row of the target image.
2. The method of high-performance real-time image scaling based on FPGA according to claim 1, characterized in that: In S1, the determination of the weight coefficients of the original image and the target image specifically includes: The geometric center of the image data is selected as the reference point to obtain the relationship between the original image coordinates and the target image coordinates, and 0.5 is added to both the original image coordinates and the target image coordinates. Then, the weight coefficient is optimized to obtain the weight coefficient, and the weight coefficient is amplified. The optimized formula is as follows: srcX=((2*destX+1)*srcWidth / destWidth-1) / 2; srcY=((2*destY+1)*srcHeight / destHeight-1) / 2; Among them, srcWidth is the width of the source image, destWidth is the width of the destination image, srcHeight is the height of the source image, destHeight is the width of the source image, srcX is the position of the source image row pixel, destX is the position of the destination image row pixel, srcY is the position of the source image field pixel, and destY is the position of the destination image field pixel.
3. The method for high-performance real-time image scaling based on FPGA according to claim 2, characterized in that: The amplification of the weight coefficient specifically includes: The weight coefficients of the original image coordinates and the target image coordinates are simultaneously enlarged by 2 N times, and at the same time <h2 style=";text-align:left;direction:ltr">para_l=2<h2 style=";text-align:left;direction:ltr"> N <h2 style=";text-align:left;direction:ltr"> *srcWidth / destWidth; para_h=2 N *srcHeight / destHeight; Among them, para_l is the row scaling parameter, para_h is the field scaling parameter, and both parameters are integers; by calculating the values of para_l and para_h, the division operation is converted into a multiplication operation, and the formula is simplified to: srcX=((2*destX+1)*para_l-2 N )>>(N+1); srcY=((2*destY+1)*para_h-2 N )>>(N+1).
4. The method of high-performance real-time image scaling based on FPGA according to claim 1, characterized in that: In S3, the cache calculation of the original image data specifically includes: S31, writing the first row of the original image data into BRAM0, and the second row into BRAM1, and setting the depth of the BRAM according to the number of pixels in the row direction of the original image; S32, after two consecutive lines of original image data are written, the data in BRAM0 and BRAM1 are read out simultaneously using f0 and f0+1 as addresses, the data read out of BRAM0 are dout00 and dout01, and the data read out of BRAM1 are dout10 and dout11; S33, send dout00, dout01 and coe0 to the calculation unit at the same time, and amplify dout00 by 2 N+1 times, then use coe0*(dout01-dout00), and finally use an adder to add the two together, discard the lower 13 bits of the result of the calculation, and get the first pixel q0(x,y0) after scaling. As the address is continuously accumulated, the intermediate data q1(x,y0), q2(x,y0), ..., qw(x,y0) of all pixels in the first row are obtained; similarly, the intermediate data q0(x,y1), q1(x,y1), ..., qw(x,y1) of all pixels in the second row are obtained. S34, repeating the above steps, through continuous ping-pong caching, completing the scaling operation from the original image to the target image column, and obtaining two consecutive rows of target image scaling data each time.
5. The method for high-performance real-time image scaling based on FPGA according to claim 4, characterized in that: In S4, the calculation of complete data specifically includes: S41, the intermediate data of the two rows of pixels calculated are spliced and sent to the BRAM2 for buffering, and the data is read out at the coordinates of the target pixel while being buffered, and sent to the calculation unit together with the field weight coefficient coe1, and the depth of the BRAM is set according to the total number of columns of the target image data; S42, send dout0, dout1 and coe1 to the calculation unit at the same time, and amplify dout0 by 2 N+1 times, then use coe1*(dout1-dout0), and finally use an adder to add the two together, discard the lower 13 bits of the result of the calculation, and get the first pixel q(x,y) after scaling. As the address is continuously accumulated, the intermediate data q1(x,y), q2(x,y), …, qw(x,y) of all pixels in the first row are obtained; S43, repeat the above steps to complete the scaling operation from the original image to the target image sequence, and obtain all the target image data.
Citation Information
Patent Citations
FPGA-based image scaling processing method and device
CN106910162A