FPGA-Based Guided Filter Weighted Aggregation Method and System
Through the boot filter weighted aggregation method on the FPGA platform, the linear regression coefficient and error allocation weight are used to optimize the processing flow, and the halo artifact problem of the boot filter algorithm at the edge is solved, achieving efficient and low-latency image processing effect.
Patent Information
- Application Number
- CN202111432755.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-29
AI Technical Summary
The existing boot filtering algorithms produce halo artifacts at the edge of the image and have high computational complexity, which cannot be applied in real-time on miniaturized, low-cost processors.
The FPGA-based guide filter weighted aggregation method is adopted to obtain the linear regression coefficients of neighboring blocks of pixel points, allocate weights according to the error size, and perform weighted aggregation, combining fixed-pointing processing, cassette filtering and Pipeline acceleration calculation modules to optimize the processing flow.
Effectively reduce or eliminate halo artifacts, maintain edge clarity, and at the same time realize high parallel and low latency image processing to meet the real-time needs of small camera hardware processors.
Smart Images

Figure CN114463193B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technologies, and in particular, to a guided filter weighted aggregation method and system based on FPGA. Background Art
[0002] With the popular application of video surveillance cameras, especially in application scenarios such as night scenes, long-distance observation, and personnel temperature measurement, the application of infrared thermal imaging cameras is becoming more and more extensive. The imaging quality of infrared thermal imaging cameras has attracted much attention. However, due to noise reasons such as the detector manufacturing process and electronic circuits of infrared thermal imaging cameras, the video images output by thermal imaging have relatively large noise. Generally, traditional filtering and noise reduction algorithms, such as mean filtering, median filtering, and Gaussian filtering, have unsatisfactory noise reduction effects or blur the edge contours of the images. Bilateral filtering, non-local means filtering, and BM3D 3D filtering algorithms have high complexity, consume too many resources, and have large delays, and are not very suitable for deployment on small and low-cost processors at the edge. Guided filtering has relatively more advantages in edge-preserving filtering algorithms.
[0003] However, during the guided filtering process, among all the neighborhood blocks of each pixel, the regularization parameter (smoothing factor) is given a priori and remains unchanged. In the image, for edge contours or flat regions, this strategy is adopted. After the edge contours are locally smoothed, halation artifacts will inevitably occur. In order to optimize the edge-preserving performance of the guided filter algorithm and at the same time solve or reduce the halation artifact phenomenon generated during the local smoothing process, a new weight weighted aggregation method is proposed.
[0004] The guided filter uses a local linear model as Figure 1 shown. This model believes that a point on a certain function and the points in its adjacent part are in a linear relationship, and a complex function can be represented by many local linear functions. When the value of a certain point on the function needs to be calculated, only the values of all the linear functions containing this point need to be calculated and averaged. This model is very useful for representing non-analytical functions. We can consider an image as a two-dimensional function and cannot write an analytical expression. Therefore, we assume that the output of this function and the input satisfy a linear relationship within a two-dimensional window:
[0005] q = mean a .*I + mean b
[0006] where q is the output of the guided filter, mean a , mean bis the average coefficient of the linear function within the local two-dimensional window neighborhood, and I is the input image. The guided filter image algorithm is an edge-preserving filtering algorithm compared to image filtering algorithms such as mean filtering and Gaussian filtering. However, as a local filter, the guided image filter has the problem of halo artifacts, that is, there are circular artifact contours in the display effect, seriously affecting the visual effect.
[0007] At the same time, when calculating the output of each pixel, this filtering algorithm needs to go through processes such as calculating the linear model parameters multiple times in multiple adjacent neighborhoods centered on the current pixel, then multiplying, accumulating, summing, and averaging. Moreover, the larger the radius of the neighborhood window, the longer the calculation time within the neighborhood, the greater the image output delay, and the more hardware resources are consumed. And it cannot be applied to our display device in real time. Summary of the Invention
[0008] Aiming at the defects in the prior art, the purpose of the present invention is to provide a guided filter weighted aggregation method and system based on FPGA.
[0009] According to a guided filter weighted aggregation method based on FPGA provided by the present invention, it includes the following steps:
[0010] Step S1: Use guided filter denoising to process the image, and obtain the linear regression coefficients of the neighborhood blocks of pixel points;
[0011] Step S2: Configure weights for different neighborhoods of pixel points according to the linear regression coefficients, perform weighted aggregation on the new weights, and generate a denoised image.
[0012] Preferably, the linear regression coefficients in step S1 include:
[0013]
[0014]
[0015] Among them, σ 2 i is the variance of the coefficient window, n 2 is the number of coefficient windows, ε is the smoothing factor, is the average value within the coefficient window.
[0016] Preferably, in step S2, weight coefficients are assigned according to the error magnitude generated by the pixel point calculation within the neighborhood, and the larger the error range, the smaller the weight coefficient.
[0017] Preferably, the error is the difference between the value output after the neighborhood is processed by guided filter denoising and the original true value, and the mean square error e i is used to correct the weight of the neighborhood;
[0018]
[0019] The denoised image I′ is represented as:
[0020]
[0021] The weight coefficient γ i = exp(-e i . / μ), where μ > 0 is a scalar coefficient, guiding the image G as the input image I, and simplifying the mean square error formula to obtain:
[0022] e i = (1 - a i ) 2 σ 2 i = a i ε(1 - a i )
[0023] where σ 2 i is the variance of the input image in the i neighborhood, I is the input image, ε is the regularization parameter, and μ is the scalar parameter.
[0024] Preferably, the linear coefficients a i and b i of each neighborhood are multiplied according to the new weight coefficient to obtain the reallocated linear coefficients g a and g b , and the pixel grayscales are multiplied and weighted according to the newly allocated coefficients to obtain the final output result.
[0025] According to a guided filtering weighted aggregation system based on FPGA introduced in the present invention, it includes the following modules:
[0026] Fixed-point processing module: converting the process of calculating the weight of the exp exponential power into a look-up table and integer-type operation process suitable for processing on an FPGA hardware processor;
[0027] Box filtering processing module: statistically summing the data within a neighborhood space to calculate the mean value;
[0028] Pipeline acceleration calculation module: processing calculation tasks in parallel;
[0029] Delay optimization module: analyzing and optimizing the time consumption of the system to reduce data processing delay.
[0030] Preferably, the fixed-point processing module amplifies, normalizes, and converts the linear regression coefficient a i , the mean square error e i and the weight coefficient γ i , and fixes the regularization parameter εf and the scalar parameter μ f , substitute a f into the fixed-point calculation module, calculate the conversion to a look-up table and then perform a power-of-2 and multiplication operation to obtain the calculation result γ of the weight coefficient f .
[0031] Preferably, the box filtering processing module receives image data in a row-column data stream manner and enters it into the box frame, and when one data enters the box frame, one calculation result is output.
[0032] Preferably, the pipeline acceleration calculation module synchronously executes each calculation step. After the data calculation in the first calculation process is completed, the result is passed to the second calculation process for calculation, and at the same time, the first calculation process receives the input calculation of the next data.
[0033] Preferably, the image data is divided into row data and column data and flows into the pipeline acceleration calculation module in sequence. The pipeline acceleration calculation module includes multiple calculation structures, and the image data enters different calculation structures in sequence, and each calculation structure performs calculations synchronously.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1. A new weighted aggregation method in the guided filtering process. For the edge contours or flat regions in the image, the modified weights are adopted, which solves or alleviates the halo artifact phenomenon generated during the local smoothing process, effectively preserves the edges, and can also achieve a good smoothing and noise reduction effect in the edge contour regions of the image, improving the visual display effect of the image.
[0036] 2. An acceleration calculation scheme is designed in the FPGA processing platform. The image is divided into multiple neighborhoods and processed in parallel at the same time, which can meet the requirements of higher frame rate and lower latency.
[0037] 3. Solve the problems of algorithm implementation and deployment, and can be compatible, efficient and low-latency in a small camera hardware processor platform, meeting the requirements of the frame rate and delay of the overall system. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objects and advantages of the present invention will become more obvious:
[0039] Figure 1 It is a schematic diagram of a local linear model in the prior art;
[0040] Figure 2 It is a schematic diagram of the guided filtering sliding window in the embodiment of the present invention;
[0041] Figure 3 It is a schematic diagram of edge step grayscale in the embodiment of the present invention;
[0042] Figure 4 It is a schematic diagram of the change of mean square error and variance in the embodiment of the present invention;
[0043] Figure 5 It is a schematic diagram of calculation steps in the embodiment of the present invention;
[0044] Figure 6 It is a schematic diagram of the weighted aggregation framework in the embodiment of the present invention;
[0045] Figure 7 It is a schematic diagram of the fixed-point processing process in the embodiment of the present invention;
[0046] Figure 8 It is a schematic diagram of box filtering in the embodiment of the present invention;
[0047] Figure 9 It is a schematic diagram of the comparison between the ordinary structure and the pipeline acceleration calculation module in the embodiment of the present invention;
[0048] Figure 10 It is a schematic diagram of the analysis of the delay optimization module in the embodiment of the present invention. Detailed implementation manners
[0049] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0050] The present invention introduces a guided filter weighted aggregation method based on FPGA, including the following steps:
[0051] Step S1: Use guided filter denoising to process the image, and obtain the linear regression coefficients of the neighborhood blocks of pixel points;
[0052] Step S2: Configure weights for different neighborhoods of pixel points according to the linear regression coefficients, and perform weighted aggregation on the new weights to generate a denoised image.
[0053] The present invention will be further introduced below.
[0054] The traditional guided filter denoising process is as Figure 2As shown. In the sliding window schematic diagram, the filter window size is set to 9*9, and the coefficient window size is 5*5. The coefficient window is sequentially moved in the filter window from left to right and from top to bottom, and a total of r*r = 25 coefficient windows are obtained. In each coefficient window, a set of linear coefficients a i and b i .
[0055]
[0056]
[0057] where σ 2 i is the variance within the current coefficient window, ε is the smoothing factor, is the average value within the current coefficient window. Therefore, according to the guided filter formula, when the guidance map is the original image, q = mean a .*I + mean b , I is the central value of the filter window, and q is the filtered output value.
[0058] In the above guided filter noise reduction process, when ε = 0, a = 1, b = 0, then the original image is output without filtering. When ε > 0, in the flat area a ≈ 0, that is, weighted filtering is achieved. At the edge jump, a ≈ 1, b ≈ 0, then the weight of the original pixel is large, effectively protecting the edge. When the value of ε is larger, the filtering intensity is greater. However, when ε is set smaller, the overall image filtering effect is poor. When the image noise is severe, ε is set larger, and the overall smoothing effect is better. But for pixels located on the edge or a single step edge, such as Figure 3 shown, under the framework of the guided filter algorithm, the result of point p in the w i neighborhood is close to the result of point p in the w j neighborhood, but it can be seen from the figure that the gray value of point p is closer to the average gray value in the w i neighborhood and far from the average gray value in the w j neighborhood, thus generating an error. Especially when the parameter r radius and ε smoothing factor are large, the estimation of this pixel in the w i , w j and other neighborhood blocks is averaged and weighted and far from the true value, resulting in a blurred edge and causing halo artifacts. To solve this problem, it is proposed to configure different weights for different neighborhoods and use the new weights for weighted aggregation to generate the final output result. Experiments show that the algorithm in the present invention can not only effectively preserve the edge, but also achieve a good smoothing and noise reduction effect in the area of the edge contour of the image, so as to improve the visual display effect of the image.
[0059] Therefore, in the linear regression coefficient a of the neighborhood blocki and b i After obtaining, perform appropriate secondary correction to implement a method of configuring different weights for different neighborhoods, rather than using the weighted aggregation of all weights averaged. The correction strategy is: allocate weight coefficients according to the error magnitude generated by calculating the pixel points in the neighborhood, and the larger the error range, the smaller the weight coefficient. When the error generated by calculating at point p in the w i neighborhood is small, a large weight coefficient is allocated, and when the error generated by calculating at point p in the w j neighborhood is large, a smaller weight coefficient is allocated. Adjust the weights of the neighborhood blocks again according to the error magnitude to make the result after weighted calculation closer to the true value, and at the same time, the halo artifact phenomenon is eliminated. During the image processing process, usually the value calculated by the traditional guided filter is compared with the original true value to measure the error magnitude in the calculation process. Therefore, when correcting the weights, it is represented by the square of the difference between the value calculated by the guided filter denoising and the input true value.
[0060]
[0061] If the mean square error e i is smaller, then the weight coefficient γi of this neighborhood block is larger., then the denoised image I′ is expressed as:
[0062]
[0063] Weight coefficient γ i = exp(-e i · / μ), μ > 0 is a scalar coefficient. Assume that the guidance image G is the input image I, and simplify the mean square error formula to get:
[0064] e i = (1 - a i ) 2 σ 2 i = a i ε(1 - a i )
[0065] where, σ 2 i is the variance of the input image in the i neighborhood, I is the input image, ε is the regularization parameter, and μ is the scalar parameter. After fixing the ε parameter, as σ 2 i increases, when σ 2 i ≥ ε, e i monotonically decreases as Figure 4 shown. When point p is at the edge or a single step edge, the variance of the neighborhood block containing point p is larger than that of the flat area at this time (σ 2 i>> ε), and the resulting mean squared error MSE is small, the greater the assigned weight. After all neighborhood blocks containing point p are corrected according to this, the calculated output value is closer to the original input value. On the contrary, if point p is in a flat area, the variance is small (σ 2 i << ε) and the resulting mean squared error MSE is large, the smaller the assigned weight. After all neighborhood blocks containing point p are corrected according to this, the calculated output value is closer to the neighborhood average value.
[0066] To sum up, assuming the input image is I, the radius is r, the regularization parameter is ε, and the scalar parameter is μ, then the output image is I′. The calculation steps are as follows Figure 5 shown. After refinement and arrangement, the schematic diagram of the weighted aggregation framework is as Figure 6 shown. After the video image data stream comes in, it is necessary to delay 4 rows and 4 columns to construct a 5*5 neighborhood template. Calculate the sum and sum of squares of the gray values within the neighborhood template, and calculate the mean value, mean square value, variance, and linear coefficient a i , b i within the template neighborhood along the data flow direction, and obtain the mean squared error e i . Through the mean squared error e i , obtain the weight coefficient γ i . Then perform aggregation, multiply the linear coefficients a i , b i of each neighborhood by the new weight coefficient to obtain the reallocated linear coefficients, and perform weighted multiplication of the pixel gray values according to the newly allocated coefficients to obtain the final result output.
[0067] The present invention also introduces a guided filtering weighted aggregation system based on FPGA. This system can complete Figure 6 the algorithm described in. It is necessary to design a reasonable acceleration scheme in combination with the hardware characteristics of FPGA to achieve the effects of high parallelism and low latency. Specifically, the acceleration scheme includes: a fixed-point processing module, a box filtering processing module, a Pipeline acceleration calculation module, and a delay optimization module
[0068] Fixed-point processing module: Convert the process of calculating weights with exp exponential powers into a look-up table and integer-type operation process suitable for processing on the FPGA hardware processor. In the FPGA processor, only DSP48E is suitable for hardware resources such as integer-type multipliers, registers, and look-up tables, and is not suitable for direct exp exponential calculation. In some high-end FPGA devices, the cordic IP core can be set to calculate, but there are problems such as complex calculation, large consumption of logic resources, and large latency. Therefore, in this solution, a fixed-point processing module is designed, as Figure 7 shown, for the linear coefficient a i derived above and the mean squared error e iand the new weight coefficient γ i , after amplification, normalization, and transformation processing, with a fixed ε f parameter and μ f parameter (the subscript f represents the fixed-point state), for a f After substituting into the fixed-point calculation module, the calculation is converted to look-up table and then power-of-2 and multiplication operations are performed, and the calculation result γ of the weight coefficient can be quickly obtained f , and the values stored in the look-up table are all the amplified floating-point values, and the variables in the calculation process maintain the precision of floating-point calculation. This module greatly simplifies the calculation process, shortens the calculation time, and is more conducive to deployment on the FPGA processor.
[0069] Box filter processing module: Statistically sum the data within a domain space and calculate the mean value. As can be seen from Figure 6 , the calculation processes of box filter 1, box filter 2, and the mean value calculation within the neighborhood block when calculating the variance all have certain similarities and repetitions. Construct a neighborhood matrix, and then sum up all the data in the matrix and take the average. This calculation process is relatively simple but time-consuming, because for each pixel value, a matrix needs to be constructed once, and then the values in the matrix are summed up and averaged. Usually, in a serial processor system, for a neighborhood space of (2r + 1) * (2r + 1) (in this invention, the neighborhood radius r = 2, and the neighborhood space size is a 5 * 5 matrix) to calculate the mean value, it is necessary to read the pixel values (25 values) of a 5 * 5 matrix from the memory centered on this pixel for calculation, and the time-consuming is mainly consumed in memory reading and writing. Therefore, for an image array with a resolution of 640 * 512 or larger, if each pixel point is processed in this way, the time-consuming is relatively large and cannot meet the real-time requirement. Therefore, in this invention, an efficient, flexibly configurable bit width, and reusable box filter module is designed. As Figure 8As shown, in the box filter module, image data flows into the box in the form of row and column data streams in sequence. Except for the first construction and filling of the neighborhood space (5*5 matrix), which has a delay time of 4 rows + 4 columns, as the data stream flows in subsequently, for each data that comes into the box, a calculation result is completed and output. There is always data in the box being calculated and output, without the process of reading and loading data from memory, thus greatly reducing the calculation delay. At the same time, in the present invention, in the steps of calculating box filter 1, box filter 2, and variance calculation, the bit widths of the data filtered and input into the box are different. The bit width of the input image pixel data data is N. After constructing the neighborhood space and accumulating and summing the 25 pixel points to obtain sum, the bit width of sum is extended by 5 bits based on the bit width of data, which is N + 5 bits. Finally, when calculating the average result, the data bit width is restored to N bits. Therefore, in engineering applications, by configuring different bit width parameters, this box filter module can be applied to different calculation steps and processes without repeated design and coding.
[0070] Pipeline acceleration calculation module: Parallel processing of calculation tasks. The high parallelism and low latency efficiency benefit from the pipeline acceleration calculation module, such as Figure 6 the calculation processes of neighborhood mean value, mean square value, variance, linear coefficient, mean square error, box filter, etc. These calculation processes on a serial processor
[0071] are in sequence, and after the previous step is completed, the next step is executed. At the beginning and end of each execution step, there are processes of reloading data and writing back to memory. When executed serially like this, the calculation time or delay of the system is not only the accumulation of the delays of each execution step, but also the time for loading and writing back the intermediate process data, resulting in a relatively large calculation time for the system, as shown in Figure 9 (a) shown. However, in the pipeline acceleration calculation module, each calculation step can be executed simultaneously, as shown in Figure 9As shown in Figure (b), each calculation process can partially overlap. After the data calculation of the first calculation process is completed, the result is passed to the second calculation process for calculation, while the first calculation process receives the input calculation of the next data. Similarly, the second and third calculation processes follow this pattern. In the pipeline acceleration calculation structure, the entire image data is divided into row and column data and flows into the pipeline structure in sequence. When the first calculation process calculates the data of the first row and first column, the second calculation process can simultaneously calculate the data of the first row and second column, and the third calculation process can calculate the data of the first row and third column, and the calculation processes proceed in sequence. In a local time period, the data of each pixel point is executed sequentially one after another, while in a certain period of time, each pixel point is processed simultaneously, and there is no process of loading and writing back intermediate results. Therefore, the time consumption or delay for calculating each pixel point is the cumulative sum of the delays of the non-overlapping parts between the previous calculation module and the next calculation module, and the overall delay can be made very small, and the calculation efficiency is very high.
[0072] Delay optimization module: Analyze and optimize the time consumption of the system to reduce the data processing delay. In the system delay optimization module, further optimize the time consumption analysis of the system to reduce the data processing delay. From Figure 6 It can be seen from the overall weighted aggregation process schematic diagram that there are many calculation steps, but the overall data flow and intermediate process are clear. The states and processes that need to be calculated involve mean, mean square value, variance, linear coefficients a, b, weight coefficient r, box filtering, and final aggregation, etc. Therefore, in this data flow process, for the variables required for the final aggregation and that are not related to each other, further parallel processing can be carried out. Here, after the calculation of the linear coefficients a, b and the weight coefficient r is completed, further parallel processing is carried out on the product results of ra and rb. In box filtering 2, the box filtering is simultaneously performed on the product results of ra and rb to obtain mean(ra) and mean(rb), and finally they are aligned and enter the result aggregation calculation. This can be extended to multiple variables. If the algorithm involves multi-dimensional variables and they are not related in time, parallel processing can be carried out to reduce the impact of the calculation time consumption of each variable on the system delay. For example Figure 10Schematic diagram of the delay optimization module analysis. The main parts with relatively large delays in the system are the summation in the 5*5 neighborhood module, as well as the box filter 1 and box filter 2 modules. For other intermediate calculation processes and result calculation processes, they are all mathematical operation processes such as addition, subtraction, multiplication, and division, and the consumed hardware resources and delay times are both fixed and relatively small. Before the result calculation, the input data needs to be delayed by 8line + 34clk to align with the weight data output by the box filter 2 module, and then multiply and aggregate with the weights output by the box filter for the final result output. Through statistics of the entire process, the delay time for the system to calculate the output of the first pixel is 8line + 40clk. When the size of the image data is 384*288, that is, when a new guided filter weighted aggregation processing operation is performed, the time to delay the output of the first pixel point is (8 * 384 + 40)clk = 3112clk. Then, it continuously outputs until the processing of the last pixel point is completed, and the delay time is very short and the efficiency is very high.
[0073] Those skilled in the art know that in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be regarded as a kind of hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structures within the hardware component; it can also be regarded that the devices, modules, and units for implementing various functions are both software modules for implementing the method and the structures within the hardware component.
[0074] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.
[0075] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments and features in the embodiments of the present application can be arbitrarily combined with each other.
Claims
1. A guided filtering weighted aggregation method based on FPGA, characterized in that, It includes the following steps: Step S1: Process the image using guided filter denoising to obtain the linear regression coefficients of the neighborhood blocks of pixel points; Step S2: Configure weights for different neighborhoods of pixel points according to the linear regression coefficients, and perform weighted aggregation on the new weights to generate a denoised image; In the said Step S2, weight coefficients are allocated according to the error magnitudes calculated for pixel points within the neighborhood, and the larger the error range, the smaller the weight coefficient; The error is the difference between the value output after denoising the neighborhood using guided filtering and the original true value, and the mean square error e is adopted. i Modify the weights of the neighborhood; N is the total number of pixels in the coefficient window; w i is the i-th neighboring block; the denoised image I ′ is expressed as: p is the point within the filter window; Weight coefficient γ i = exp(-e i . / μ), where μ is a scalar parameter, μ > 0. Guided by the image G as the input image I, the mean square error formula is simplified to obtain: e i = (1 - a i ) 2 σ 2 i = a i ε(1 - a i ) Among them, σ 2 i is the variance of the input image in the i neighborhood, I is the input image, and ε is the regularization parameter.
2. The guided filter weighted aggregation method based on FPGA according to claim 1, wherein: Multiply the linear coefficients a i and b i of each neighborhood according to the new weight coefficients to obtain the re-distributed linear coefficients g a and g b , and multiply and weight the pixel grayscales according to the newly assigned coefficients to obtain the final output result.
3. A system applying the FPGA-based guided filtering weighted aggregation method according to claim 1, characterized in that: It includes the following modules: Fixed-point processing module: Convert the process of calculating weights with exp exponential powers into a process of look-up table and integer-type operations suitable for processing on an FPGA hardware processor; Box filtering processing module: Statistically sum the data within a neighborhood space and calculate the mean value; Pipeline acceleration calculation module: Process calculation tasks in parallel; Delay optimization module: Analyze and optimize the time consumption of the system to reduce data processing delay.
4. The system according to claim 3, wherein: The said box filtering processing module receives image data in the form of row and column data streams and enters it into the box frame, and when one data enters the box frame, one calculation result is output.
5. The system according to claim 3, wherein: The said Pipeline acceleration calculation module synchronously executes each calculation step. After the data in the first calculation process is calculated, the result is passed to the second calculation process for calculation, and at the same time, the first calculation process receives the input calculation of the next data.
6. The system according to claim 5, wherein: The image data is divided into row data and column data and flows into the Pipeline acceleration calculation module in sequence. The said Pipeline acceleration calculation module includes multiple calculation structures, and the image data enters different calculation structures in sequence, and each calculation structure performs calculations synchronously.
Citation Information
Patent Citations
FPGA based guide filter and achieving method thereof
CN104063847A
Implementing device of high-speed guiding filter on FPGA (field programmable gate array) platform
CN104571401A