NLM image processing method and device based on FPGA
By implementing an NLM-based image processing method on the FPGA platform and using a pulsating processor array for streaming operations, the problems of slow computing speed and high hardware power consumption in the traditional BM3D algorithm are solved, and efficient and low-power image noise reduction processing is achieved.
Patent Information
- Application Number
- CN202510297486.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional BM3D algorithms are slow in image noise reduction processing, difficult to process large-scale images, and have high hardware power consumption, making them not suitable for embedded scenarios.
Using the NLM image processing method based on FPGA, the image matrix is operated through the pulsating processor array, variance calculation and weight calculation are performed, resource overhead and logical series are reduced, and algorithm accuracy is improved.
It realizes the reduction of resource overhead and logic series, improves algorithm accuracy, and has low hardware power consumption, which is suitable for real-time processing and embedded scenarios.
Smart Images

Figure CN120147102A_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of image processing, and particularly relates to a non-local means (NLM) image processing method and device based on a field programmable gate array (FPGA). Background Art
[0002] In the imaging, processing, and transmission of images, noise is inevitably introduced. These situations make the rapid and high-quality image denoising function become increasingly important. At the same time, since both the image noise spectrum and the image detail spectrum are parts of the high-frequency components of the image, it is necessary to suppress noise and retain the image edges as much as possible during denoising, which increases the processing difficulty of image denoising.
[0003] Among the common denoising algorithms, the BM3D algorithm is a traditional denoising algorithm with the best performance. By separately processing the spatial and temporal domains, it can effectively reduce the static noise and spatial Gaussian white noise of the image, better retain the motion changes and edge details, and improve the image imaging quality. The spatial and temporal domain denoising operators of BM3D mostly use the NLM operator to implement. The NLM operator performs image filtering through the variance weight of the matching block and the reference block within a search interval to achieve the purpose of suppressing noise and retaining edges. However, the traditional BM3D algorithm has a slow calculation speed, is difficult to process large-scale images, and at the same time, the hardware power consumption for operation is relatively high, which is not suitable for embedded scenarios. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a non-local means (NLM) image processing method and device based on a field programmable gate array (FPGA) to reduce resource overhead, reduce the logic level, and improve the algorithm accuracy.
[0005] To solve the above technical problem, in a first aspect, the present invention provides a non-local means (NLM) image processing method based on a field programmable gate array (FPGA), including: processing the image to be processed into an image matrix; inputting the image matrix into a systolic processor array; the systolic processor array obtaining the accumulated difference index value and the accumulated reference index value of the required image according to the image matrix, and inputting the accumulated difference index value and the accumulated reference index value into a division unit; the division unit dividing the accumulated reference index value by the accumulated difference index value to obtain a filtered pixel value; and processing the image matrix of the image to be processed through the filtered pixel value to obtain an output image.
[0006] Further, the systolic processing array includes a plurality of systolic processors. The systolic processors perform a pipelined operation on the image matrix. After each level of systolic processor calculates and obtains the accumulated difference index value and the accumulated reference index value, the obtained data and the image matrix are delayed and input to the next level of systolic processor for processing. The primary systolic processor calculates according to a preset value and the image matrix.
[0007] Further, the pulsating processor processes variance calculation and weight calculation.
[0008] Further, the variance calculation includes: calculating the variance of the corresponding column vectors at the corresponding positions in the matching block and the reference block in the image matrix, and obtaining the difference index value through exponential operation after cumulative processing of the variance calculation result.
[0009] Further, the weight calculation includes: accumulating the difference index value obtained in the current calculation step and the difference index value obtained in the upper-level calculation to obtain the cumulative difference index value; accumulating the reference index value obtained in the current calculation step and the reference index value obtained in the upper-level calculation to obtain the cumulative reference index value.
[0010] Further, when performing the variance calculation and the weight calculation, the timing data of the current calculation is also recorded, and the timing data adjusts the operation steps of the pulsating processor and enables the processed image data to be output in time sequence.
[0011] Further, the number of the pulsating processors is determined according to the range of the image search interval.
[0012] In a second aspect, the present invention provides an FPGA-based NLM image processing apparatus, including: a matrix mapping unit: configured to process the image to be processed into an image matrix; a pulsating processor array: configured to include a plurality of pulsating processors, and each pulsating processor includes a variance calculation unit and a weight calculation unit; a water flow manager: configured to manage the input and output data and the corresponding timing data of the pulsating processors; a division unit: configured to obtain the filtered pixel value through the output result of the pulsating processor array; an image output unit: configured to output the processed image by combining the timing data, the filtered pixel value, and the image to be processed.
[0013] Further, the pulsating processor array further includes an accumulator, and the accumulator is used for cumulative calculation of the operation results at each level.
[0014] Further, the variance calculation unit further includes a difference calculation unit and an exponential operation unit; the difference calculation unit sends the difference between the variance calculation result and the statistical variance of the image noise to the exponential operation unit; the exponential operation unit outputs the difference index value.
[0015] Compared with the prior art, the present invention has the following beneficial effects: reducing resource overhead, decomposing the NLM operator for calculation, multiplexing adders and multipliers, and reducing the hardware area overhead. Improving the timing convergence through timing data, enabling the NLM to perform pipelined operation, reducing the logic level, and improving the algorithm accuracy. The hardware has low power consumption and low latency, is suitable for real-time processing and embedded scenarios, and can be customized. Description of the Drawings
[0016] The accompanying drawings are included to provide a further understanding of the present invention, and they are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the present invention and, together with the description, serve to explain the principles of the present invention. In the drawings:
[0017] Figure 1 is a schematic flowchart of the NLM image processing method based on FPGA according to an embodiment of the present invention;
[0018] Figure 2 is a schematic structural diagram of the NLM image processing apparatus based on FPGA according to an embodiment of the present invention;
[0019] Figure 3 is a schematic diagram of the systolic processor array in an embodiment of the present invention. Detailed Embodiments
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, the present invention can also be applied to other similar scenarios based on these drawings. Unless obvious from the language context or otherwise stated, the same reference numerals in the figures represent the same structure or operation.
[0021] It should be understood that when a component is referred to as "on another component", "connected to another component", "coupled to another component", or "in contact with another component", it can be directly on, connected to, or coupled to, or in contact with the other component, or there may be an intervening component. In contrast, when a component is referred to as "directly on another component", "directly connected to", "directly coupled to", or "directly in contact with" another component, there is no intervening component. Similarly, when the first component is referred to as "electrically in contact with" or "electrically coupled to" the second component, there is an electrical path allowing current to flow between the first component and the second component. The electrical path may include capacitors, coupled inductors, and / or other components allowing current to flow, even if there is no direct contact between the conductive components.
[0022] Reference Figure 1As shown, the FPGA-based NLM image processing method in this embodiment mainly includes the following steps: processing the image to be processed into an image matrix; inputting the image matrix into the systolic processor array; the systolic processor array obtaining the accumulated difference index value and the accumulated reference index value of the required image based on the image matrix, and inputting the accumulated difference index value and the accumulated reference index value into the division unit; the division unit dividing the accumulated reference index value by the accumulated difference index value to obtain the filtered pixel value; and processing the image matrix of the image to be processed through the filtered pixel value to obtain the output image.
[0023] In some embodiments, the number of systolic processors is determined according to the range of the image search interval, that is, the required number of systolic processors can be divided according to the size of the search interval required by NLM. Exemplarily, the systolic processing array includes 25 systolic processors, that is, an image matrix with an output of 5×5. The systolic processors perform a pipelined operation on the image matrix. After each level of systolic processor calculates and obtains the accumulated difference index value and the accumulated reference index value, the obtained data and the image matrix are delayed and input to the next-level systolic processor for processing. The primary systolic processor calculates based on the preset value and the image matrix.
[0024] In some embodiments, the processing of the systolic processor includes variance calculation and weight calculation. The variance calculation can be: calculating the variance of the corresponding column vectors in the matching block and the reference block in the image matrix, and performing an exponential operation on the result of the variance calculation after accumulation to obtain the difference index value. The weight calculation can be: accumulating the difference index value obtained in the current calculation step with the difference index value obtained in the upper-level calculation to obtain the accumulated difference index value, and accumulating the reference index value obtained in the current calculation step with the reference index value obtained in the upper-level calculation to obtain the accumulated reference index value. When performing variance calculation and weight calculation, the timing data of the current calculation can also be recorded, and the timing data adjusts the operation steps of the systolic processor and enables the processed image data to be output in sequence.
[0025] In one implementation, the FPGA-based NLM image processing method specifically includes: the image data to be processed generates an image matrix through the matrix mapping unit, and at the same time inputs the matrix into the first-level systolic processor; after receiving the image matrix, the first-level systolic processor performs relevant calculations through the preset value, and at the same time delays and outputs the same image matrix to the second-level systolic processor. When the calculation is completed, the calculation result is output to the second-level systolic processor. Except for the first-level systolic processor, the input of the remaining systolic processors is the output of the previous-level systolic processor; when the calculation result output of the 24th-order systolic processor is valid, the calculation result is sent to the divider for calculation to obtain the filtered pixel value; the image data of the NLM calculation is obtained through the filtered pixel value; the image data is restored to the original image and output in combination with the timing output.
[0026] Another embodiment of the present invention is an FPGA-based NLM image processing device, which mainly includes: a matrix mapping unit: configured to process the image to be processed into an image matrix; a systolic processor array: configured to include a plurality of systolic processors, each systolic processor including a variance calculation unit and a weight calculation unit; a water flow manager: configured to manage the input and output data and corresponding timing data of the systolic processors; a division unit: configured to obtain a filtered pixel value through the output result of the systolic processor array; an image output unit: configured to output the processed image by combining the timing data, the filtered pixel value, and the image to be processed.
[0027] In some embodiments, the systolic processor array further includes an accumulator, which is used for cumulative calculation of the operation results at each level.
[0028] In some embodiments, the variance calculation unit further includes a difference calculation unit and an exponential operation unit. The difference calculation unit sends the difference between the variance calculation result and the statistical variance of the image noise to the exponential operation unit, and the exponential operation unit outputs a difference exponential value.
[0029] The present invention can configure the search interval size, the matching block size, and the statistical variance value of the image noise according to the hardware environment and the algorithm effect, achieving low hardware requirements, simple structure, convenient control, and good processing effects.
[0030] For those skilled in the art, the above disclosure of the invention is only an example and does not constitute a limitation to the present invention. Although not explicitly stated here, those skilled in the art may make various modifications, improvements, and corrections to the present invention. Such modifications, improvements, and corrections are proposed in the present invention, so such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of the present invention.
[0031] Although the present invention has been described with reference to the current specific embodiments, those of ordinary skill in the art in this technical field should recognize that the above embodiments are only used to illustrate the present invention, and various equivalent changes or substitutions can be made without departing from the spirit of the present invention. Therefore, as long as the changes and modifications to the above embodiments fall within the scope of the spirit of the present invention, they will fall within the scope of the claims of the present invention.
Claims
1. A NLM image processing method based on FPGA, characterized in that: include: Process the image to be processed into an image matrix; inputting the image matrix into a systolic processor array; The systolic processor array obtains a difference index accumulation value and a reference index accumulation value of a required image according to the image matrix, and inputs the difference index accumulation value and the reference index accumulation value into a division unit; The division unit divides the reference index accumulated value by the difference index accumulated value to obtain a filtered pixel value; The image matrix of the image to be processed is processed by using the filtered pixel values to obtain an output image.
2. The FPGA-based NLM image processing method according to claim 1, characterized in that: The systolic processing array includes a plurality of systolic processors, and the systolic processors perform pipeline operations on the image matrix. After the systolic processors at each level calculate and obtain the difference index accumulated value and the reference index accumulated value, the obtained data and the image matrix are delayed and input to the next level systolic processor for processing. The primary systolic processor performs calculations based on preset values and the image matrix.
3. The FPGA-based NLM image processing method according to claim 2, characterized in that: The systolic processor processing includes variance calculation and weight calculation.
4. The FPGA-based NLM image processing method according to claim 3, characterized in that: The variance calculation includes: The column vectors corresponding to the positions of the matching block and the reference block in the image matrix are subjected to variance calculation, and the variance calculation results are subjected to exponential operation after accumulation processing to obtain the difference index value.
5. The FPGA-based NLM image processing method according to claim 4, characterized in that: The weight calculation includes: Accumulating the difference index value obtained in the current calculation step and the difference index value obtained in the previous calculation step to obtain the difference index accumulated value; The reference index accumulated value is obtained by accumulating the reference index value obtained in the current calculation step with the reference index value obtained in the previous calculation step.
6. The FPGA-based NLM image processing method according to any one of claims 3 to 5, characterized in that: When performing the variance calculation and the weight calculation, the time series data of the current calculation is also recorded. The time series data adjusts the operation steps of the pulsation processor and enables the processed image data to be output in time series.
7. The FPGA-based NLM image processing method according to claim 2, characterized in that: The number of systolic processors is determined according to the range of the image search interval.
8. An FPGA-based NLM image processing device, characterized in that: include: Matrix mapping unit: configured to process the image to be processed into an image matrix; A systolic processor array is configured to include a plurality of systolic processors, each systolic processor including a variance calculation unit and a weight calculation unit; A water flow manager: configured to manage the input and output data and corresponding timing data of the pulsation processor; A division unit configured to obtain a filtered pixel value through the output result of the systolic processor array; Image output unit: configured to output a processed image through the time series data, the filtered pixel values and the image to be processed.
9. The FPGA-based NLM image processing device according to claim 8, characterized in that: The systolic processor array further includes an accumulator, which is used for accumulating calculations of operation results at each level.
10. The FPGA-based NLM image processing device according to claim 8, characterized in that: The variance calculation unit also includes a difference calculation unit and an exponential operation unit; the difference calculation unit sends the difference between the variance calculation result and the image noise statistical variance to the exponential operation unit; the exponential operation unit outputs the difference exponential value.