A hardware architecture system of a bilateral filter and its weight optimization method

By segment fitting the value-domain core weight of the bilateral filter, optimizing the bilateral filter weight, solving the problem of difficult storage consumption and accuracy in the prior art, and achieving more efficient computing and storage.

CN114359116BActive Publication Date: 2025-06-20SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111440496.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-06-20
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

The existing bilateral filtering methods are difficult to effectively compromise between storage consumption and accuracy while reducing calculation consumption.

Method used

By obtaining the filter mask of the bilateral filter, the airspace core weight and the value domain core weight are obtained, and the value domain core weight is segmented and fitted to obtain the optimal value domain core fit weight, thereby optimizing the bilateral filter weight.

Benefits of technology

Reduce storage consumption while retaining accuracy and accelerate the computing process by optimizing the hardware architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359116B_ABST
    Figure CN114359116B_ABST
Patent Text Reader

Abstract

The present invention discloses a hardware architecture system of a bilateral filter and its weight optimization method. The system includes: obtaining a filtering mask of the bilateral filter, and obtaining a spatial domain kernel weight according to the filtering mask, wherein the filtering mask is used to represent the weight of the filter; obtaining a range domain kernel weight according to the filtering mask, and performing a piecewise fitting operation on the range domain kernel weight to obtain an optimal range domain kernel fitting weight; obtaining an optimized bilateral filter weight according to the spatial domain kernel weight and the optimal range domain kernel fitting weight. The present invention reduces the storage consumption while retaining the accuracy by performing a piecewise fitting operation on the range domain kernel weight of the bilateral filter. At the same time, the weight is selected on the hardware architecture, and the weight normalization process is optimized to further achieve an acceleration effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic technology, and particularly to a hardware architecture system of a bilateral filter and a method for optimizing its weights. Background Art

[0002] As a popular research direction in the field of computer vision, image denoising aims to reduce noise in digital images. However, existing denoising methods cannot maintain the characteristics of edges. The bilateral filtering method is limited in application because its weights cannot be simply pre-calculated and stored like Gaussian filtering. Although existing improved bilateral filtering methods reduce the consumption of weight calculation, it is difficult to make a compromise between the storage consumption of weights and accuracy.

[0003] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a hardware architecture system of a bilateral filter and a method for optimizing its weights in view of the above-mentioned defects of the existing technology, aiming to solve the problem that although the existing improved bilateral filtering method reduces the calculation consumption, it is difficult to make a compromise between the storage consumption and accuracy.

[0005] The technical solution adopted by the present invention to solve the problem is as follows:

[0006] In a first aspect, an embodiment of the present invention provides a method for optimizing the weights of a bilateral filter, where the method includes:

[0007] Obtain a filtering mask of the bilateral filter, and obtain a spatial domain kernel weight according to the filtering mask; wherein, the filtering mask is used to represent the weights of the filter;

[0008] Obtain a range domain kernel weight according to the filtering mask, and perform a piecewise fitting operation on the range domain kernel weight to obtain an optimal range domain kernel fitting weight;

[0009] Obtain an optimized bilateral filter weight according to the spatial domain kernel weight and the optimal range domain kernel fitting weight.

[0010] In an implementation manner, the obtaining a spatial domain kernel weight according to the filtering mask includes:

[0011] Obtain a first preset value;

[0012] Obtain a spatial domain kernel weight based on the pixel coordinate values in the filtering mask and the first preset value.

[0013] In an implementation manner, the obtaining a range domain kernel weight according to the filtering mask includes:

[0014] Obtain a second preset value;

[0015] Obtain the range kernel weight based on the pixel coordinate value in the filtering mask and the second preset value.

[0016] In one implementation, the performing piecewise fitting operation on the range kernel weight to obtain the optimal range kernel fitting weight includes:

[0017] Divide the range kernel weight into several segments evenly;

[0018] Based on the least squares method, perform fitting calculation on the range kernel weight in each segment to obtain the optimal range kernel segment fitting weight;

[0019] Form the optimal range kernel fitting weight by combining several optimal range kernel segment fitting weights.

[0020] In one implementation, the obtaining the optimized bilateral filter weight according to the spatial domain kernel weight and the optimal range kernel fitting weight includes:

[0021] Multiply the spatial domain kernel weight by the optimal range kernel fitting weight to obtain the optimized bilateral filter weight.

[0022] In a second aspect, an embodiment of the present invention further provides a hardware architecture system of a bilateral filter based on a bilateral filter weight optimization method, wherein the system includes: a weight selection module, configured to select several normalized optimized bilateral filter weights through a multiplexer;

[0023] A weight operation module, configured to perform weighted summation on pixel values to sum the normalized optimized bilateral filter weights selected by the multiplexer;

[0024] A divider module, configured to convert the division operation of a 16-bit dividend and an 8-bit divisor into a multiplication operation of an 8-bit multiplier and an 18-bit multiplier.

[0025] In one implementation, the weight operation module includes several selectors, several adder trees, and several multipliers.

[0026] In one implementation, the divider module includes several lookup tables, a multi-input selector, an adder, and several multipliers.

[0027] In a third aspect, an embodiment of the present invention further provides an intelligent terminal, including a memory, and one or more programs, wherein one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include instructions for executing the bilateral filter weight optimization method as described in any one of the above.

[0028] Fourthly, an embodiment of the present invention further provides a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the bilateral filter weight optimization method described in any one of the above.

[0029] Advantages of the present invention: In the embodiments of the present invention, first, a filtering mask of a bilateral filter is obtained, and an airspace kernel weight is obtained according to the filtering mask, where the filtering mask is used to represent the weight of the filter. Then, a range kernel weight is obtained according to the filtering mask, and a piecewise fitting operation is performed on the range kernel weight to obtain an optimal range kernel fitting weight. Finally, an optimized bilateral filter weight is obtained according to the airspace kernel weight and the optimal range kernel fitting weight. It can be seen that in the embodiments of the present invention, by performing a piecewise fitting operation on the range kernel weight of the bilateral filter, the storage consumption is reduced while maintaining the accuracy. At the same time, the weights are selected in the hardware architecture, and the weight normalization process is optimized to further achieve an acceleration effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0031] Figure 1 Schematic flowchart of the bilateral filter weight optimization method provided by the embodiment of the present invention.

[0032] Figure 2 Frame diagram of the piecewise fitting method in an implementation manner provided by the embodiment of the present invention.

[0033] Figure 3 Graph of the relationship between the output absolute difference and the pixel difference (W = 7.32) in an implementation manner provided by the embodiment of the present invention.

[0034] Figure 4 Range kernel function image in an implementation manner provided by the embodiment of the present invention.

[0035] Figure 5 Overall framework diagram of the bilateral filter system provided by the embodiment of the present invention.

[0036] Figure 6 Architecture diagram of the weight selection module in an implementation manner provided by the embodiment of the present invention.

[0037] Figure 7The architecture diagram of the pixel value weighted summation and weight sum calculation module for an implementation manner provided by an embodiment of the present invention.

[0038] Figure 8 The architecture diagram of the divider module for an implementation manner provided by an embodiment of the present invention.

[0039] Figure 9 The internal structure principle block diagram of the intelligent terminal provided by an embodiment of the present invention. Detailed implementation manners

[0040] The present invention discloses a hardware architecture system of a bilateral filter and its weight optimization method. To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0041] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0042] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0043] In the prior art, Gaussian filtering and median filtering are typical and simple methods for reducing Gaussian noise and salt-and-pepper noise. Among them, the bilateral filter was proposed to protect the smooth edges and details of the Gaussian filter during the denoising process. It is a local non-linear filter that takes into account geometric distance and radiometric difference and has the property of preserving edges. In addition, more advanced denoising methods have been proposed, such as the non-local mean algorithm, sparse module, etc. Deep learning techniques have been applied to many of the latest denoising methods. However, more advanced algorithms such as deep learning methods or non-local methods have a high computational cost. The bilateral filter can well preserve the edge information of the image, and the algorithm complexity is suitable for real-time processing. However, since its weights cannot be simply pre-computed and stored like Gaussian filtering, the utilization of the bilateral filter is limited, especially in application scenarios with high resolution or severe noise. To accelerate the bilateral filter to break this limitation, many hardware architecture-based designs have been proposed: the hardware design architectures are mainly divided into two categories: a. Simplify the calculation of bilateral filter weights. b. Pre-compute the weights and cache them in a look-up table (LUT), and replace the weight calculation by looking up. For the first method, it can significantly reduce the computational complexity, but compared with the standard bilateral filter, it will also cause certain errors. For the second method, storing the weights in the look-up table can significantly reduce the computational consumption, but the trade-off between storage consumption and accuracy still needs to be considered: more weights need to be stored to achieve higher accuracy.

[0044] To solve the problems of the prior art, this embodiment provides a hardware architecture system for a bilateral filter and its weight optimization method. By the above method, the range kernel weights of the bilateral filter are subjected to piecewise fitting operations, so as to reduce the storage consumption while retaining the accuracy. At the same time, the weights are selected on the hardware architecture, and the weight normalization process is optimized to further achieve the acceleration effect. Specifically in implementation, first obtain the filtering mask of the bilateral filter, and obtain the spatial domain kernel weights according to the filtering mask; wherein, the filtering mask is used to represent the weights of the filter; then obtain the range kernel weights according to the filtering mask, and perform piecewise fitting operations on the range kernel weights to obtain the optimal range kernel fitting weights; finally, obtain the optimized bilateral filter weights according to the spatial domain kernel weights and the optimal range kernel fitting weights.

[0045] Exemplary method

[0046] This embodiment provides a method for optimizing the weights of a bilateral filter, and this method can be applied to intelligent terminals in electronic technology. Specifically as Figure 1 shown, the method includes:

[0047] Step S100: Obtain the filtering mask of the bilateral filter, and based on the filtering mask, obtain the spatial domain kernel weights; wherein, the filtering mask is used to represent the weights of the filter.

[0048] Specifically, the algorithm principle of the bilateral filter (abbreviated as BF) is as follows: The bilateral filter consists of a spatial domain kernel and a range domain kernel. The spatial domain kernel acts as a low-pass filter to reduce noise by averaging the pixels in the filtering window, and the weights are determined by the Euclidean distance from the pixel to the center. The spatial domain kernel is based on the Gaussian function, and the formula of the spatial domain kernel is shown in (1):

[0049]

[0050] where (k, l) are the coordinates of the central pixel, and (i, j) are the coordinates of the other pixels. σ s is the standard deviation of the Gaussian function.

[0051] The role of the range domain kernel is to protect the edges of the image. The weights are determined by the intensity difference between the pixel under consideration and the central pixel of the window, representing the similarity degree of these two pixels. The range domain kernel is shown in formula (2):

[0052]

[0053] where f(k, l) is the intensity of the central pixel, f(i, j) is the intensity of the other pixels, and σ r is the standard deviation of the Gaussian function.

[0054] The weights of BF are obtained by multiplying the spatial domain kernel and the range domain kernel. The weights of BF are shown in formula (3):

[0055] w(i, j, k, l) = g s (i, j, k, l) · g r (i, j, k, l) (3)

[0056] This filter denoises through the spatial domain kernel g s similar to a Gaussian filter, and preserves the edges through the range domain kernel g r The range domain kernel g r reduces the weights as the gray difference increases. Therefore, the edge pixels will play a smaller role in the smoothing process. Compared with the brightness difference caused by noise, the edge pixels have a larger brightness difference.

[0057] The complete output of the bilateral filter is shown in formula (4):

[0058]

[0059] where f(k, l) is the intensity of the central pixel, ∑ k,lThe sum of the products of the intensity of the central pixel and the weights of the BF, ∑ k,l The sum of the weights of the BF is obtained by summing up the weights of the BF. The piecewise fitting method of the present invention does not simplify the bilateral filter to reduce the calculation cost, nor pre-compute the weights of all filters to avoid calculation, but pre-computes the Gaussian function and approximately calculates several fitting points as weights. In this way, both the optimization of the calculation cost and the saving of memory are achieved. By analyzing the action process of the BF weights, the least squares method is used to fit as few weights as possible. Therefore, in practice, an n×n filter mask of a bilateral filter needs to be obtained first. In this embodiment, a 5×5 mask is obtained, and then the spatial domain kernel weights and the range domain kernel weights can be obtained by calculating the filter mask; wherein, the filter mask is used to represent the weights of the filter. This is to prepare for the subsequent piecewise fitting operation of the range domain kernel weights.

[0060] To obtain the spatial domain kernel weights, obtaining the spatial domain kernel weights according to the filter mask includes the following steps: obtaining a first preset value; obtaining the range domain kernel weights based on the pixel coordinate values in the filter mask and the second preset value.

[0061] Specifically, to obtain the first preset value, in this embodiment, the first preset value is that the standard deviation of the spatial domain kernel is 1.1, which is determined by the 3σ rule. Then, substituting the pixel coordinate values in the filter mask and the first preset value into a preset first formula, the spatial domain kernel weights can be obtained. In this embodiment, the preset first formula is formula (1). The pixel coordinate values in the filter mask include the central pixel coordinate values (k, l) and the coordinate values (i, j) of the remaining pixels. The standard deviation σ of the spatial domain kernel s is 1.1, and the spatial domain kernel weights g can be obtained through formula (1). s .

[0062] To obtain the range domain kernel weights, obtaining the range domain kernel weights according to the filter mask includes the following steps: obtaining a second preset value; obtaining the range domain kernel weights based on the pixel coordinate values in the filter mask and the second preset value.

[0063] First, obtain the second preset value. In this embodiment, the second preset value is the standard deviation of the range domain kernel, with a value of 15. The preset second formula is formula (2). The pixel coordinate values in the filter mask include the central pixel coordinate values (k, l) and the coordinate values (i, j) of the remaining pixels. The range domain kernel weights g can be obtained through formula (2). r .

[0064] After obtaining the range domain kernel weights, the following can be executed as Figure 1The following steps are shown: S200. Perform piecewise fitting operation on the value range kernel weight to obtain the optimal value range kernel fitting weight. Correspondingly, the step of performing piecewise fitting operation on the value range kernel weight to obtain the optimal value range kernel fitting weight includes the following steps: Divide the value range kernel weight into several segments evenly; Based on the least squares method, perform fitting calculation on the value range kernel weight in each segment to obtain the optimal value range kernel segment fitting weight; Several of the optimal value range kernel segment fitting weights form the optimal value range kernel fitting weight.

[0065] Specifically, as Figure 2 shown, in the case of 8-bit pixel values, the pixel value range is from 0 to 255, so there are 256 value range kernel weights. Using the least squares method, first find the points after the critical point, that is, the points in the value range kernel, and then perform piecewise fitting on the value range kernel weight. In this embodiment, in order to determine the critical point, an effect analysis is performed on a specific range standard deviation. First, consider performing bilateral filtering calculation on any pixel with intensity f(k, l). When there is a Δx difference between this pixel and other pixels in the mask, where Δx is ||f(i, j)-f(k, l)|| and the range of Δx is (0 - 255), from Equation (4), the output of the bilateral filter BF with a Δx pixel difference is I′(i, j), as shown in Equation (5):

[0066]

[0067] where f(k, l) is the intensity of the central pixel, ∑ k,l (f(k, l) ± Δx)w(i, j, k, l) is the cumulative sum of the product of the intensity and the weight of BF for pixels with a difference of Δx from the central pixel, and ∑ k,l w(i, j, k, l) is the cumulative sum of the weights of BF. Therefore, the influence on the output (|ΔI|) can be deduced from Equations (4) and (5), as shown in Equation (6):

[0068] |I′(i, j)-I(i, j)| = |ΔI| (6)

[0069] where I'(i, j) is the complete output of the bilateral filter BF, I(i, j) is the output of the bilateral filter BF with a Δx pixel difference, the value of the fixed weight w, and the relationship between ΔI and Δx is as Figure 3As shown. It can be seen that there is a peak value, which is the critical point. Before this point, the spatial kernel is dominant, so the larger the pixel difference, the greater the impact on the result; after the critical point, the value domain kernel is dominant, so the larger the pixel difference, the smaller the impact on the result, so the edge of the image can be protected. After obtaining the critical point, the value domain kernel weight is fitted using the least squares method. Since the left side of the critical point has little effect on the result dominated by the spatial kernel, as few fitting points as possible can be selected. On the right side of the critical point, the fitting range weight curve needs to be refined. Therefore, the number of fitting points on the left can be less than the number of fitting points on the right. In formula (2), Δx, that is, ||f(i,j)-f(k,l)||, is the pixel difference, which is also the horizontal coordinate value in the function graph, and the vertical coordinate is the weight g(x). The function graph is divided into two parts, left and right, by the critical point. In this embodiment, taking 6 fitting values ​​as an example, the number of fitting points is taken in a ratio of 1:2, that is, 2 points are fitted on the left and 4 points are fitted on the right. In practice, the range kernel weight is divided into several segments: the left part is divided into two parts according to the change value of g(x), and the right part is divided into four parts according to the change value of g(x). Figure 4 Then, based on the least squares method, the range kernel weight in each segment is fitted and calculated to obtain the optimal range kernel segment fitting weight; first, assume that the optimal fitting point of each segment is G i , the corresponding coordinates of the segmentation points on the x-axis are [k1-k6]. g(x) is the value of x in (k i-1 , k i ) The weight corresponding to a point in the range. The optimal weight G i It is obtained by the least square method in formula (7):

[0070] min∑(G i -g(x)) (7)

[0071] That is, (k i-1 , k i ) and G i When the absolute value of the difference is accumulated and the value is the smallest, the G i That is (k i-1 , k i ) is the optimal solution in the segment. Finally, several of the optimal value domain kernel segment fitting weights G i Composition of optimal range kernel fitting weights.

[0072] After obtaining the optimal range kernel fitting weights, you can execute Figure 1The following steps are shown: S300. Obtain an optimized bilateral filter weight according to the spatial domain kernel weight and the optimal range domain kernel fitting weight. Correspondingly, the step of obtaining an optimized bilateral filter weight according to the spatial domain kernel weight and the optimal range domain kernel fitting weight includes the following steps: Multiply the spatial domain kernel weight by the optimal range domain kernel fitting weight to obtain an optimized bilateral filter weight. For example, the optimized bilateral filter weight w i = g s * G i . In practice, the optimized bilateral filter weight is a floating-point number and must be converted to an integer when input into the hardware system of the bilateral filter. Therefore, according to formula (8), convert the floating-point value of the optimized bilateral filter weight to an integer:

[0073]

[0074] where w i ′ is the optimized bilateral filter weight after normalization, w i is the optimized bilateral filter weight, ∑ k,l w(i, j, k, l) is the sum of the optimized bilateral filter weights, W max is the maximum number when quantifying 1 to an integer with a bit width of n. For example, assume that the bit width of the integer after converting the optimized bilateral filter weight is 8, then W max is 255.

[0075] Exemplary device

[0076] As Figure 5 shown in, an embodiment of the present invention provides a hardware architecture system of a bilateral filter for a bilateral filter weight optimization method. The system includes: a weight selection module for selecting a plurality of optimized bilateral filter weights after normalization through a multiplexer;

[0077] a weight operation module for performing weighted summation on pixel values and summing the optimized bilateral filter weights after normalization selected by the multiplexer;

[0078] a divider module for converting a division operation of a 16-bit dividend and an 8-bit divisor into a multiplication operation of an 8-bit multiplier and an 18-bit multiplier.

[0079] Specifically, the architecture includes three modules: a weight selection module, a weight operation module, and a divider module. For a bilateral filter mask of size m×m with a pixel-level pipeline, the input pixel stream is transmitted to the filter mask module through an m - 1 row buffer. The filter mask module consists of m×m shift registers, and each shift register is connected to the weight operation module respectively. The hardware architecture system of the bilateral filter further includes a weight selection module for selecting a number of optimized bilateral filter weights after normalization through a multiplexer; in one implementation, the weight operation module includes a number of selectors, a number of adder trees, and a number of multipliers. In this embodiment, the calculation performance is further improved by storing approximate weights. Through the piecewise fitting method, the number of fitting points of the range weights has been significantly reduced, and at the same time, the storage capacity of the LUT has also been reduced. Six fitting points are the minimum number of range kernel weights. As Figure 6 shown, by comparing the absolute pixel difference and the value of the pixel difference (Δx), six optimized bilateral filter weights after normalization, that is, approximate bilateral weights (Approx BilateralWeight, abbreviated as ABW), are selected by the multiplexer according to the piecewise interval values of (k i-1 , k i ), where the piecewise interval value is the rounded value. The weight operation module is used to perform weighted summation on the pixel values and sum the optimized bilateral filter weights after normalization selected by the multiplexer; in one implementation, the weight operation module includes a number of selectors, a number of adder trees, and a number of multipliers. In practice, the size of the filter mask determines the number of adders, multipliers, and the number of parallel weight selection modules. A larger filter mask will not only cause an increase in the delay of the critical path but also result in a better smoothing effect. Therefore, this is a trade-off between performance and cost. In this embodiment, considering resource utilization and performance, we selected a 5×5 filter mask. The 5×5 filter window is divided into 5 linear summation (line_sum) columns, each column contains 5 weight selection modules, and each module selects the weights required for the corresponding pixel calculation. Then, as Figure 7As shown, the product of each pixel and the optimized bilateral filter weights after normalization is calculated, and the weighted sum of the pixel values and the sum of the optimized bilateral filter weights after normalization selected by the multiplexer are calculated to meet the throughput of the pixel-level pipeline. In this weight calculation module, the main resource consumption and critical path come from the multiplier and adder, because both the weighted sum of the pixel values and the sum of the optimized bilateral filter weights after normalization selected by the multiplexer require an adder tree for accumulation. To more effectively add these accumulated values, the adder tree is designed as a highly parallel two-stage structure, with each stage being a five-input one-output adder tree. The hardware architecture system of the bilateral filter further includes: a divider module for converting the division operation of a 16-bit dividend and an 8-bit divisor into a multiplication operation of an 8-bit multiplier and an 18-bit multiplier. In one implementation, as Figure 8 shown, the divider module includes several lookup tables, a multi-input selector, an adder, and several multipliers. Since general dividers usually consume a large amount of hardware resources and have a high latency, the present invention designs a lookup table-based divider to convert the division operation of a 16-bit dividend and an 8-bit divisor into a multiplication operation of an 8-bit multiplier and an 18-bit multiplier. Figure 8 Fig. shows the structure of the divider, where the value after weighted summation of the pixel values input is a 16-bit unsigned number, and the sum of the optimized bilateral filter weights after normalization input is an 8-bit unsigned number.

[0080] For the sum of the optimized bilateral filter weights after normalization, eight LUTs and an eight-input selector are used to convert the sum of the optimized bilateral filter weights after normalization into its reciprocal. Since multipliers consume much fewer resources than dividers, the divisor B (the sum of the optimized bilateral filter weights after normalization) is converted into a fraction to change the division into multiplication. For the eight LUTs, each LUT stores 16 reciprocals, which are selected by the last four bits of the weight sum (2 4 = 16, four bits can store 16 reciprocals), and the eight LUTs respectively correspond to the first three bits of the sum of the optimized bilateral filter weights after normalization (2 3 = 8, three bits correspond to eight LUTs). Therefore, the divider divides the lookup process into two parts: selecting one of the eight LUTs through the first three bits of the sum of the optimized bilateral filter weights after normalization, and selecting the corresponding reciprocal from the 16 reciprocals stored in the LUT by the last four bits of the sum of the optimized bilateral filter weights after normalization to improve the path speed.

[0081] The number obtained by rounding the reciprocal of the sum of the optimized bilateral filter weights after 8-bit normalization to 18-bit unsigned is converted, while the value obtained by weighted summation of pixel values is a 16-bit unsigned number and remains unchanged. After conversion, the multiplier with 18-bit and 16-bit input bits is expanded into two multipliers with 18-bit and 8-bit input bits, which is equivalent to changing a large multiplier into two small multipliers, thus shortening the critical path of the multiplication operation process. The expansion process is as shown in formula (9):

[0082]

[0083] In the above formula, a and b represent two multipliers, and n is the bit width of a. It can be seen from the above formula that the multiplier is divided into two. Among them, the high n bits of one multiplication are a, and are the remaining low bits of a. In this embodiment, the value obtained by weighted accumulation of 16-bit pixel values is divided into two 8-bit multipliers. Finally, the output pixel stream of this architecture is connected to a selector. If the sum of weights is zero, the selector is equal to zero, otherwise it is equal to the output of the divider module. Therefore, the LUT-based frequency divider significantly reduces the latency period of the data stream of the BF. Although parallelization requires clock cycles, the remaining calculations only require 9 cycles. Among them, the divider module in this embodiment only has 4 latency periods, and the latency period is smaller than the traditional latency period.

[0084] Based on the above embodiments, the present invention also provides an intelligent terminal, and its principle block diagram can be as Figure 9 shown. The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a temperature sensor connected through a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a method for optimizing bilateral filter weights. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen. The temperature sensor of the intelligent terminal is pre-set inside the intelligent terminal and is used to detect the operating temperature of internal devices.

[0085] Those skilled in the art can understand that Figure 9 in the schematic diagram, it is only a block diagram of part of the structure related to the solution of the present invention, and does not constitute a limitation on the intelligent terminal to which the solution of the present invention is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0086] In one embodiment, an intelligent terminal is provided, including a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: obtaining a filtering mask of a bilateral filter, and obtaining a spatial domain kernel weight according to the filtering mask; wherein the filtering mask is used to characterize the weight of the filter.

[0087] Obtaining a range domain kernel weight according to the filtering mask, and performing a piecewise fitting operation on the range domain kernel weight to obtain an optimal range domain kernel fitting weight.

[0088] Obtaining an optimized bilateral filter weight according to the spatial domain kernel weight and the optimal range domain kernel fitting weight.

[0089] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0090] In summary, the present invention discloses a hardware architecture system of a bilateral filter and its weight optimization method. The method includes: obtaining a filtering mask of the bilateral filter, and obtaining a spatial domain kernel weight according to the filtering mask, where the filtering mask is used to represent the weight of the filter; obtaining a range domain kernel weight according to the filtering mask, and performing a piecewise fitting operation on the range domain kernel weight to obtain an optimal range domain kernel fitting weight; obtaining an optimized bilateral filter weight according to the spatial domain kernel weight and the optimal range domain kernel fitting weight. By performing a piecewise fitting operation on the range domain kernel weight of the bilateral filter, the present invention reduces the storage consumption while retaining the accuracy. At the same time, the weights are selected in the hardware architecture, and the weight normalization process is optimized to further achieve an acceleration effect.

[0091] Based on the above embodiments, the present invention discloses a method for optimizing the weights of a bilateral filter. It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A bilateral filter weight optimization method, characterized in that, The method includes: Obtaining a filtering mask of a bilateral filter, and obtaining a spatial domain kernel weight according to the filtering mask; wherein, the filtering mask is used to represent the weight of the filter; Obtaining a range domain kernel weight according to the filtering mask, and performing a piecewise fitting operation on the range domain kernel weight to obtain an optimal range domain kernel fitting weight; Obtaining an optimized bilateral filter weight according to the spatial domain kernel weight and the optimal range domain kernel fitting weight; The obtaining the range domain kernel weight according to the filtering mask includes: Obtaining a second preset value; Obtaining a range domain kernel weight based on the pixel coordinate value in the filtering mask and the second preset value; The second preset value is the standard deviation of the range domain kernel, and the pixel coordinate value in the filtering mask includes the central pixel coordinate value (k, l) and the coordinate (i, j) values of the remaining pixels; the formula for the range domain kernel weight is: where f(k, l) is the intensity of the central pixel, f(i, j) is the intensity of the remaining pixels, and σ r is the standard deviation of the Gaussian function; The performing a piecewise fitting operation on the range domain kernel weight to obtain an optimal range domain kernel fitting weight includes: Dividing the range domain kernel weight into several segments evenly; Performing a fitting calculation on the range domain kernel weight in each segment based on the least squares method to obtain an optimal range domain kernel segment fitting weight; Forming an optimal range domain kernel fitting weight by several of the optimal range domain kernel segment fitting weights.

2. The bilateral filter weight optimization method according to claim 1, characterized in that, The obtaining the spatial domain kernel weight according to the filtering mask includes: Obtaining a first preset value; Obtaining a spatial domain kernel weight based on the pixel coordinate value in the filtering mask and the first preset value.

3. The bilateral filter weight optimization method according to claim 1, characterized in that, The obtaining an optimized bilateral filter weight according to the spatial domain kernel weight and the optimal range domain kernel fitting weight includes: Multiplying the spatial domain kernel weight by the optimal range domain kernel fitting weight to obtain an optimized bilateral filter weight.

4. A hardware architecture system of a bilateral filter based on the bilateral filter weight optimization method according to any one of claims 1-3, characterized in that, The system includes: A weight selection module, configured to select several normalized optimized bilateral filter weights through a multiplexer; A weight operation module, configured to perform a weighted sum on pixel values to sum the normalized optimized bilateral filter weights selected by the multiplexer; A divider module, configured to convert a division operation of a 16-bit dividend and an 8-bit divisor into a multiplication operation of an 8-bit multiplier and an 18-bit multiplier.

5. The hardware architecture system of the bilateral filter according to the bilateral filter weight optimization method of claim 4, characterized in that, The weight operation module includes several selectors, several adder trees, and several multipliers.

6. The hardware architecture system of the bilateral filter according to the bilateral filter weight optimization method of claim 4, characterized in that, The divider module includes several look-up tables, a multi-input selector, an adder, and several multipliers.

7. An intelligent terminal, characterized in that, It includes a memory, and one or more programs, wherein one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for executing the method according to any one of claims 1-3.

8. A non-transitory computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Guided-filtering optimization speed-up method based on CUDA

    CN104899840A

  • Telemetering spectrum noise suppression method suitable for infrared small target recognition

    CN105654447A