Image Processing Method and Apparatus, Electronic Device, and Storage Medium
By processing median filtering of images in parallel on multi-threaded hardware, combined with the sorting of rows, columns and diagonal directions, the problem of high complexity of the median filtering method in large filtering windows is solved, and more efficient image processing is achieved.
Patent Information
- Application Number
- CN202510187383.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The existing median filtering method has high complexity in large filtering windows, resulting in poor performance and is not suitable for parallel computing on multi-threaded hardware computing resources.
The median filtering operation is performed in parallel by multi-threaded hardware. By dividing the input images into multiple block images, and using multiple hardware execution units to perform median filtering, combining the sorting of rows, columns and diagonal directions, pixel points that cannot be median are removed, and the median filtering result of each pixel point is determined.
The processing speed and execution efficiency of the median filtering algorithm are improved, especially in the case of a large filtering window, which significantly improves performance and makes full use of multi-threaded hardware computing resources.
Smart Images

Figure CN119648575B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to an image processing method, a processor, an electronic device, and a non-transitory computer-readable storage medium. Background Art
[0002] Median filtering is a commonly used signal processing technique, mainly used to remove noise in data, especially salt-and-pepper noise. Salt-and-pepper noise is a randomly occurring black-and-white pixel points, looking like salt and pepper sprinkled on the image.
[0003] The working principle of median filtering is to replace the value of each pixel point with the median of all pixel point values in the neighborhood of the pixel point. Here, the "neighborhood" can be a 3x3, 5x5 or other sized matrix, specifically depending on the size of the filtering window. Median filtering does not blur the image like some low-pass filters, so it is very effective in removing noise while preserving the image edges. Summary of the Invention
[0004] At least one embodiment of the present disclosure provides an image processing method, including: obtaining M images to be processed, where the M images to be processed are obtained by respectively performing border expansion on M segmented images, the M segmented images are obtained by evenly dividing an input image into M blocks in the column direction, and M is a positive integer; respectively performing median filtering operations on the M images to be processed by M hardware execution units to obtain a median filtering result of the input image, where each hardware execution unit includes L thread hardwares, and the L thread hardwares perform the median filtering operation on the image to be processed input to the hardware execution unit in parallel, and each thread hardware performs the median filtering operation on n pixel points at a single time, and n is a positive integer, where performing the median filtering operation on the n pixel points by the thread hardware includes: determining a filtering region corresponding to the n pixel points; sorting the pixel points in the filtering region in the row direction, column direction, and diagonal direction, and removing the pixel points that cannot be the median in the sorted filtering region; determining a median filtering result of each pixel point among the n pixel points based on the remaining pixel points.
[0005] For example, in the image processing method provided by at least one embodiment of the present disclosure, sorting the pixel points in the filtering area in the row direction, column direction, and diagonal direction, and removing the pixel points that cannot be the median in the sorted filtering area includes: sorting the pixel points in the filtering area in the row direction and the column direction to obtain a first filtering area, where in the first filtering area, the pixel values of the pixel points in each row are in an increasing order in the row direction, and the pixel values of the pixel points in each column are in an increasing order in the column direction; removing the pixel points in the first filtering area whose coordinates do not satisfy the first median condition to obtain a second filtering area; sorting the second filtering area in the diagonal direction to obtain a third filtering area, where in the third filtering area, the pixel values of the pixel points on the diagonal are in an increasing order in the diagonal direction; removing the pixel points in the third filtering area that do not satisfy the second median condition.
[0006] For example, in the image processing method provided by at least one embodiment of the present disclosure, the pixel points whose coordinates satisfy the following relationship are determined not to satisfy the first median condition:
[0007] x×y - 1 ≥ (c×d) / 2
[0008] (c - x + 1)×(d - y + 1) - 1 ≥ (c×d) / 2
[0009] where x and y represent the coordinates of the pixel point, c and d represent the size of the filtering area, and x, y, c, and d are positive integers.
[0010] For example, in the image processing method provided by at least one embodiment of the present disclosure, the pixel points that satisfy the following relationship are determined not to satisfy the second median condition:
[0011] Plt > (R + 1) / 2 or Prt > (R + 1) / 2
[0012] where R represents the number of pixel points in the second filtering area, Plt represents the number of pixel points in the third filtering area whose pixel values are greater than or equal to the pixel value of the pixel point and are located on the first side of the pixel point, Prt represents the number of pixel points in the third filtering area whose pixel values are less than or equal to the pixel value of the pixel point and are located on the second side of the pixel point, and along the direction from the first side to the second side, the pixel values of the pixel points increase.
[0013] For example, in the image processing method provided by at least one embodiment of the present disclosure, in response to n being greater than 1, the n pixel points include a×b pixel points arranged in an array in the to-be-processed image, where a and b are positive integers and at least one of a and b is greater than 1. Determining the filtering region corresponding to the n pixel points includes: determining the filtering window regions respectively corresponding to the n pixel points, where the filtering window region corresponding to each pixel point includes a region centered on the pixel point and defined by the filtering window of the median filtering operation; determining, according to the filtering window regions respectively corresponding to the n pixel points, the overlapping part of the filtering window regions respectively corresponding to the n pixel points as the filtering region corresponding to the n pixel points, where the size of the filtering region corresponding to the n pixel points is c×d, and the size of the filtering window is h×w, and c, d, h, and w are positive integers, and c is less than h, and d is less than w.
[0014] For example, in the image processing method provided by at least one embodiment of the present disclosure, determining the median filtering result of each pixel point in the n pixel points based on the remaining pixel points includes: for any one of the n pixel points: performing at least one first operation on the pixel points in the filtering window region corresponding to the any one pixel point that do not belong to the filtering region and the remaining pixel points to obtain the median point in the filtering window region corresponding to the any one pixel point as the median filtering result of the any one pixel point. Wherein, in each first operation, a first queue and a second queue are obtained, and at least one merge sorting operation is performed on the first queue and the second queue to obtain an output queue, where the number of elements in the output queue is r + 1, and r is the number of pixel points in the filtering window region corresponding to the any one pixel point that have not undergone the first operation, and the output queue is the r + 1 elements with the middle values in the first queue and the second queue.
[0015] For example, in the image processing method provided by at least one embodiment of the present disclosure, the first queue is the remaining pixel points or the output queue obtained from the previous first operation, and the second queue is c pixel points that have not undergone the first operation and are adjacent to the filtering region in the row direction or d pixel points that are adjacent to the filtering region in the column direction, where the c pixel points belong to the same column and the d pixel points belong to the same row.
[0016] For example, in the image processing method provided by at least one embodiment of the present disclosure, when r = 1, sorting the 2 elements in the output queue obtained from the current first operation and the 1 element in the filtering window region corresponding to the any one pixel point that has not undergone the first operation, and taking the median point of the 3 sorted elements as the median filtering result of the any one pixel point.
[0017] For example, in the image processing method provided by at least one embodiment of the present disclosure, the merge sorting operation includes: obtaining two input pixel queues, where each input pixel queue is ordered; performing merge sorting on the two input pixel queues to remove the smallest t elements and the largest t elements in the two input pixel queues, where t is a positive integer greater than 1; using the remaining elements in the two input pixel queues as a merged pixel queue, where the merged pixel queue is at least partially ordered.
[0018] For example, in the image processing method provided by at least one embodiment of the present disclosure, in response to n being equal to 1, determining the filtering region corresponding to the n pixel points includes: based on the filtering window of the median filtering operation, defining a region centered on the one pixel point in the image to be processed as the filtering region.
[0019] For example, in the image processing method provided by at least one embodiment of the present disclosure, determining the median filtering result of each pixel point among the n pixel points based on the remaining pixel points includes: sorting the remaining pixel points and determining the median point after sorting as the median filtering result of the one pixel point.
[0020] For example, in the image processing method provided by at least one embodiment of the present disclosure, different thread hardwares in the same hardware execution unit perform the median filtering operation on different pixel points in the same execution cycle.
[0021] For example, in the image processing method provided by at least one embodiment of the present disclosure, each hardware execution unit includes a vector operation unit buffer for data caching of the thread hardwares in the hardware execution unit. Obtaining M images to be processed includes: evenly dividing the input image into the M sub-images in the column direction; prefetching the M sub-images into the vector operation unit buffers of the M hardware execution units respectively; loading the pixel points in the edge regions obtained by performing border expansion on the M sub-images into the vector operation unit buffers of the corresponding hardware execution units respectively, so that when performing the median filtering operation in each hardware execution unit, the corresponding image to be processed can be directly obtained from the vector operation unit buffer of the hardware execution unit.
[0022] For example, in the image processing method provided by at least one embodiment of the present disclosure, loading the pixel points in the edge regions obtained by respectively performing the border expansion on the M sub-block images into the vector operation unit buffer of the corresponding hardware execution unit includes: for each sub-block image: expanding the upper side and the lower side of the sub-block image by p rows of pixel points in a direction away from the sub-block image, and expanding the left side and the right side of the sub-block image by q columns of pixel points in a direction away from the sub-block image, to obtain the image to be processed and the edge region of the sub-block image, where the edge region is the region in the image to be processed that does not belong to the sub-block image, and where p and q are positive integers; determining the coordinates of multiple pixel points in the edge region that belong to the upper-side adjacent sub-block image or the lower-side adjacent sub-block image of the sub-block image, and loading the multiple pixel points into the vector operation unit buffer of the corresponding hardware execution unit; obtaining the pixel values of the remaining pixel points in the edge region that do not belong to the upper-side adjacent sub-block image and the lower-side adjacent sub-block image according to a preset filling rule, and loading the remaining pixel points into the vector operation unit buffer of the corresponding hardware execution unit.
[0023] For example, in the image processing method provided by at least one embodiment of the present disclosure, splicing the median filtering results obtained by performing the median filtering operation on the L thread hardwares according to the positional relationship of the pixel points to obtain the filtering result of each hardware execution unit for the input image to be processed for the median filtering operation; splicing the filtering results respectively generated by the M hardware execution units according to the relative positional relationship of the M images to be processed in the row direction to obtain the median filtering result of the input image.
[0024] At least one embodiment of the present disclosure provides an image processing apparatus, including: an acquisition module configured to acquire M images to be processed, where the M images to be processed are obtained by respectively performing border expansion on M sub-block images, the M sub-block images are obtained by evenly dividing an input image into M blocks in the column direction, and M is a positive integer; M hardware execution units configured to respectively perform median filtering operations on the M images to be processed to obtain the median filtering result of the input image, where each hardware execution unit includes L thread hardwares, the L thread hardwares are configured to perform the median filtering operation on the image to be processed input to the hardware execution unit in parallel, and each thread hardware is configured to perform the median filtering operation on n pixel points at a time, where n is a positive integer, and where when the thread hardware performs the median filtering operation on the n pixel points, it includes performing the following operations: determining the filtering region corresponding to the n pixel points; sorting the pixel points in the filtering region in the row direction, column direction, and diagonal direction, and removing the pixel points that cannot be the median in the sorted filtering region; determining the median filtering result of each pixel point among the n pixel points based on the remaining pixel points.
[0025] For example, in the image processing apparatus provided by at least one embodiment of the present disclosure, the hardware execution unit includes a plurality of computing units, each computing unit includes a plurality of execution units, and each execution unit includes a plurality of thread hardwares.
[0026] At least one embodiment of the present disclosure provides an electronic device, including: a memory that stores computer-executable instructions non-transiently; a processor configured to run the computer-executable instructions, wherein when the computer-executable instructions are run by the processor, an image processing method according to any embodiment of the present disclosure is implemented.
[0027] At least one embodiment of the present disclosure provides a non-transient computer-readable storage medium, wherein the non-transient computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, an image processing method according to any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0029] Figure 1 It is a schematic structural diagram of a general-purpose graphics processing unit (GPGPU);
[0030] Figure 2 It is a schematic flowchart of an image processing method provided by at least one embodiment of the present disclosure;
[0031] Figure 3 It is a schematic diagram of the relationship between an image to be processed and a segmented image provided by an embodiment of the present disclosure;
[0032] Figure 4 It is a schematic diagram of the processing order of pixel points in an image to be processed provided by an embodiment of the present disclosure;
[0033] Figure 5 It is a schematic flowchart of a median filtering operation provided by at least one embodiment of the present disclosure;
[0034] Figure 6 It is a schematic diagram of the process of step S202 provided by an embodiment of the present disclosure;
[0035] Figure 7 It is a schematic diagram of a filtering region provided by at least one embodiment of the present disclosure;
[0036] Figure 8 It is a schematic block diagram of an image processing apparatus provided by at least one embodiment of the present disclosure;
[0037] Figure 9 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure;
[0038] Figure 10 Schematic block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0039] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0040] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of the present disclosure clear and concise, the detailed descriptions of some known functions and known components are omitted in the present disclosure.
[0041] Figure 1 Schematic diagram of a general-purpose graphics processor.
[0042] As Figure 1 shown, a general-purpose graphics processor is actually an array of programmable multi-processors. For example, the programmable multi-processors may be Streaming Processor Clusters (SPCs), for example, including Figure 1 the Streaming Processor Cluster 1 shown in... the Streaming Processor Cluster M, where M is a positive integer greater than 1. In a general-purpose graphics processor, 1 Streaming Processor Cluster processes one computing task, or multiple Streaming Processor Clusters process one computing task. Data sharing is performed among multiple Streaming Processor Clusters through a global cache or global memory.
[0043] As shown Figure 1 in the figure, taking the streaming processor cluster 1 as an example, one streaming processor cluster includes multiple computing units, such as Figure 1 the computing unit 1, computing unit 2,..., computing unit N in it, where N is a positive integer. Each computing unit (Compute Unit, abbreviated as CU) is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A computing unit includes multiple cores (also called computing cores or computing kernels), and each computing core includes an arithmetic logic unit (ALU), a floating-point computing unit, etc., and the computing core is used to perform specific computing tasks. In addition, the computing unit also includes registers (such as Figure 1 the register file in it) and shared memory, which are used to hierarchically store the source data and destination data related to the computing task. The shared memory in one computing unit is used to share data among the cores of this computing unit. As shown Figure 1 in the figure, each streaming processor cluster also includes a vector operation unit buffer (GEMM-Main-Buffer, abbreviated as GMB), which is shared by the computing units in the same streaming processor cluster and is used to share data among the computing units belonging to one streaming processor cluster.
[0044] For example, the register file can be an extended thread-local register (Thread-Local-Register, abbreviated as TLR), which is the closest to the computing unit and has a small capacity; the shared memory (Global-Shared-Memory, abbreviated as GSM) is relatively close to the computing unit, but farther from the computing unit than TLR and has a small capacity; the vector operation unit buffer (GEMM-Main-Buffer, abbreviated as GMB) is closer to the computing unit than the global cache (such as the L2 cache), but farther from the computing unit than the shared memory and has a large capacity.
[0045] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before being executed in a general-purpose graphics processor (or called a parallel computing processor), and then the multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in the figure). All threads in one thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundle (or simply called a thread bundle, warp), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.
[0046] In each computing unit, a warp scheduling / distribution module ( Figure 1 not shown in the figure) schedules and allocates warps so that multiple computing cores in the computing unit can execute warps. According to the number of computing cores in the computing unit, multiple warps in a thread block can be executed simultaneously or time-shared. Multiple threads in each warp will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the intermediate-level cache (such as the vector operation unit buffer) or the global cache or the global memory for read / write operations, etc.
[0047] Currently, there are two mainstream processing methods for median filtering of images. One is the conventional method of sorting to obtain the median, and the other is the method of using a histogram to obtain the median for images with a larger resolution.
[0048] The principle of the sorting-based median method is to first select a neighborhood centered on the current pixel, then sort all the pixel values in the neighborhood, and finally select the median value after sorting as the filtering value of the current pixel.
[0049] The histogram-based median method regards the entire picture as a sliding window. During the moving process (for example, from left to right), there is a shared part in the middle part between two adjacent filterings, and the newly added part is the part on the right side of the neighborhood. Therefore, a histogram can be used to update the pixel values.
[0050] For the first sorting-based median filtering method, the complexity of median filtering depends on the complexity of the sorting algorithm. Moreover, when the filtering window is large, a lot of repeated calculations are introduced, resulting in serious algorithm time consumption and poor performance.
[0051] For the second histogram-based median filtering method, in the horizontal direction, the filtering result of each pixel point needs to wait until the filtering calculation of the adjacent previous (for example, left) pixel point is completed to be obtained. That is, there is a data dependency in the row direction of the filtering process, and it is not suitable for parallel calculation on multi-threaded hardware computing resources (such as a graphics processing unit or a general-purpose graphics processing unit).
[0052] At least one embodiment of the present disclosure provides an image processing method, a processor, an electronic device, and a non-transitory computer-readable storage medium. The image processing method includes: obtaining M images to be processed, where the M images to be processed are obtained by respectively performing border expansion on M block images, the M block images are obtained by equally dividing an input image into M blocks in the column direction, and M is a positive integer; respectively performing median filtering operations on the M images to be processed by M hardware execution units to obtain a median filtering result of the input image, where each hardware execution unit includes L thread hardwares, and the L thread hardwares perform the median filtering operations on the images to be processed of the input hardware execution unit in parallel, and each thread hardware performs the median filtering operation on n pixel points once, and n is a positive integer. Wherein, performing the median filtering operation on n pixel points by the thread hardware includes: determining a filtering region corresponding to the n pixel points; sorting the pixel points in the filtering region in the row direction, column direction, and diagonal direction, and removing the pixel points that cannot be the median in the sorted filtering region; determining the median filtering result of each pixel point among the n pixel points based on the remaining pixel points.
[0053] In at least one embodiment, the image processing method can utilize multiple thread hardwares to perform the median filtering operation on pixel points in parallel, thereby fully utilizing the computing resources of the multi-thread hardwares, accelerating the processing speed of the median filtering operator in the image algorithm, and improving the execution efficiency of the image algorithm; in addition, the image processing method combines the characteristics of the multi-thread hardware computing resources and provides a median filtering operation (or median filtering operator) to accelerate the sorting process.
[0054] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.
[0055] Figure 2 It is a schematic flowchart of the image processing method provided by at least one embodiment of the present disclosure.
[0056] As Figure 2 shown, the image processing method provided by at least one embodiment of the present disclosure includes steps S10 - S20.
[0057] In step S10, obtain M images to be processed.
[0058] For example, the M images to be processed are obtained by respectively performing border expansion on M block images, and the M block images are obtained by dividing the input image into M blocks in the column direction (for example, the vertical direction), and M is a positive integer.
[0059] For example, the number of rows of the M block images may be the same or different.
[0060] For example, the input image can be evenly divided into M blocks in the column direction to obtain M block images. For example, each block image can include H / M rows and W columns, where H is the number of rows of the input image, W is the number of columns of the input image, and H and W are positive integers. For example, if H is not an integer multiple of M, the number of rows of some block images can be set to be greater than H / M so that all pixel points in the input image can be covered by the M block images. Thus, the hardware execution units can make full use of their respective computing powers, balance the median filtering calculation amount, and obtain the median filtering result of the final input image at the fastest speed.
[0061] For example, referring to Figure 2 the general graphics processor shown, each hardware execution unit includes a vector operation unit buffer for data caching of the thread hardware in the hardware execution unit. For example, N computing units belonging to one hardware execution unit can share the vector operation unit buffer, and the vector operation unit buffer has a caching mechanism different from that of a conventional L1 cache or L2 cache and can provide a data storage function.
[0062] For example, in this embodiment, step S10 may include: evenly dividing the input image into M block images in the column direction; prefetching the M block images into the vector operation unit buffers of the M hardware execution units respectively; loading the pixel points in the edge regions obtained by respectively expanding the borders of the M block images into the vector operation unit buffers of the corresponding hardware execution units, so that when the median filtering operation is performed in each hardware execution unit, the corresponding image to be processed can be directly obtained from the vector operation unit buffer of the hardware execution unit.
[0063] In this embodiment, the border expansion process does not require creating a new image data by allocating memory space. Instead, the "in-place mapping" method is used for border expansion. Specifically, the coordinate positions of the pixel points belonging to the edge region are calculated through logical judgment and recalculation, and then directly read. Thus, the memory overhead and the time consumption caused by allocating memory can be saved.
[0064] Figure 3 It is a schematic diagram showing the relationship between the image to be processed and the block images provided by an embodiment of the present disclosure.
[0065] As Figure 3 shown, the gray rectangular area is the block image, Figure 3 and three adjacent block images are shown in
[0066] Taking the block image 2 as an example, the dashed box around the block image 2 is the range of the image to be processed obtained by expanding the border of the block image 2. For example, the area defined by the dashed box is the image to be processed. The image to be processed includes the middle block image 2. The part outside the block image 2 is called the edge area. The edge area is the area in the image to be processed that does not belong to the block image 2, that is, Figure 3 the black stripe area outside the block image 2 in Figure 3 .
[0067] For example, expand the upper and lower sides of the block image 2 by p rows of pixel points in the direction away from the block image 2. For example, p is related to the size of the filtering window of the median filtering operation. For example, p = h / 2; expand the left and right sides of the block image 2 by q columns of pixel points in the direction away from the block image 2. For example, q is related to the size of the filtering window of the median filtering operation. For example, q = w / 2. Here, q and p are positive integers, and h and w are the height and width of the filtering window. Thus, the image to be processed and the edge area of the block image 2 are obtained.
[0068] After that, calculate the coordinates of multiple pixel points in the edge area that belong to the upper-side adjacent block image (i.e., block image 1) or the lower-side adjacent block image (i.e., block image 3) of the block image 2, and load these pixel points from the block image 1 and block image 3 stored in the memory to the vector operation unit buffer of the corresponding hardware execution unit.
[0069] After that, obtain the pixel values of the remaining pixel points in the edge area that do not belong to the block image 1 and block image 3 according to the preset filling rule, and load them to the vector operation unit buffer of the corresponding hardware execution unit. For example, the preset filling rule can adopt border replicate, border constant, border reflect, border reflect101, etc. The present disclosure does not make specific limitations on this.
[0070] In the above embodiment, the stored in the memory are the respective block images. The image to be processed after border expansion is obtained by calculating the mapping coordinates and stored in the vector operation unit buffer. Thus, it is not necessary to store each complete image to be processed, saving the memory overhead and the time-consuming caused by allocating memory.
[0071] For example, in some other embodiments, traditional border expansion processing can also be adopted, that is, store the complete data of the image to be processed obtained by border expansion in the memory. The specific border expansion process of the image to be processed can refer to the relevant part of the above embodiment; after that, directly load it from the memory to the vector operation unit buffer or register when needed. The present disclosure does not make specific limitations on this.
[0072] In the present disclosure, by using the in-situ mapping method as described above or directly loading the image to be processed from the memory, etc., the required data (image to be processed) is pre-loaded into the vector operation unit buffer before performing the median filtering operation, so that the data can be directly obtained from the vector operation unit buffer when performing the median filtering operation, without having to load data from the memory again. The vector operation unit buffer is closer to the calculation core than the memory, thereby accelerating the data reading process and reducing the memory access cost.
[0073] For example, in some other embodiments, if the hardware execution unit does not provide a vector operation unit buffer, after obtaining the region range of the image to be processed, the block image and the above-mentioned multiple pixel points can be read into the register in the conventional way of reading from the memory, and the pixel points after boundary filling according to the filling rule can also be read into the register, which will not be described in detail here.
[0074] In step S20, M hardware execution units respectively perform median filtering operations on M images to be processed to obtain the median filtering result of the input image.
[0075] For example, each hardware execution unit performs the median filtering operation on 1 image to be processed. Thus, M hardware execution units can be used to perform the median filtering operations on M images to be processed in parallel, making full use of the multi-threaded hardware computing resources to perform the median filtering operation in parallel and improving the processing efficiency.
[0076] For example, the image processing method can be applied to multi-threaded hardware computing resources, such as graphics processors, general-purpose graphics processors, etc. For example, the multi-threaded hardware computing resources provide multiple thread hardwares, which can be used to execute multiple instructions or computing tasks in parallel. For example, taking a general-purpose image processor as an example, a thread is the basic unit for executing a computing task, and a thread hardware refers to the hardware structure that supports the execution of these threads. Referring to the foregoing Figure 1 related content, for example, the hardware execution unit can be a programmable multi-processor, such as a streaming processor cluster. The hardware execution unit includes multiple computing units, each computing unit includes multiple execution units, each execution unit includes multiple thread hardwares, and one thread runs in one thread hardware. For example, one thread hardware performs the median filtering operation on n pixel points at a time, where n is a positive integer. Here, "at a time" means the computing task performed by the thread hardware at a single time, that is, after completing the median filtering operation on n pixel points, the data is re-loaded for the next computing task.
[0077] For example, different thread hardwares in the same hardware execution unit perform the median filtering operations on different pixel points at the same time.
[0078] For example, in a specific example, a general image processor includes M hardware execution units. Each hardware execution unit includes 4 computing units, each computing unit includes 4 execution units, and each execution unit includes 32 thread hardwares. Thus, one hardware execution unit can execute 512 threads simultaneously and can perform median filtering operations on 512×n pixel points simultaneously, thereby greatly improving the median filtering processing efficiency.
[0079] Figure 4 Schematic diagram of the processing order of pixel points in the image to be processed provided by an embodiment of the present disclosure.
[0080] As Figure 4 shown, assume that the hardware execution unit for processing the image to be processed includes T thread hardwares, and each thread hardware can perform median filtering operations on n pixel points at a time. Therefore, in one execution cycle (executing one computing task), median filtering results of T×n pixel points can be obtained. For example, first, median filtering operations are performed on T×n pixel points in the first rectangular frame of the first row in Figure 4 , and then, in the next execution cycle, median filtering operations are performed on T×n pixel points in the second rectangular frame of the first row rectangular frame in Figure 4 , and so on.
[0081] In Figure 4 , T1 is a positive integer. Assume that n pixel points include a rows and b columns of pixel points, that is, a×b pixel points arranged in an array in the row and column directions. If the number of pixel points in the first a rows of the image to be processed that have not undergone median filtering operations is less than T×b, then continue to select the first (T - T1)×b pixel points from the (a + 1)-th row to the 2×a-th row for median filtering operations.
[0082] For example, when obtaining the median filtering result of the input image, the median filtering results obtained by performing median filtering operations on L thread hardwares are stitched according to the positional relationship of the pixel points to obtain the filtering result of each hardware execution unit for performing median filtering operations on the input image to be processed; the filtering results generated by each of the M hardware execution units are stitched according to the relative positional relationship of the M images to be processed in the row direction to obtain the median filtering result of the input image.
[0083] The following specifically describes the specific process of the thread hardware performing median filtering operations on n pixel points at a single time in conjunction with the accompanying drawings.
[0084] Figure 5 Schematic flowchart of the median filtering operation provided by at least one embodiment of the present disclosure. For example, this median filtering operation is used to obtain the median filtering result of n pixel points.
[0085] As Figure 5As shown, the median filtering operation includes steps S201 - S203.
[0086] For example, in step S201, a filtering region corresponding to n pixel points is determined.
[0087] For example, in step S202, the pixel points in the filtering region are sorted in the row direction, column direction, and diagonal direction, and the pixel points that cannot be the median in the sorted filtering region are removed.
[0088] For example, in step S203, the median filtering result of each of the n pixel points is determined based on the remaining pixel points.
[0089] For example, in some embodiments, if n = 1, step S201 may include: based on the filtering window of the median filtering operation, a filtering region centered on this one pixel point is defined in the image to be processed. That is, when n = 1, the filtering region is the region defined by the filtering window in the image to be processed with the center of the filtering window overlapping the center of the pixel point.
[0090] For example, step S202 may include: sorting the pixel points in the filtering region in the row direction and column direction to obtain a first filtering region, where in the first filtering region, the pixel values of each row of pixel points are in increasing order in the row direction, and the pixel values of each column of pixel points are in increasing order in the column direction; removing the pixel points in the first filtering region whose coordinates do not satisfy the first median condition to obtain a second filtering region; sorting the second filtering region in the diagonal direction to obtain a third filtering region, where in the third filtering region, the pixel values of the pixel points on the diagonal are in increasing order in the diagonal direction; removing the pixel points in the third filtering region that do not satisfy the second median condition.
[0091] For example, pixel points whose coordinates satisfy the following relationship are determined not to satisfy the first median condition:
[0092] x×y - 1≥(c×d) / 2
[0093] (c - x + 1)×(d - y + 1) - 1≥(c×d) / 2
[0094] Where x and y represent the coordinates of the pixel point, c and d represent the size of the filtering region, and x, y, c, d are positive integers.
[0095] For example, pixel points that satisfy the following relationship are determined not to satisfy the second median condition:
[0096] Plt>(R + 1) / 2 or Prt>(R + 1) / 2
[0097] Wherein, R represents the number of pixel points in the second filtering region, Plt represents the number of pixel points in the third filtering region whose pixel values are greater than the pixel value of the pixel point and are located on the first side of the pixel point, Prt represents the number of pixel points in the third filtering region whose pixel values are smaller than the pixel value of the pixel point and are located on the second side of the pixel point, and along the direction from the first side to the second side, the pixel value of the pixel point increases.
[0098] Figure 6 It is a schematic diagram of the process of step S202 provided by an embodiment of the present disclosure.
[0099] As Figure 6 shown, the size of the filtering region is 7×7, which is determined by a 7×7 filtering window.
[0100] As Figure 6 shown, first, the filtering region is sorted in the row direction so that the pixel values of the pixel points in each row are in an increasing order in the row direction. For example, as Figure 6 shown, in the filtering region sorted in the row direction, the pixel points in each row increase in order from left to right.
[0101] After that, the filtering region is sorted in the column direction so that the pixel values of the pixel points in each column are in an increasing order in the column direction. For example, as Figure 6 shown, in the first filtering region sorted in the column direction, the pixel points in each column increase in order from top to bottom.
[0102] Of course, Figure 6 shown is an example, and it is also possible to sort in the column direction first and then in the row direction, and the sorting order is not specifically limited. In addition, Figure 6 taking the pixel point (149) in the upper left corner of the filtering region as the origin, starting from the pixel point in the upper left corner, the left-to-right direction is the row direction, and the top-to-bottom direction is the column direction for description. When the origin position is modified (for example, taking the pixel point (75) in the lower left corner as the coordinate origin), the sorting process and the subsequent median condition judgment process can also be adjusted accordingly, which will not be elaborated here.
[0103] After that, the pixel points in the first filtering region whose coordinates do not meet the first median condition are removed, and the pixel points that are most likely to be median points are retained.
[0104] For example, taking the pixel point (3) in the upper left corner of the first filtering region as the coordinate origin, the pixel points whose coordinates satisfy the following relationship are determined not to meet the first median condition:
[0105] x×y - 1 ≥ (c×d) / 2
[0106] (c - x + 1)×(d - y + 1) - 1 ≥ (c×d) / 2
[0107] Taking Figure 6Taking the first filtering region shown as an example, at this time c = 7, d = 7, (c × d) / 2 = 25 (round up when the result of c × d is odd). For example, when y = 1, the pixel points where x - 1 < 25 and (7 - x + 1) × (7 - 1 + 1) - 1 < 25 need to be satisfied simultaneously, that is, the pixel points where x > 4 are retained. For example, when y = 3, 3x - 1 < 25 and (7 - x + 1) × 5 - 1 < 25 need to be satisfied, that is, the pixel points where x >= 3 are retained, and so on.
[0108] Thus, remove Figure 6 the pixel points with a dark background in the second filtering region. At this time, there are 29 remaining pixel points in the second filtering region that may be median pixel points, that is, Figure 6 the pixel points with a white background in the second filtering region.
[0109] After that, perform a diagonal sorting on the second filtering region so that the pixel values of the pixel points in the second filtering region are in an increasing order in the diagonal direction. For example, as Figure 6 shown, in the third filtering region after diagonal sorting, the pixel points on the diagonal increase sequentially in the order from the upper left to the upper right.
[0110] Of course, it should be noted here that the first median condition judgment process can also be executed after diagonal sorting, that is, after sorting the filtering region in the row direction, column direction, and diagonal direction, remove the pixel points whose coordinates do not meet the first median condition in the sorted filtering region.
[0111] After that, remove the pixel points in the third filtering region that do not meet the second median condition, and retain the pixel points that are most likely to be median points.
[0112] For example, the pixel points that satisfy the following relationship are determined not to meet the second median condition:
[0113] Plt > (R + 1) / 2 or Prt > (R + 1) / 2
[0114] For example, for Figure 6 the example shown, R = 29. For any pixel point P(x, y) in the third filtering region to become a median point, it must not fall into the smallest 29 / 2 numbers, and at the same time, it must not fall into the largest 29 / 2 numbers. Also, since after diagonal sorting, the rows, columns, and diagonals are all in order at the same time, that is, for the current pixel point P, divide it into three parts to count the number of pixel points that must be greater than or equal to P:
[0115] 1. The pixel values of the pixel points on the diagonal to the right of point P must be greater than or equal to the pixel value of P, and the total number of pixel points that meet this condition is denoted as rt0
[0116] 2. For the pixel points on the diagonal line to the right of point P, the pixel values of the pixel points below them in the column direction must also be greater than or equal to the pixel value of P. The total number of pixel points that meet this condition is denoted as rt1.
[0117] 3. For the pixel points below point P in the column direction, the pixel values must be greater than or equal to the pixel value of P. The total number of pixel points that meet this condition is denoted as rt2.
[0118] For pixel point P, among the pixel points retained in the third filtering region, the number of pixel points with pixel values greater than or equal to that of P is counted and denoted as Prt. Here, Prt = rt0 + rt1 + rt2;
[0119] Similarly, among the retained data, the number of pixel points with pixel values less than or equal to the pixel value of P is counted and denoted as Plt.
[0120] Similar to the principle of counting the number of points greater than or equal to P, but with the comparison direction exactly opposite, it is also divided into three parts for counting:
[0121] 1. For the pixel points on the diagonal line to the left of point P, the pixel values must be less than or equal to the pixel value of P. The total number of pixel points that meet this condition is denoted as lt0
[0122] 2. For the pixel points on the diagonal line to the left of point P, the pixel values of the pixel points above them in the column direction must also be less than or equal to the pixel value of P. The total number of pixel points that meet this condition is denoted as lt1.
[0123] 3. For the pixel points above point P in the column direction, the pixel values must be less than or equal to the pixel value of P. The total number of pixel points that meet this condition is denoted as lt2.
[0124] For pixel point P, among the retained data, the number of pixel points with pixel values less than or equal to that of this point is counted and denoted as Plt = lt0 + lt1 + lt2.
[0125] Referring to the above process, for Figure 6 the example shown, finally 11 points are retained in the third filtering region, that is, Figure 6 the white background pixel points in the third filtering region in
[0126] Of course, the above median condition is an example. Those skilled in the art can also use other algorithms or sorting means to obtain the median condition to determine the pixel points that cannot be median points. The present disclosure does not make specific limitations on this.
[0127] After that, in step S203, based on the remaining pixel points, the median filtering results of each of the n pixel points are determined.
[0128] For example, in some embodiments, if n = 1, step S203 may include: sorting the remaining pixel points and determining the median point after sorting as the median filtering result of a pixel point.
[0129] Since the number of remaining pixel points is limited and has been significantly reduced compared to the number of pixel points in the filtering region, the median point can be determined through a finite number of comparison operations. For example, the median filtering result of this pixel point can be obtained through bubble sort or merge sort, etc.
[0130] For example, in some other embodiments, if n is greater than 1. For example, the n pixel points include a×b pixel points arranged in an array of a rows and b columns, that is, a×b pixel points arranged in an array in the row direction and the column direction, and at least one of a and b is greater than 1.
[0131] At this time, in some embodiments, a method of calculating the n pixel points independently of each other can be adopted. For example, referring to the median filtering process when n = 1, each pixel point is regarded as an independent pixel point for median filtering operation.
[0132] In some other embodiments, the overlapping parts of the filtering window regions of the n pixel points can be reused. For example, the n pixel points can reuse the sorting results of the overlapping regions, reducing the computational complexity in the process of determining the median point.
[0133] For example, if n is greater than 1, at this time step S201 may include: determining the filtering window regions respectively corresponding to the n pixel points, where the filtering window region corresponding to each pixel point includes a region defined by the filtering window of the median filtering operation with the pixel point as the center; according to the filtering window regions respectively corresponding to the n pixel points, determining the overlapping part of the n filtering window regions as the filtering region, where the size of the filtering region is c×d, the size of the filtering window is h×w, and c, d, h, and w are positive integers, and c is less than h, and d is less than w.
[0134] For example, the filtering window region corresponding to a pixel point is the region defined by the filtering window in the image to be processed with the center of the filtering window overlapping the center of the pixel point.
[0135] Figure 7 It is a schematic diagram of the filtering region provided by at least one embodiment of the present disclosure. For example, Figure 7 each small square represents a pixel point, Figure 7 shows a schematic diagram of the filtering region when the median filtering operation of 4 pixel points is executed once by the thread hardware. The 4 pixel points are respectively Figure 7 the pixel point p1, the pixel point p2, the pixel point p3, and the pixel point p4 in the black rectangular frame in
[0136] For example, assume that the size of the filtering window is 7×7, as Figure 7As shown, it shows 8×8 pixel points in the image to be processed, and is used to perform median filtering operations on pixel point p1, pixel point p2, pixel point p3, and pixel point p4.
[0137] For example, the filtering window area of pixel point p1 is as Figure 7 shown by the dashed box in, which is a 7×7 area centered on pixel point p1. The filtering window areas of pixel point p2, pixel point p3, and pixel point p4 are similar to that of pixel point p1, and are not shown again here.
[0138] For example, the filtering areas corresponding to the 4 pixel points are the overlapping parts of the 4 filtering window areas, that is, Figure 7 the 6×6 white-background rectangular boxes in the center here, with a total of 6×6 pixel points.
[0139] Figure 7 The following is an example. When the positions and quantities of n pixel points are different, the filtering areas also change accordingly. For example, if the n pixel points include 1×2 pixel points, such as two adjacent pixel points in the same column, the filtering area corresponding to the 2 pixel points at this time includes 7×6 pixel points.
[0140] After that, step S202 is executed to sort the pixel points in the filtering area corresponding to the n pixel points in the row direction, column direction, and diagonal direction, and remove the pixel points that cannot be the median in the sorted filtering area. The specific process can refer to the relevant description in step S202 above, and the repeated parts will not be elaborated here.
[0141] After that, step S203 is executed to determine the median filtering result of each pixel point among the n pixel points based on the remaining pixel points.
[0142] For example, in this embodiment, step S203 may include: for any one of the n pixel points: perform at least one first operation on the pixel points that do not belong to the filtering area and the remaining pixel points in the filtering window area corresponding to any one of the pixel points, and obtain the median point in the filtering window area corresponding to any one of the pixel points as the median filtering result of any one of the pixel points. Among them, in each first operation, a first queue and a second queue are obtained, and at least one merge sorting operation is performed on the first queue and the second queue to obtain an output queue. The number of elements in the output queue is r + 1, where r is the number of pixel points that have not undergone the first operation in the filtering window area corresponding to any one of the pixel points, and the output queue is the r + 1 elements in the middle of the values in the first queue and the second queue.
[0143] For example, the first queue is the remaining pixel points or the output queue obtained from the previous first operation, and the second queue is c pixel points adjacent to the filtering area in the row direction or d pixel points adjacent to the filtering area in the column direction that have not undergone the first operation. The c pixel points belong to the same column, and the d pixel points belong to the same row.
[0144] In at least one embodiment of the present disclosure, the process of finding the median point is based on the following theorem:
[0145] For a sequence with a length of n, r numbers are taken out from it (n - r > r), and then the r + 1 middlemost numbers are retained from the remaining n - r numbers. Then the median of the new sequence formed by these r + 1 numbers and the taken-out r numbers is equal to the median of the original sequence.
[0146] Assume that in step S202, after removing the pixel points that cannot be median points, the number of remaining pixel points in the third filtering area is P, and P is a positive integer. Based on this theorem, P should be greater than h × w - c × d.
[0147] After that, in step S203, the median filtering results of each of the n pixel points are calculated in sequence.
[0148] For example, for pixel point 1 among the n pixel points, these P pixel points are used as the first queue, and c pixel points adjacent to the filtering area in the row direction in the filtering window area of pixel point 1 are used as the second queue. At least one merge sort operation is performed on the first queue and the second queue to obtain an output queue. The output queue is the r + 1 elements with the middle pixel values among the total P + c pixel points of the first queue and the second queue. Here, r = w.
[0149] For example, the merge sort operation includes: obtaining two input pixel queues, where each input pixel queue is ordered; performing a merge sort on the two input pixel queues to remove the smallest t elements and the largest t elements in the two input pixel queues, where t is a positive integer greater than 1; and using the remaining elements in the two input pixel queues as a merged pixel queue, where the merged pixel queue remains locally ordered.
[0150] In the present disclosure, locally ordered means that the queue includes one or more sub-queues, and each sub-queue is ordered (for example, from smallest to largest or from largest to smallest). The input pixel queue should be ordered, and if it is locally ordered, it can be converted into an ordered queue through a finite number of comparisons. For example, the input pixel queue can be an ordered sub-queue in the merged pixel queue.
[0151] For example, perform a sorting operation on the first queue to make it ordered, thereby obtaining the first input pixel queue, and use the second queue as the second input pixel queue. Perform a merge sorting operation on the first input pixel queue and the second input pixel queue to fuse the two input pixel queues into one queue, and at the same time remove the largest t elements and the smallest t elements in the queue to obtain the output queue. For example, t is half of the total number of elements in the shorter queue among the first input pixel queue and the second input pixel queue, and if the total number of elements is odd, round down.
[0152] If the number of elements in the output queue is greater than r + 1, continue the above merge sorting process. For example, use a locally ordered sub-queue in the output queue as the first input pixel queue, and another sub-queue as the second input pixel queue for merge sorting. The specific process will not be elaborated here.
[0153] Through the above merge sorting process, among the total of P + c pixel points of the first queue and the second queue, the r + 1 pixel points with intermediate pixel values can be obtained, and these r + 1 pixel points are used as the r + 1 elements in the output queue.
[0154] After that, use these r + 1 elements as the first queue, and use the d pixel points adjacent to the filtering area in the column direction in the filtering window area of pixel point 1 as the second queue. Perform at least one merge sorting operation on the first queue and the second queue to obtain the output queue. The output queue is the r + 1 elements with intermediate pixel values among the total of w + 1 + d pixel points of the first queue and the second queue. Here, r = 1.
[0155] The specific process of merge sorting can refer to the above steps and will not be elaborated here.
[0156] For example, when r = 1, sort the 2 elements in the output queue obtained from the current first operation and the 1 element in the filtering window area corresponding to pixel point 1 that has not undergone the first operation. For example, this element is the element at the diagonal vertex, and use the median point of the 3 sorted elements as the median filtering result of any pixel point.
[0157] Of course, it is also possible to use all the remaining pixel points that have not undergone the first operation as the second queue during the second first operation and directly output the median filtering result, which will not be elaborated here.
[0158] Thus, the median filtering result of pixel point 1 is obtained through the above process. Similarly, for the remaining pixel points, similar processes can be used to obtain the median filtering results. The median filtering processes of each pixel point reuse the result of step S202, and only the pixel points in the non-filtered area need to be considered during subsequent sorting. This reduces the sorting calculation amount, has a low sorting implementation complexity, fully reuses the existing results, improves the median filtering efficiency, and especially significantly improves the performance for a relatively large filtering window (such as 5×5 or 7×7, etc.).
[0159] Next, taking the Figure 7 shown filtering area as an example, the execution process of step S203 will be described in detail.
[0160] As mentioned above, for the Figure 7 shown 4 pixel points, excluding the 6×6 = 36 pixel points in the filtering area corresponding to the 4 pixel points, there are still 7×7 - 36 = 13 points left. Therefore, only the 13 + 1 = 14 pixel points with the middle pixel values need to be found in the 6×6 area.
[0161] Referring to step S202 as described above, the pixel points that cannot be median points are removed through step S202, and finally 14 pixel points that may be median points are obtained.
[0162] After that, referring to step S203, the median filtering results of each of the 4 pixel points are determined based on these 14 pixel points.
[0163] Taking pixel point p1 as an example, referring to Figure 7 , first, for these 14 pixel points and the 6 pixel points that are not in the filtering area but on the left side of the filtering window area located at pixel point p1, that is, Figure 7 the 6 pixel points of the checkerboard background in the leftmost column of
[0164] a first operation is performed. Specifically, for the filtering window area of pixel point p1, there are still 7 pixel points in the first row that have not undergone the first operation. Therefore, according to the above theorem, the 8 middle pixel points among these 14 + 6 = 20 points need to be retained.
[0165] For example, taking these 14 pixel points as the first queue and these 6 pixel points as the second queue, at least one merge sorting operation is performed to obtain the output queue, and the output queue is the 8 middle pixel points among these 20 points.
[0166] Specifically, assume that the 6 pixel points are respectively represented as v0, v1, ..., v5, and the 14 pixel points are represented as u0, u1, ..., u13. In the first merge sorting operation, v0, v1, ..., v5 serve as the first input pixel queue, where v0, v1, ..., v5 increase in sequence, and u0, u1, ..., u13 serve as the second input pixel queue. The 14 pixel points are sorted to make them ordered, and u0, u1, ..., u13 increase in sequence.
[0167] First, compare the magnitudes of v0, v1, and v2 with u0, u1, and u2. Retain the larger 3 values among the 6 elements in u0, u1, and u2, and retain the smaller 3 elements in v0, v1, and v2 and remove them.
[0168] Similarly, compare the magnitudes of v3, v4, and v5 with u11, u12, and u13. Retain the smaller 3 values among the 6 elements in u11, u12, and u13, and retain the larger 3 elements in v3, v4, and v5 and remove them. Thus, the merged pixel queue u0, u1, ..., u13 is obtained, from which the largest 3 elements and the smallest 3 elements among the 20 elements are removed. This completes one merge sorting.
[0169] The current merged pixel queue u0, u1, ..., u13 is locally ordered. The magnitude relationship between u2 and u3 is unknown, and the magnitude relationship between u10 and u11 is unknown. Therefore, u0, u1, u2 can be used as the first input pixel queue, and u3, u4, and u5 can be used as the second input pixel queue to perform the merge sorting as described above, and the larger 3 elements among the 6 elements are retained in u3, u4, and u5.
[0170] Similarly, use u8, u9, u10 as the first input pixel queue, and u11, u12, and u3 as the second input pixel queue to perform the merge sorting as described above, and the smaller 3 elements among the 6 elements are retained in u8, u9, u10.
[0171] Thus, the output queue u3, u4, u5, u6, u7, u8, u9, u10 is obtained.
[0172] After that, use the 8 elements in the above output queue as the first queue, and use the 6 pixel points that are located in the filtering window area of pixel point p1, do not belong to the filtering area but are adjacent to the upper part of the filtering area, that is Figure 7 the 6 pixel points of the diamond pattern background in the first row in
[0173] Specifically, for the filtering window area of pixel p1, there is still 1 remaining pixel point with a twill background in the upper left corner that has not undergone the first operation. Therefore, according to the above theorem, in this first operation, the 2 middle pixel points among the 8 + 6 = 14 points need to be retained.
[0174] The specific process is similar to the merge sorting process in the above first operation, and the repeated parts will not be elaborated here.
[0175] After that, these 2 pixel points and the finally remaining vertex pixel point in the upper left corner are sorted to obtain the median point, thereby obtaining the median filtering result of pixel p1.
[0176] For pixels p2, p3, and p4, the 14 pixel points obtained in step S202 are reused, and their respective median filtering results are obtained by referring to the above process, which will not be elaborated here.
[0177] In the image processing method provided by at least one embodiment of the present disclosure, a low-complexity median filtering operation that can be parallelly computed in multi-threaded hardware computing resources is provided. Especially, the performance is significantly improved on median operators with larger filtering windows (such as 5×5 / 7×7, etc.); moreover, the hardware characteristics are fully utilized, and the high-speed buffer provided by each hardware execution unit is used to further increase the parallel computing ability and reduce the memory access overhead.
[0178] At least one embodiment of the present disclosure further provides an image processing device. Figure 8 It is a schematic block diagram of an image processing device provided by at least one embodiment of the present disclosure.
[0179] As Figure 8 shown, the image processing device 100 includes an acquisition module 101 and M hardware execution units 102. Of course, the image processing device 100 may further include more hardware execution units, and the present disclosure does not make specific limitations on this.
[0180] For example, the acquisition module 101 is configured to acquire M images to be processed. The M images to be processed are obtained by respectively performing border expansion on M sub-block images, and the M sub-block images are obtained by evenly dividing the input image into M blocks in the column direction, where M is a positive integer.
[0181] For example, the M hardware execution units 102 are configured to respectively perform median filtering operations on the M images to be processed to obtain the median filtering result of the input image.
[0182] For example, each hardware execution unit includes L thread hardwares, and the L thread hardwares are configured to parallelly perform median filtering operations on the images to be processed of the input hardware execution unit. Each thread hardware is configured to perform a median filtering operation on n pixel points at a time, where n is a positive integer.
[0183] For example, the image processing device 100 is, for example, a graphics processing unit or a general-purpose graphics processing unit. Referring to the relevant description above Figure 1 the hardware execution unit may be a programmable multi-processor, such as a stream processor cluster.
[0184] For example, the hardware execution unit includes a plurality of computing units, each computing unit includes a plurality of execution units, and each execution unit includes a plurality of thread hardwares. Regarding the structure of the hardware execution unit, reference can be made to the relevant description above Figure 1 and will not be elaborated here.
[0185] Of course, the present disclosure is not limited thereto, and the graphics processing device 100 may also be other processing devices capable of providing a plurality of thread hardwares.
[0186] For example, when the thread hardware performs a median filtering operation on n pixel points, it includes performing the following operations: determining a filtering region corresponding to the n pixel points; sorting the pixel points in the filtering region in the row direction, column direction, and diagonal direction, and removing the pixel points that cannot be the median in the sorted filtering region; determining the median filtering result of each of the n pixel points based on the remaining pixel points.
[0187] For example, the image processing device 100 further includes a memory and a vector operation unit buffer ( Figure 8 not shown in the figure). For example, the input image is stored in the memory of the image processing device 100. Each hardware execution unit has a vector operation unit buffer for loading the to-be-processed image that the hardware execution unit needs to process from the memory to the vector operation unit buffer before performing the median filtering operation, so that when each hardware execution unit performs the median filtering operation, the corresponding to-be-processed image can be directly obtained from the vector operation unit buffer of the hardware execution unit without reading from the memory, reducing the memory access overhead.
[0188] It should be noted that the acquisition module 101 can be used to implement Figure 2 the step S10 shown in the figure; the hardware execution unit 102 can be used to implement Figure 2 the step S20 shown in the figure. Therefore, for the specific description of the functions that the acquisition module 101 can implement, reference can be made to the relevant description of step S10 in the embodiments of the above image processing method. For the specific description of the functions that the hardware execution unit 102 can implement, reference can be made to the relevant description of step S20 in the embodiments of the above image processing method. The repeated parts will not be elaborated. In addition, the image processing device 100 can achieve technical effects similar to those of the foregoing image processing method, which will not be elaborated here.
[0189] It should be noted that, in at least one embodiment of the present disclosure, the image processing apparatus 100 may include more or fewer circuits or units, and the connection relationship between the respective circuits or units is not limited and may be determined according to actual needs. The specific constitution manner of each circuit or unit is not limited and may be constituted by analog devices according to circuit principles, or may be constituted by digital chips, or in other applicable manners.
[0190] For example, the image processing apparatus 100 may be implemented in a manner combining hardware, software, or both hardware and software, and the present disclosure does not make specific limitations thereto.
[0191] Figure 9 Schematic diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure. For example, as Figure 9 shown, the storage medium 200 may be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 210 may be non-temporarily stored on the storage medium 200. For example, when the computer-readable instructions 210 are executed by a processor, one or more steps of the image processing method described above may be executed.
[0192] For example, the storage medium 200 may be applied to an electronic device 300. For example, the storage medium 200 may include a storage device 308 in the electronic device 300.
[0193] For example, the storage device may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions may be stored on the computer-readable storage medium, and the processor may run the computer-readable instructions to implement various functions of the processor. Various application programs and various data may also be stored in the storage medium.
[0194] For example, the storage medium may include a memory card of a smart phone, a cache component of a tablet computer, a hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, and may also be other applicable storage media.
[0195] Figure 10Schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 10 shown, the electronic device 300 is, for example, suitable for implementing the image processing method provided by the embodiments of the present disclosure. It should be noted that Figure 10 the components of the electronic device 300 shown are exemplary and not restrictive. According to actual application needs, the electronic device 300 may also have other components.
[0196] As Figure 10 shown, the electronic device 300 may include a processing device (such as a graphics processor, a central processing unit, etc.) 301, which can perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to implement various functions.
[0197] For example, when the computer-readable instructions are run by the processing device 301, one or more steps of the image processing method described in any of the above embodiments may be executed. It should be noted that the detailed description of the processing process of the image processing method can refer to the relevant descriptions in the embodiments of the above image processing method.
[0198] For example, the central processing unit in the electronic device can be configured to schedule and control the graphics processor during image processing. For example, the central processing unit is responsible for processing complex logical operations and overall control, while the graphics processor focuses on graphics processing and data parallel computing.
[0199] For example, the memory may include any combination of one or more computer program products. The computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may, for example, include random access memory (RAM) 303 and / or cache memory, etc. For example, the computer-readable instructions can be loaded from the storage device 308 into the random access memory (RAM) 303 to run the computer-readable instructions. Non-volatile memory may, for example, include read-only memory (ROM) 302, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. Various application programs and various data can also be stored in the computer-readable storage medium, such as style images, and various data used and / or generated by the application programs, etc.
[0200] For example, the processing device 301, ROM 302, and RAM 303 are connected to each other through the bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0201] Typically, the following devices can be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, magnetic tape, a hard disk, a flash memory, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 10 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the electronic device 300 can alternatively implement or have more or fewer devices. For example, the processing device 301 can control other components in the electronic device 300 to perform desired functions. The processing device 301 can be a device with data processing capabilities and / or program execution capabilities such as a central processing unit (CPU), a tensor processing unit (TPU), or a graphics processing unit GPU. The central processing unit (CPU) can be of X86, ARM, RISC-V architecture, etc. The GPU can be directly integrated into the SOC, directly integrated onto the motherboard, or built into the northbridge chip of the motherboard.
[0202] For example, the electronic device 300 can be set on the server side (or cloud).
[0203] For example, in some embodiments, the electronic device 300 can be a mobile phone, a tablet computer, an electronic paper, a television, a monitor, a laptop computer, a digital photo frame, a navigator, a wearable electronic device, a smart home device, etc.
[0204] For example, the electronic device 300 can include a display panel, and the display panel can be used to divide images, etc. For example, the display panel can be a rectangular panel, a circular panel, an oval panel, or a polygonal panel, etc. In addition, the display panel can be not only a flat panel, but also a curved panel, or even a spherical panel.
[0205] For example, the electronic device 300 can have a touch function, that is, the electronic device 300 can be a touch device.
[0206] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0207] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. In this case, the name of the unit does not constitute a limitation on the unit itself.
[0208] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0209] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.
[0210] Furthermore, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0211] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
[0212] The following points also need to be noted regarding the present disclosure:
[0213] (1) The accompanying drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures may refer to the general design.
[0214] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other to obtain new embodiments.
[0215] The above are only the specific implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.
Claims
1. An image processing method, comprising: Acquire M images to be processed, wherein the M images to be processed are obtained by respectively expanding the edges of M block images, and the M block images are obtained by dividing the input image into M blocks in a column direction, and M is a positive integer; M hardware execution units respectively perform median filtering operations on the M images to be processed to obtain the median filtering results of the input image, wherein each hardware execution unit includes L thread hardware, and the L thread hardware performs the median filtering operation on the image to be processed input to the hardware execution unit in parallel, and each thread hardware independently performs the median filtering operation on n pixels at a time, and in each operation cycle, the L thread hardware performs the median filtering operation on L*n pixels in the image to be processed input to the hardware execution unit in parallel, and the L thread hardware performs the median filtering operation on different L*n pixels in different operation cycles, n and L are positive integers, L is greater than 1, The median filtering operation of the n pixels is performed by the thread hardware, including: Determine the filtering area corresponding to the n pixels; Sorting the pixels in the filtering area in the row direction, the column direction and the diagonal direction, and removing the pixels in the filtering area that cannot be the median after the sorting; Determine a median filtering result for each pixel in the n pixels based on the remaining pixels; Each hardware execution unit includes a vector operation unit cache area, and the vector operation unit cache area is used for data caching of each thread hardware in the hardware execution unit. Get M images to be processed, including: Dividing the input image into the M block images in the column direction; Pre-fetching the M block images stored in the memory into the vector operation unit buffers of the M hardware execution units respectively; The pixel points of the edge areas obtained by respectively performing the edge expansion on the M block images are loaded into the vector operation unit cache of the corresponding hardware execution unit, so that when each hardware execution unit performs the median filtering operation, the corresponding image to be processed is directly obtained from the vector operation unit cache of the hardware execution unit.
2. The image processing method according to claim 1, wherein: Sorting the pixels in the filtering area in the row direction, the column direction and the diagonal direction, and removing the pixels in the filtering area that cannot be the median after the sorting, including: Sorting the pixel points in the filtering area in the row direction and the column direction to obtain a first filtering area, wherein in the first filtering area, the pixel values of the pixel points in each row are in increasing order in the row direction, and the pixel values of the pixel points in each column are in increasing order in the column direction; Remove the pixel points whose coordinates do not satisfy the first median condition in the first filtering area to obtain a second filtering area; Sorting the second filtering area in the diagonal direction to obtain a third filtering area, wherein in the third filtering area, pixel values of pixels on the diagonal are in increasing order in the diagonal direction; Pixels in the third filtering area that do not satisfy the second median condition are removed.
3. The image processing method according to claim 2, wherein: The pixel points whose coordinates satisfy the following relationship are determined not to satisfy the first median condition: x×y-1≥(c×d) / 2, (c-x+1)×(d-y+1)-1≥(c×d) / 2; Wherein, x and y represent the coordinates of the pixel point, c and d represent the size of the filtering area, and x, y, c, and d are positive integers.
4. The image processing method according to claim 2, wherein: Pixel points that satisfy the following relationship are determined not to satisfy the second median condition: Plt>(R+1) / 2 or Prt>(R+1) / 2; Among them, R represents the number of pixel points in the second filtering area, Plt represents the number of pixel points in the third filtering area whose pixel values are greater than or equal to the pixel value of the pixel point and are located on the first side of the pixel point, and Prt represents the number of pixel points in the third filtering area whose pixel values are less than or equal to the pixel value of the pixel point and are located on the second side of the pixel point. Along the direction from the first side to the second side, the pixel values of each pixel point increase.
5. The image processing method according to claim 1, wherein: In response to n being greater than 1, the n pixels include a rows and b columns of pixels arranged in an array in the image to be processed, a and b are positive integers, and at least one of a and b is greater than 1, Determining the filtering area corresponding to the n pixel points includes: Determine the filter window areas corresponding to the n pixel points respectively, wherein the filter window area corresponding to each pixel point includes an area centered on the pixel point and defined by the filter window of the median filter operation; According to the filter window areas corresponding to the n pixel points, the overlapping parts of the filter window areas corresponding to the n pixel points are determined as the filter areas corresponding to the n pixel points, wherein the size of the filter area corresponding to the n pixel points is c×d, and the size of the filter window of the median filtering operation is h×w, c, d, h, and w are positive integers, c is less than h, and d is less than w.
6. The image processing method according to claim 5, wherein: Determining a median filtering result of each pixel in the n pixels based on the remaining pixels, comprising: For any pixel among the n pixels: Performing at least one first operation on the pixel points that do not belong to the filtering area in the filtering window area corresponding to any one pixel point and the remaining pixel points to obtain a median point in the filtering window area corresponding to any one pixel point as a median filtering result of the any one pixel point, In which, in each first operation, the first queue and the second queue are obtained, and at least one merge sort operation is performed on the first queue and the second queue to obtain an output queue, wherein the number of elements of the output queue is r+1, r is the number of pixels in the filter window area corresponding to any pixel point that has not undergone the first operation, and the output queue is the r+1 elements with values in the middle between the first queue and the second queue.
7. The image processing method according to claim 6, wherein: The first queue is the remaining pixel points or the output queue obtained by the last first operation, and the second queue is the c pixel points that have not undergone the first operation and are adjacent to the filter area in the row direction or the d pixel points adjacent to the column direction, the c pixel points belong to the same column, and the d pixel points belong to the same row.
8. The image processing method according to claim 6, wherein: When r=1, the two elements in the output queue obtained by the current first operation and the one element in the filtering window area corresponding to any pixel point that has not undergone the first operation are sorted, and the median point of the sorted three elements is used as the median filtering result of any pixel point.
9. The image processing method according to claim 6, wherein: The merge sort operation includes: Obtain two input pixel queues, wherein each input pixel queue is ordered; Merge and sort the two input pixel queues to remove the smallest t elements and the largest t elements in the two input pixel queues, where t is a positive integer greater than 1; The remaining elements in the two input pixel queues are used as a merged pixel queue, wherein the merged pixel queue is at least partially ordered.
10. The image processing method according to claim 1, wherein: In response to n being equal to 1, determining the filtering area corresponding to the n pixels includes: Based on the filtering window of the median filtering operation, an area centered on the one pixel point is defined in the image to be processed as the filtering area.
11. The image processing method according to claim 10, wherein: Determining a median filtering result of each pixel in the n pixels based on the remaining pixels, comprising: The remaining pixel points are sorted, and the sorted median point is determined as the median filtering result of the one pixel point.
12. The image processing method according to any one of claims 1 to 11, wherein: Different thread hardware in the same hardware execution unit executes the median filtering operation on different pixels in the same execution cycle.
13. The image processing method according to claim 1, wherein: The step of loading pixel points of edge areas obtained by respectively performing the edge expansion on the M block images into a vector operation unit buffer of a corresponding hardware execution unit includes: For each tile image: The upper side and the lower side of the block image are respectively extended by p rows of pixels in a direction away from the block image, and the left side and the right side of the block image are respectively extended by q columns of pixels in a direction away from the block image, so as to obtain an image to be processed and an edge region of the block image, wherein the edge region is a region of the image to be processed that does not belong to the block image, wherein p and q are positive integers; Determine coordinates of a plurality of pixel points in the edge region belonging to an upper side adjacent block image or a lower side adjacent block image of the block image, and load the plurality of pixel points from the upper side adjacent block image or the lower side adjacent block image stored in the memory to a vector operation unit buffer of the corresponding hardware execution unit; The pixel values of the remaining pixels in the edge area that do not belong to the upper side adjacent block image and the lower side adjacent block image are obtained according to a preset filling rule, and the remaining pixels are loaded into the vector operation unit buffer of the corresponding hardware execution unit.
14. The image processing method according to any one of claims 1 to 11, wherein: The median filtering results obtained by the median filtering operation performed by the L thread hardware are spliced according to the position relationship between the pixel points to obtain the filtering result of each hardware execution unit performing the median filtering operation on the input image to be processed; The filtering results generated by the M hardware execution units are respectively spliced according to the relative position relationship of the M images to be processed in the column direction to obtain a median filtering result of the input image.
15. An image processing device, comprising: An acquisition module is configured to acquire M images to be processed, wherein the M images to be processed are obtained by respectively expanding the edges of M block images, and the M block images are obtained by dividing the input image into M blocks in a column direction, and M is a positive integer; M hardware execution units are configured to perform median filtering operations on the M images to be processed respectively to obtain median filtering results of the input images, Each hardware execution unit includes L thread hardware, and the L thread hardware is configured to perform the median filtering operation on the image to be processed input to the hardware execution unit in parallel. Each thread hardware is configured to independently perform the median filtering operation on n pixels at a time. In each operation cycle, the L thread hardware performs the median filtering operation on L*n pixels in the image to be processed input to the hardware execution unit in parallel, and the L thread hardware performs the median filtering operation on different L*n pixels in different operation cycles. n and L are positive integers, and L is greater than 1. Wherein, when the thread hardware performs the median filtering operation of the n pixels, the following operations are performed: Determine the filtering area corresponding to the n pixels; Sorting the pixels in the filtering area in the row direction, the column direction and the diagonal direction, and removing the pixels in the filtering area that cannot be the median after the sorting; Determine a median filtering result for each pixel in the n pixels based on the remaining pixels; Each hardware execution unit includes a vector operation unit cache area, and the vector operation unit cache area is used for data caching of each thread hardware in the hardware execution unit. Get M images to be processed, including: Dividing the input image into the M block images in the column direction; Pre-fetching the M block images stored in the memory into the vector operation unit buffers of the M hardware execution units respectively; The pixel points of the edge areas obtained by respectively performing the edge expansion on the M block images are loaded into the vector operation unit cache of the corresponding hardware execution unit, so that when each hardware execution unit performs the median filtering operation, the corresponding image to be processed is directly obtained from the vector operation unit cache of the hardware execution unit.
16. The image processing apparatus according to claim 15, wherein: The hardware execution unit includes multiple computing units, each computing unit includes multiple execution units, and each execution unit includes multiple thread hardware.
17. An electronic device comprising: A memory non-transitorily stores computer executable instructions; a processor configured to execute the computer executable instructions, Wherein, when the computer executable instructions are executed by the processor, the image processing method according to any one of claims 1-14 is implemented.
18. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores computer-executable instructions, When the computer executable instructions are executed by a processor, the image processing method according to any one of claims 1 to 14 is implemented.
Citation Information
Patent Citations
Graphic processor-based weighted median filtering method and device
CN115797206A