Series Filter Structure Optimization Method, Device and Its Computer Equipment
By using the method of parallel processing of multiple filters in the first-level filter for digital image processing, the problem of line cache delay caused by series filters is solved, and the effect of eliminating line cache and reducing memory requirements is achieved.
Patent Information
- Application Number
- CN202311074516.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-08-24
AI Technical Summary
In digital image processing, the line cache delay problem caused by series filters, especially when image resolution is increased, the cost per line cache data represents a high cost.
By using multiple filters in the first-stage filter in parallel processing, k rows of data are output simultaneously, thereby exchanging the reduction of cache area in area and eliminating the line cache delay.
It achieves complete elimination of line cache, reduces memory requirements, reduces processing costs, and improves image processing efficiency.
Smart Images

Figure CN116862753B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of filtering structures, and particularly to an optimization method and device for a series filtering structure and a computer device thereof. Background Art
[0002] In digital image processing, it is usually necessary to use an n x m data window for a filtering algorithm, which will cause a delay of n / 2 rows in the output of the filter. When two filters are connected in series, the second-stage filter uses the output (k x j) of the first stage as the input, which will cause a delay of (n + k) / 2 rows in the final output relative to the original input. If the final filtering result needs to be subjected to a certain operation processing with the original input, we will need to use a memory space of (n + k) / 2 rows to store the original input data in order to wait for the output result of the second-stage filter. To reduce the memory requirement, a common method is to select the necessary data from the n x m data window of the first-stage filter for storage, which can avoid a delay of n / 2 rows and only needs to store k / 2 rows of data. Although this method can slightly reduce the cache, there is still a large cache requirement remaining. With the trend of continuous improvement of digital image resolution, each row of cached data represents a great cost. Summary of the Invention
[0003] The main object of the present invention is to provide an optimization method and device for a series filtering structure and a computer device thereof. In order to completely eliminate the row cache, a method of parallel processing of multiple filters in the first-stage filtering is adopted to output k rows of data simultaneously, achieving the effect of exchanging the area of k filters for the cache area of k / 2 rows.
[0004] To achieve the above object, the present invention provides an optimization method for a series filtering structure, including the following steps:
[0005] Obtain first image data;
[0006] Retrieve a first image data window from the first image data according to a preset retrieval interval, and obtain a number of second image data windows that have not been retrieved;
[0007] Perform beat delay processing on the first image data window, and simultaneously send a number of the second image data windows to a first filter, and arrange the number of second image data windows by the first filter one by one;
[0008] Input the number of second image data windows arranged by the first filter into a second filter for normalization operation to generate a third image data window;
[0009] Fuse the first image data window and the third image data window to generate second image data;
[0010] Output the second image data.
[0011] Further, in the step of retrieving a first image data window from the first image data according to a preset retrieval interval and obtaining a plurality of second image data windows that have not been retrieved, the retrieval interval includes an n×m interval of the central part of the first image data.
[0012] Further, the step of performing a beat delay process on the first image data window includes:
[0013] Perform a pause and wait process on the first image data window to wait for the third image data window generated after calculation for fusion processing.
[0014] Further, the step of inputting the plurality of second image data windows after the first filter is sorted into a second filter for normalization calculation to generate a third image data window includes:
[0015] Use the second filter to splice the pixels in rows and columns of the plurality of second image data windows to generate the third image data window.
[0016] Further, the step of performing a fusion process on the first image data window and the third image data window to generate second image data includes:
[0017] Perform a fusion process on the first image data window and the third image data window according to a preset weight ratio to generate second image data.
[0018] Further, the weight ratio includes taking 50% of the pixel rows and columns of the first image data window and 50% of the pixel rows and columns of the third image data window, which are combined to form the second image data.
[0019] Further, in the step of obtaining the first image data and outputting the second image data, it includes:
[0020] Obtain the first image data pixel by pixel through the data_in interface and output the output second image data through the data_out interface.
[0021] Further, the first image data window and the third image data window are represented row by row and column by column by RGBtoken strings. In the step of performing a fusion process on the first image data window and the third image data window, it includes:
[0022] Fuse the RGBtoken strings corresponding to the first image data window and the third image data window respectively using a color mixing unit, where the color mixing unit includes a color mixing module of ps.
[0023] The present invention also proposes a series filtering structure optimization device, including:
[0024] An acquisition unit for acquiring first image data;
[0025] An extraction unit for extracting a first image data window from the first image data according to a preset extraction range, and obtaining a plurality of second image data windows that have not been extracted;
[0026] A first filtering unit for performing beat delay processing on the first image data window, simultaneously sending the plurality of second image data windows to a first filter, and arranging the plurality of second image data windows one by one by the first filter;
[0027] A second filtering unit for inputting the plurality of second image data windows arranged by the first filter into a second filter for normalization operation to generate a third image data window;
[0028] A fusion unit for performing fusion processing on the first image data window and the third image data window to generate second image data;
[0029] An output unit for outputting the second image data.
[0030] The present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the series filtering structure optimization method described in any one of the above are implemented.
[0031] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the series filtering structure optimization method described in any one of the above are implemented.
[0032] The series filtering structure optimization method, device and computer device provided by the present invention have the following beneficial effects:
[0033] 1. Significantly shorten the model R & D cycle of a single project, thereby improving the project delivery efficiency;
[0034] 2. More projects can be carried out in parallel, significantly improving the R & D efficiency from another dimension;
[0035] 3. Significantly improve the response efficiency to requirement changes;
[0036] 4. The large model makes the model training standard efficient and easy to maintain, and also makes software R & D more standardized;
[0037] 5. The detection ability of the large model on new tasks is continuously enhanced, and the speed at which new projects reach the detection index based on fine-tuning training of the large model parameters is continuously accelerating. Description of the Drawings
[0038] Figure 1 It is a schematic diagram of a solution in the prior art;
[0039] Figure 2 It is a schematic diagram of the principle of the method for optimizing the series filtering structure in an embodiment of the present invention;
[0040] Figure 3 It is a schematic flowchart of the method for optimizing the series filtering structure in an embodiment of the present invention;
[0041] Figure 4 It is a block diagram of the structure of the device for optimizing the series filtering structure in an embodiment of the present invention;
[0042] Figure 5 It is a schematic block diagram of the structure of a computer device in an embodiment of the present invention.
[0043] The realization, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0044] In order to make the object, technical solution, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0045] Refer to Figure 1 It is a schematic diagram of a solution in the prior art. The data_in is input pixel by pixel. The row buffer does not change the original data, but it needs to accumulate n rows of data (usually starts to output after n / 2 data, and the missing rows are filled in by the mirror method), so there will be a row delay. The filter will change the data, but the input-output delay is relatively small, and it will usually output after several or more than a dozen clock cycles.
[0046] Refer to the attached Figure 2 , which is a schematic diagram of the principle of the method for optimizing the series filtering structure proposed by the present invention. By improving according to the prior art, after improvement, a buffer with n + k rows is used to replace the previous two row buffers. The total size remains the same, but because they are concentrated together, the area will be a little smaller. The output data window of (n + k) × m can extract k data windows of n × m respectively and send them to the corresponding A filters. The outputs of the k A filters form k rows of data and are sent to the filter B for operation. The required original data can be extracted from the (n + k) × m data window, and after being delayed by beating, it is aligned with data_B and then fused. The number of beats depends on the total delay clock cycle number processed by the two-stage filter. Because of parallel processing, the power consumption and total area of the filter A will increase, but it is much smaller than the area of the row buffer.
[0047] Specifically,
[0048] Refer to the attachedFigure 3 It is a schematic flowchart of an optimization method for a series filtering structure proposed by the present invention, including the following steps:
[0049] S1. Obtain first image data;
[0050] S2. Retrieve a first image data window from the first image data according to a preset retrieval interval, and obtain a number of second image data windows that have not been retrieved;
[0051] S3. Perform beat delay processing on the first image data window, and at the same time send the number of second image data windows to a first filter, and the first filter arranges the number of second image data windows one by one;
[0052] S4. Input the number of second image data windows arranged by the first filter into a second filter for normalization operation to generate a third image data window;
[0053] S5. Perform fusion processing on the first image data window and the third image data window to generate second image data;
[0054] S6. Output the second image data.
[0055] Specifically, in Step 1, we obtain a set of initial image data from a certain source, which is called "First Image Data". In Step 2, according to a pre-set selection interval, we extract a specific data subset from this First Image Data, and this subset is called "First Image Data Window". During this process, we also create some unselected data windows, which are called "Several Second Image Data Windows". The next Step 3 involves two parallel processes. First, we process the First Image Data Window to achieve a so-called "beat delay processing". The specific processing method may vary according to specific applications. At the same time, we send those unselected Second Image Data Windows to the First Filter. The task of the First Filter is to arrange these data windows one by one respectively for the next step of processing. In Step 4, the arranged Second Image Data Windows are input into the Second Filter. The work of the Second Filter is to perform normalization calculations on these data windows, and this process may include certain standardization or normalization steps, generating "Third Image Data Windows". In Step 5, we merge the First Image Data Window and the Third Image Data Window, which may involve some processing methods such as weight allocation, image smoothing or enhancement, and finally generate the "Second Image Data" we need. In the last Step 6, after all these processing steps, we output the generated Second Image Data, which may include saving to a file, displaying on a screen or sending to a specific device or application. This process generally defines a workflow that can effectively process and fuse image data to achieve the expected image processing effect.
[0056] Preferably, the retrieval interval includes an n×m interval of the central part of the First Image Data.
[0057] In one embodiment, the step of performing beat delay processing on the First Image Data Window includes:
[0058] Perform a pause and wait process on the First Image Data Window to wait for the Third Image Data Window generated after calculation for fusion processing.
[0059] This step actually introduces a delay or pause process when processing the First Image Data Window. The purpose of this step is to synchronize two different data windows (i.e., the First Image Data Window and the Third Image Data Window) so that they can be fused simultaneously. This may be because the generation of the Third Image Data Window is relatively late in the whole process, so by performing beat delay processing on the First Image Data Window, it can ensure the temporal matching of the two for appropriate fusion operations.
[0060] In one embodiment, the step of inputting a plurality of second image data windows after the first filter is sorted into a second filter for normalization operation to generate a third image data window includes:
[0061] Using the second filter to splice the pixel rows and columns of a plurality of second image data windows to generate the third image data window.
[0062] This step is about how to use the second filter to process the second image data window and generate the third image data window. The processing method here is to splice the pixel rows and columns after the second image data window is sorted by the first filter. In other words, the second filter will process all the second image data windows, and by splicing their pixel rows and columns in a specified manner, a new and larger image data window is formed, which is the so-called third image data window. The goal of doing this is to perform the next step of processing, merging this third image data window and the first image data window together to form the final second image data.
[0063] In one embodiment, the step of fusing the first image data window and the third image data window to generate the second image data includes:
[0064] Fusing the first image data window and the third image data window according to a preset weight ratio to generate the second image data.
[0065] Specifically, the weight ratio includes taking 50% of the pixel rows and columns of the first image data window and 50% of the pixel rows and columns of the third image data window, which are combined to form the second image data.
[0066] In one embodiment, in the step of obtaining the first image data and outputting the second image data, it includes:
[0067] Obtaining the first image data pixel by pixel through the data_in interface and outputting the output second image data through the data_out interface.
[0068] In one embodiment, the first image data window and the third image data window are represented row by row and column by column by RGBtoken strings. In the step of fusing the first image data window and the third image data window, it includes:
[0069] Fusing the RGBtoken strings corresponding to the first image data window and the third image data window respectively using a color mixing unit, where the color mixing unit includes a color mixing module of ps.
[0070] In this embodiment, this step involves how to perform fusion processing on the first image data window and the third image data window. During this process, each image data window is represented row by row and column by column by an RGB token string. Among them, RGB is an identifier representing the three primary colors of red, green, and blue, which are the basis for constructing almost all colors. Token is a way of representing data used to describe this specific information. In the step of fusion processing, we send the RGB token strings corresponding to the first image data window and the third image data window respectively into the color mixing unit for fusion. The color mixing unit is a component specifically used to process image color and photometric information, and it includes a color mixing module named ps. This module may use a preset color mixing algorithm to adjust the respective red, green, and blue parts, thereby generating a new color mixing result, which constructs the second image data we finally need.
[0071] Refer to the appendix Figure 4 , which is a structural block diagram of an optimized series filtering structure device proposed by the present invention, including:
[0072] An acquisition unit 1 for acquiring first image data;
[0073] An extraction unit 2 for extracting a first image data window from the first image data according to a preset extraction interval and obtaining a number of second image data windows that have not been extracted;
[0074] A first filtering unit 3 for performing beat delay processing on the first image data window, simultaneously sending a number of the second image data windows to a first filter, and arranging the number of second image data windows respectively one by one by the first filter;
[0075] A second filtering unit 4 for inputting the number of second image data windows arranged by the first filter into a second filter for normalization operation to generate a third image data window;
[0076] A fusion unit 5 for performing fusion processing on the first image data window and the third image data window to generate second image data;
[0077] An output unit 6 for outputting the second image data.
[0078] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to that described in the above method embodiment and will not be elaborated here.
[0079] Refer to Figure 5 , and the present invention embodiment also provides a computer device, which may be a server, and its internal structure may be as Figure 5As shown in the figure. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0080] S1. Obtain first image data;
[0081] S2. Retrieve a first image data window from the first image data according to a preset retrieval interval, and obtain a number of second image data windows that have not been retrieved;
[0082] S3. Perform a beat delay process on the first image data window, and at the same time send a number of the second image data windows to a first filter, and the first filter arranges the number of second image data windows one by one;
[0083] S4. Input the number of second image data windows arranged by the first filter into a second filter for normalization operation to generate a third image data window;
[0084] S5. Perform a fusion process on the first image data window and the third image data window to generate second image data;
[0085] S6. Output the second image data.
[0086] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0087] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0088] In summary, by obtaining the first image data; retrieving a first image data window from the first image data according to a preset retrieval interval, and obtaining a plurality of second image data windows that have not been retrieved; performing a beat delay process on the first image data window, and simultaneously sending the plurality of second image data windows to a first filter, and arranging the plurality of second image data windows one by one by the first filter; inputting the plurality of second image data windows arranged by the first filter into a second filter for normalization operation to generate a third image data window; performing a fusion process on the first image data window and the third image data window to generate second image data; outputting the second image data; in order to completely eliminate the line buffer, a method of parallel processing with multiple filters in the first-stage filtering is adopted to output k rows of data simultaneously, and the area of k filters is used to exchange for the buffer area of k / 2 rows.
[0089] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0090] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, apparatus, article, or method including a series of elements not only includes those elements, but also includes other elements that are not explicitly listed, or further includes elements inherent to such a process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, apparatus, article, or method including that element.
[0091] The above are only the preferred embodiments of the present invention, and thus do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present invention.
Claims
1. An optimization method for a series filtering structure, characterized in that, it includes the following steps: Obtain first image data; Retrieve a first image data window from the first image data according to a preset retrieval interval, and obtain several second image data windows that have not been retrieved; the several second image data windows form the first image data window; Perform beat delay processing on the first image data window, and at the same time send the several second image data windows to a first filter, and the first filter arranges the several second image data windows one by one; Input the several second image data windows arranged by the first filter into a second filter for normalization operation to generate a third image data window; Fuse the first image data window and the third image data window to generate second image data; Output the second image data; The step of performing beat delay processing on the first image data window includes: Perform pause and wait processing on the first image data window to wait for fusion processing with the third image data window generated after the operation; The step of inputting the several second image data windows arranged by the first filter into a second filter for normalization operation to generate a third image data window includes: Use the second filter to splice the pixels of the rows and columns of the several second image data windows to generate the third image data window; the first filter processes the several second image data windows in parallel.
2. The optimization method for a series filtering structure according to claim 1, characterized in that, in the step of retrieving a first image data window from the first image data according to a preset retrieval interval and obtaining several second image data windows that have not been retrieved, the retrieval interval includes an n×m interval of the central part of the first image data.
3. The optimization method for a series filtering structure according to claim 1, characterized in that, the step of fusing the first image data window and the third image data window to generate second image data includes: Fuse the first image data window and the third image data window according to a preset weight ratio to generate second image data.
4. The optimization method for a series filtering structure according to claim 3, characterized in that, the weight ratio includes taking 50% of the pixel rows and columns of the first image data window and 50% of the pixel rows and columns of the third image data window, and combining them to form the second image data.
5. The optimization method for a series filtering structure according to claim 1, characterized in that, in the steps of obtaining the first image data and outputting the second image data, it includes: Obtain the first image data pixel by pixel through the data_in interface, and output the second image data through the data_out interface.
6. The optimization method for a series filtering structure according to claim 1, characterized in that, the first image data window and the third image data window are represented row by row and column by column by RGBtoken strings. In the step of fusing the first image data window and the third image data window, it includes: Fuse the RGB token strings corresponding to the first image data window and the third image data window respectively by using a color mixing unit, where the color mixing unit includes a color mixing module of Photoshop.
7. An apparatus for optimizing a cascade filtering structure Characterized in that Comprising: An acquisition unit for acquiring first image data; A retrieval unit for retrieving a first image data window from the first image data according to a preset retrieval interval and obtaining a plurality of second image data windows that have not been retrieved; the plurality of second image data windows constitute the first image data window; A first filtering unit for performing slap delay processing on the first image data window, simultaneously sending the plurality of second image data windows to a first filter, and arranging the plurality of second image data windows one by one by the first filter; A second filtering unit for inputting the plurality of second image data windows arranged by the first filter into a second filter for normalization operation to generate a third image data window; A fusion unit for performing fusion processing on the first image data window and the third image data window to generate second image data; An output unit for outputting the second image data; The step of performing slap delay processing on the first image data window includes: Performing a pause and wait process on the first image data window to wait for fusion processing with the third image data window generated after the operation; The step of inputting the plurality of second image data windows arranged by the first filter into a second filter for normalization operation to generate a third image data window includes: Using a second filter to splice the pixel rows and columns of the plurality of second image data windows to generate the third image data window; the first filter processes the plurality of second image data windows in parallel.
8. A computer device, including a memory and a processor, and a computer program is stored in the memory Characterized in that When the processor executes the computer program, the steps of the cascade filtering structure optimization method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Parallel processing method for synthesis filterbank
KR1020030080268A
Multilevel filters for cache-efficient access
US20150220573A1