A hardware implementation method for image dynamic search window block matching

By dynamically adjusting the search window and reducing the search window in the BM3D algorithm, the problem of high resource consumption cost in the hardware implementation of traditional BM3D algorithm is solved, and hardware resource saving and manufacturing cost reduction are achieved.

CN119672379BActive Publication Date: 2025-05-23BRITE SEMICON SHANGHAI CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510180280.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-23
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

The traditional BM3D algorithm has high resource consumption costs during the block matching hardware implementation process, resulting in large hardware resource utilization and increased manufacturing costs.

Method used

In image denoising processing, the search window is reduced and the search window is dynamically adjusted according to the position of the image block, thereby reducing the number of SRAM and control logic. The search window is constructed using a pipeline method. The method of dynamically adjusting the search window uses a smaller search window at the corners of the image.

Benefits of technology

It realizes that while ensuring image denoising quality, reduces hardware resources, reduces chip production costs, and simplifies control logic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672379B_ABST
    Figure CN119672379B_ABST
Patent Text Reader

Abstract

The present invention discloses a hardware implementation method for image dynamic search window block matching, which belongs to the field of image dynamic search technology. The implementation method operates to reduce the image search window while ensuring the image denoising quality, and then dynamically adjusts the search window according to the position of the image block, thereby reducing the number of SRAMs and reducing control logic in the block matching operation process. The present invention adopts dynamic search window technology and implements the block matching part of the BM3D algorithm with less SRAM hardware. The control logic of the technology is easy to implement, can reduce the occupation of hardware resources, and thus reduce the chip production cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image dynamic search, and in particular relates to a hardware implementation method for image dynamic search window block matching. Background Art

[0002] The BM3D (three-dimensional block matching) algorithm has obvious advantages in image denoising. The traditional or conventional BM3D algorithm traverses each block one by one in the step of finding similar pixel blocks by performing block matching operations. The larger the size of the matching block, the larger the size of the matching search window, and the smaller the step of the matching sliding window, the longer the software calculation takes, and the more complex the hardware implementation method is and the greater the consumption of hardware resources. The BM3D algorithm was originally implemented in software. Although some hardware circuits have been implemented, the problem is that it consumes a lot of hardware resources, resulting in higher manufacturing costs. Summary of the invention

[0003] The purpose of the present invention is to provide a hardware implementation method for image dynamic search window block matching, which not only helps to reduce resource consumption in the hardware implementation process of the BM3D algorithm, but is also easy to design and implement, and can solve the problem of high resource consumption cost in the hardware implementation process of the BM3D algorithm block matching.

[0004] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: a hardware implementation method for image dynamic search window block matching, which operates to reduce the image search window while ensuring the image denoising quality, and then dynamically adjusts the search window according to the position of the image block, thereby reducing the number of SRAMs and reducing the control logic during the block matching operation process.

[0005] Preferably, the operation of reducing the image search window and then dynamically adjusting the search window according to the position of the image block specifically includes:

[0006] The image block matching takes the minimum value of the sum of absolute differences (SAD) as the best match; then the four minimum SAD values ​​containing itself are found in the dynamic search window as the matching block; finally, the read data of SRAM is used to construct the search window.

[0007] Preferably, constructing a search window with the read data of the SRAM specifically includes:

[0008] When data is read out from SRAM, a pipeline method is used to construct the search window, and a method of dynamically adjusting the search window is adopted, and a smaller search window is used at the corners of the image.

[0009] Preferably, the image block matching size is 4x4, three pseudo dual-port SRAMs are used, the bit width is the width of the image pixels after DCT two-dimensional transformation, and the depth is the number of pixels in one row of the image.

[0010] Preferably, the three pseudo dual-port SRAMs perform read operations synchronously while writing, and the write operation of each SRAM is one clock cycle earlier than the read operation to ensure that the written data is read.

[0011] Preferably, the writing and reading method of the three pseudo dual-port SRAMs is as follows:

[0012] a. Store the DCT two-dimensional transformed image data blocks of rows 1 to 4 into SRAM;

[0013] b. Store the DCT two-dimensional transformed image data blocks of 5 to 8 rows into SRAM2. After a set of data is stored in SRAM2, start reading SRAM1 and SRAM2.

[0014] c. Store the DCT two-dimensional transformed image data blocks of 9 to 12 rows into SRAM3. After a group of data is stored in SRAM3, start reading SRAM1, SRMA2, and SRAM3;

[0015] d. The DCT two-dimensional transformed image data blocks after 12 rows are stored in SRAM1, SRAM2, and SRAM3 in groups of four rows, and SRAM1, SRAM2, and SRAM3 are read at the same time;

[0016] e. Repeat step d to the last four lines of the image. When the current frame ends, wait for a new frame to start and then go to step a again, repeating the cycle.

[0017] Preferably, the search window is dynamically adjusted according to the current position of the image block, and a frame of image data is divided into nine areas, and the nine areas are named A, B, C, D, E, F, G, H, and I, and the arrangement is as follows: ;

[0018] Among them, area A is the four blocks in the upper left corner of the image; area B is the six blocks on the upper edge of the image; area C is the four blocks in the upper right corner of the image; area D is the six blocks on the left edge of the middle area of ​​the image row; area E is the nine blocks in the middle of the image; area F is the six blocks on the left edge of the middle area column of the image row; area G is the four blocks in the lower right corner of the image; area H is the six blocks on the lower edge of the image; area I is the four blocks in the lower right corner of the image;

[0019] Among them, the size of the image block is 4x4; the execution order of a frame search window is area A, area B, area C, area D, area E, area F, area G, area H, and area I; when a frame ends, wait for the next frame to arrive, and the search window returns to area A.

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] The present invention adopts dynamic search window technology and implements the block matching part of the BM3D algorithm with less SRAM hardware. The control logic of this technology is easy to implement, which can reduce the occupation of hardware resources and thus reduce the chip production cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a processing flow chart of the BM3D basic estimation part of a hardware implementation method for image dynamic search window block matching of the present invention.

[0023] Figure 2 The present invention is a schematic diagram of a dynamic search window of an image dynamic search window block matching hardware implementation method.

[0024] Figure 3 The present invention is an operation timing diagram of a hardware implementation method for image dynamic search window block matching.

[0025] Figure 4 This is a block matching processing timing diagram of a hardware implementation method for image dynamic search window block matching of the present invention.

[0026] Figure 5 The present invention discloses a state machine state transfer condition diagram of a hardware implementation method for image dynamic search window block matching. DETAILED DESCRIPTION

[0027] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] The BM3D (three-dimensional block matching) algorithm is a technology used for image denoising and is widely used in the fields of computer vision and image processing. The basic idea of ​​the algorithm is to use the similarity of blocks in an image for denoising.

[0029] The BM3D basic estimation process is as follows Figure 1 The present invention optimizes and improves the hardware implementation structure and method of matching similar blocks, namely Figure 1 The block matching part in .

[0030] Considering the hardware resource consumption and logic complexity of the algorithm, it needs to be optimized during the implementation process. Since the block matching in the BM3D process needs to read the data after the two-dimensional DCT transformation from the SRAM, the size of the SRAM will affect the consumption of hardware resources. The larger the search window size, the greater the resolution of the video image, and the larger the number and depth of the SRAM, which will inevitably affect the overall resource occupancy and further affect the manufacturing cost of the chip. In addition, a large search window and a large number of SRAMs will inevitably increase the control logic.

[0031] To this end, in order to address this problem, the present invention reduces the search window while ensuring the quality of image denoising, and dynamically adjusts the search window according to the position of the image block, which not only reduces the number of SRAMs, but also reduces the control logic during the block matching operation.

[0032] The image block matching size of the present invention is 4x4, and three pseudo dual-port SRAMs are used. The bit width is the width of the image pixel after DCT two-dimensional transformation, and the depth is the number of pixels in one row of the image. The writing and reading method of the three SRAMs is as follows:

[0033] a. Store the DCT two-dimensional transformed image data blocks of rows 1 to 4 into SRAM1;

[0034] b. Store the DCT two-dimensional transformed image data blocks of 5 to 8 rows into SRAM2. After a set of data is stored in SRAM2, start reading SRAM1 and SRAM2.

[0035] c. Store the DCT two-dimensional transformed image data blocks of 9 to 12 rows into SRAM3. After a group of data is stored in SRAM3, start reading SRAM1, SRMA2, and SRAM3;

[0036] d. The DCT two-dimensional transformed image data blocks after 12 lines are stored in SRAM1, SRAM2, and SRAM3 in groups of four lines each (the old data is overwritten), and SRAM1, SRAM2, and SRAM3 are read at the same time;

[0037] e. Repeat step d to the last four lines of the image. When the current frame ends, wait for a new frame to start and then go to step a again, repeating the cycle.

[0038] The present invention dynamically adjusts the search window according to the current position of the image block and divides a frame of image data into nine areas. For ease of description, the nine areas are named A, B, C, D, E, F, G, H, and I, and the arrangement is as follows: . Area A is the four blocks at the upper left corner of the image. B is the six blocks at the upper edge of the image; C is the four blocks at the upper right corner of the image; D is the six blocks at the left edge of the middle area of ​​the image row; E is the nine blocks in the middle of the image; F is the six blocks at the left edge of the middle area of ​​the image row; G is the four blocks at the lower right corner of the image; H is the six blocks at the lower edge of the image; I is the four blocks at the lower right corner of the image. The block size is 4x4. The execution order of a frame search window is A—>B—>C—>D—>E—>F—>G—>H—>I. When a frame ends, wait for the next frame to arrive, and the search window returns to A.

[0039] When data is read out from SRAM, a pipeline method is used to construct a search window. Since the information of an image is mainly concentrated in the middle area, in order to obtain the final similar matching block, the present invention adopts a method of dynamically adjusting the search window. A smaller search window can be used at the corners of the image, thereby reducing the number of SRAMs and control logic.

[0040] Because the video stream is composed of multiple frames of images, for ease of description, the embodiment of the present invention is described using a Y channel of a YUV video format with a resolution of 1280x720.

[0041] The three SRAMs are pseudo dual-port type, with a data write and read bit width of 132 bits and a depth of 1280. The Y channel contains the brightness information of the image. Here, the two-dimensional DCT transform result of the Y channel is used as the block matching information. The dynamic search window is as follows: Figure 2 , the shaded part in each area is the current block, and the white area is the reference block. The operation sequence is as follows Figure 3 .

[0042] HSYNC is the line synchronization signal of a frame. CNT_FLG is a counter that counts in a cycle of 4. When it is 3, it means that the two-dimensional DCT transformation of four lines can be cached. At the same time, the BLOCK_CNT[1:0] count is increased by 1. When BLOCK_CNT[1:0] is 2, the counter returns to zero. DCT2_DATA is the data block after the two-dimensional transformation. When CNT_FLG is 3 and BLOCK_CNT[1:0] is 0, WEN1 is 1. The image block after two-dimensional DCT transformation will be written into SRAM1 in a pipeline manner according to the address WADDR1 sequence; when CNT_FLG is 3 and BLOCK_CNT[1:0] is 1, WEN2 is 1, and the image block after two-dimensional DCT transformation will be written into SRAM2 in a pipeline manner according to the address WADDR2 sequence; when CNT_FLG is 3 and BLOCK_CNT[1:0] is 2, WEN3 is 1, and the image block after two-dimensional DCT transformation will be written into SRAM3 in a pipeline manner according to the address WADDR3 sequence, where the address range of WADDR(1-3) is 0~1279, and the subsequent operations are similar.

[0043] Since the block matching uses a pseudo dual-port SRAM, a read operation is performed while writing. Each SRAM write operation is one clock cycle earlier than the read operation to ensure that the written data is read. Figure 3 The readout timing is also marked by CNT_FLG, but one clock cycle later than the write. When CNT_FLG is 3, the REN1_3 signal is pulled high, and the image block data DATA1, DATA2, DATA3 are read out in the order of RDADDR1_3. The readout image block is used to construct the search window of the matching block.

[0044] The present invention adopts a dynamic search window method, and the block matching takes the minimum value of the sum of absolute differences (SAD) as the best match. In the dynamic search window, find the four matching blocks with the minimum SAD value that include itself. The read data of SRAM is used to construct the search window. The processing sequence is as follows: Figure 4As shown. In the first step, Block1_r1, Block1_r2, Block1_r3, Block1_r4, Block2_r1, Block2_r2, Block2_r3, Block2_r4, Block3_r1, Block3_r2, Block3_r3, Block3_r4 are 12 lines of data read from SRMA. When BLOCK_CNT = 0, the valid data lines are Block1_r1, Block1_r2, Block1_r3, Block1_r4. When BLOCK_CNT = 1, the valid data lines are Block1_r1, Block1_r2, Block1_r3, Block1_r4, Block2_r1, Block2_r2, Block2_r3, Block2_r4. When BLOCK_CNT = 2, the valid data lines are Block1_r1, Block1_r2, Block1_r3, Block1_r4, Block2_r1, Block2_r2, Block2_r3, Block2_r4, Block3_r1, Block3_r2, Block3_r3, Block3_r4. In the second step, Block1~3_col1, Block1~3_col2, Block1~3_col3, Block1~3_col4, Block1~3_col5, Block1~3_col6, Block1~3_col7, Block1~3_col8, Block1~3_col9, Block1~3_col10, Block1~3_col11, and Block1~3_col12 are register samples of 12 data in each column. The purpose of this sampling operation is to write the first 12 data of the image data into the register matrix, that is, when the number of columns is greater than 12, it is rolled forward in units of 4 pixels to prepare for the next step of placing all pixels in the search window at the same time position. In the third step, S1_col1~12, S2_col1~12, S3_col1~12, S4_col1~12, S5_col1~12, S6_col1~12, S7_col1~12, S8_col1~12, S9_col1~12, S10_col1~12, S11_col1~12, and S12_col1~12 perform a second sampling on the data Block1~3_col4, Block1~3_col8, and Block1~3_col12 processed in the second step in the previous clock cycle before the change, thereby obtaining a block matching search window with a maximum of 12x12 pixels.

[0045] The next step is the block matching operation, which is to calculate the SAD value between the current block in the search window and the matching block. Because we need to get the final four matching blocks, the four image blocks at the corners of the image, that is, Figure 2 There is no need to calculate the SAD values ​​for A, C, G, and I, and the four matching blocks are directly output. The remaining parts need to compare the SAD values ​​and select the four matching blocks with the smallest values ​​for output. Figure 4 The final maximum search window is 12x12 pixels, which already includes Figure 2 In the form of 9 search windows. For example, when the current block is 1 to 4 lines, it is not necessary to wait until the 9th to 12th lines of SRAM data are read out, and only the first 8 lines of image data are sufficient, so that the block matching operation can be performed 4 lines in advance.

[0046] Figure 4 When the search window matches the block data output, combined with Figure 3 When BLOCK_CNT = 1, the 1st to 4th row data block and the 5th to 8th row data block contain Figure 2 A, B, C search windows; when BLOCK_CNT=2, 1~4 rows of data blocks, 5~8 rows of data blocks and 9~12 rows of data blocks contain Figure 2 D, E, F search windows; when BLOCK_CNT > 2 and BLOCK_CNT < (720 / 4)-1, the data block also contains Figure 2 D, E, and F search windows; when BLOCK_CNT=179, the last 717 to 720 rows and 713 to 716 rows of data contain Figure 2 Middle G, H, I search window.

[0047] The matching process can be implemented using a state machine, which corresponds to Figure 2 In the search window, the state transition conditions need to be set during design, such as Figure 5 HSYNC_d8 is Figure 3 The line synchronization signal HSYNC is delayed by 8 clock cycles; HSYNC_d9 is Figure 3The horizontal synchronization signal HSYNC is delayed by 9 clock cycles. SAD_CNT is the column counter, and VCNT is the row counter, which starts counting from 0. Because the first block match at the beginning of a frame starts when the 5th to 8th row data is output, it is necessary to generate 4 horizontal synchronization signals—HSYNC_EXT after the last HSYNC in a frame, and the number of valid cycles is the same as HSYNC. HSYNC_EXT_d8 is a signal that the horizontal synchronization signal HSYNC_EXT is delayed by 8 clock cycles, and HSYNC_EXT_d9 is a signal that the horizontal synchronization signal HSYNC_EXT is delayed by 8 clock cycles. SAD_CNT is a counter when the level of HSYNC_d9 or HSYNC_EXT_d9 is high, and it starts counting from 0. When HSYNC_d8 is high and HSYNC_d9 is low, and BLOCK_CNT=1, the state machine enters the ROW1_BLK_FST state from the IDLE state; when SAD_CNT=3, the state machine jumps to ROW1_BLK_NEQ1; when SAD_CNT=1275, the state machine jumps to ROW1_BLK_END; when SAD_CNT =1279, the state machine jumps to ROWN_BLK_FST, and in the ROWN_BLK_FST state, SAD_CNT has been restored to 0 and continues to count; when SAD_CNT=3, the state machine jumps to ROWN_BLK_NEQ1; when SAD_CNT=1275, the state machine jumps to ROWN_BLK_END; when SAD_CNT=1279, if VCNT = 0, and the field synchronization signal VSYNC is high, then jump to ROWEND_BLK_FST, otherwise jump back to ROWN_BLK_FST. After jumping to ROWN_BLK_END, when SAD_CNT = 3, the state machine jumps to ROWEND_BLK_NEQ1; when SAD_CNT = 1275, the state machine jumps to ROWEND_BLK_END, and when SAD_CNT = 1279, the state machine jumps back to IDLE state and waits for the start of the next frame. ROW1_BLK_FST corresponds to Figure 2 In area A, ROW1_BLK_NEQ1 corresponds to Figure 2 In area B, ROW1_BLK_END corresponds to Figure 2 In area C, ROWN_BLK_FST corresponds to Figure 2 In area D, ROWN_BLK_NEQ1 corresponds to Figure 2 In area E, ROWN_BLK_END corresponds to Figure 2 Region F in ROWEND_BLK_FST corresponds to Figure 2 Middle area G, ROWEND_BLK_NEQ1 corresponds Figure 2In area H, ROWEND_BLK_END corresponds to Figure 2 Middle area I. In ROW1_BLK_FST, ROW1_BLK_END, ROWEND_BLK_FST, and ROWEND_BLK_END states, four matching blocks are directly output; in ROW1_BLK_NEQ1, ROWN_BLK_FST, ROWN_BLK_NEQ1, ROWN_BLK_END, and ROWEND_BLK_NEQ1 states, the SAD values ​​of the current block and the reference block are calculated to obtain the final four matching blocks.

[0048] The present invention adopts dynamic search window technology and implements the block matching part of the BM3D algorithm with less SRAM hardware. The control logic of this technology is easy to implement, which can reduce the occupation of hardware resources and thus reduce the chip production cost.

[0049] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0050] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A hardware implementation method for image dynamic search window block matching, characterized in that: The implementation method reduces the image search window while ensuring the image denoising quality, and then dynamically adjusts the search window according to the position of the image block, thereby reducing the number of SRAMs and reducing control logic in the block matching operation process; The operation of reducing the image search window and then dynamically adjusting the search window according to the position of the image block specifically includes: The image block matching takes the minimum value of the sum of absolute differences (SAD) as the best match; then find the four minimum SAD values ​​containing itself in the dynamic search window as the matching block; finally, construct the search window with the read data of SRAM; The search window is dynamically adjusted according to the current position of the image block, and a frame of image data is divided into nine areas. The nine areas are named A, B, C, D, E, F, G, H, and I, and the arrangement is as follows: ; Among them, area A is the four blocks in the upper left corner of the image; area B is the six blocks on the upper edge of the image; area C is the four blocks in the upper right corner of the image; area D is the six blocks on the left edge of the middle area of ​​the image; area E is the nine blocks in the middle of the image; area F is the six blocks on the right edge of the middle area of ​​the image; area G is the four blocks in the lower left corner of the image; area H is the six blocks on the lower edge of the image; area I is the four blocks in the lower right corner of the image; The size of the image block is 4x4; the execution order of the search window of one frame is area A, area B, area C, area D, area E, area F, area G, area H, and area I; when one frame ends, wait for the next frame to arrive, and the search window returns to area A; The step of constructing a search window with the read data of the SRAM specifically includes: When data is read out from SRAM, the search window is constructed in a pipeline manner, and a method of dynamically adjusting the search window is adopted, and a smaller search window is used at the corners of the image; The image block matching size is 4x4, using three pseudo dual-port SRAMs, the bit width is the width of the image pixels after DCT two-dimensional transformation, and the depth is the number of pixels in one row of the image; The three pseudo dual-port SRAMs perform read operations synchronously while writing, and the write operation of each SRAM is one clock cycle earlier than the read operation to ensure that the written data is read.

2. The method for implementing image dynamic search window block matching hardware according to claim 1, characterized in that: The writing and reading method of the three pseudo dual-port SRAMs is as follows: a. Store the DCT two-dimensional transformed image data blocks of rows 1 to 4 into SRAM; b. Store the DCT two-dimensional transformed image data blocks of 5 to 8 rows into SRAM2. After a set of data is stored in SRAM2, start reading SRAM1 and SRAM2. c. Store the DCT two-dimensional transformed image data blocks of 9 to 12 rows into SRAM3. After a group of data is stored in SRAM3, start reading SRAM1, SRMA2, and SRAM3; d. The DCT two-dimensional transformed image data blocks after 12 rows are stored in SRAM1, SRAM2, and SRAM3 in groups of four rows, and SRAM1, SRAM2, and SRAM3 are read at the same time; e. Repeat step d to the last four lines of the image. When the current frame ends, wait for a new frame to start and then go to step a again, repeating the cycle.

Citation Information

Patent Citations

  • Method for estimating block matching motion of H.264 encode

    CN101378504A