Working buffer for storing loop filter intermediate data
The combination of software and hardware with a scratch buffer for piecewise macroblock processing in video decompression systems addresses the inefficiencies of conventional methods, reducing computational and memory demands while effectively eliminating blocking artifacts.
Patent Information
- Application Number
- DE112006000270
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2005-01-25
- Filing Date
- 2006-01-17
- Publication Date
- 2025-06-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Conventional video decompression systems face high computational intensity and memory bandwidth requirements due to inefficient overlap smoothing and in-loop deblocking filtering operations, particularly in hardware-based decoding, which introduces blocking artifacts and requires large memory resources.
A flexible decompression system combining software and hardware uses a processor for upstream operations and a video accelerator for downstream tasks, employing a scratch buffer for piecewise processing of macroblocks to reduce memory bandwidth needs, with an in-loop filter and scratch buffer performing overlap smoothing and deblocking in a macroblock-based manner.
This approach significantly reduces processing requirements and memory needs by allowing efficient, pipelined processing of video data, minimizing the size of on-chip memory, and improving memory access times while effectively reducing blocking artifacts.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Background of the inventionField of the invention
[0001] The present invention relates to video processing technology. In one aspect, the present invention relates to decompressing digital video information. Description of the state of the art
[0002] Because video information requires a large amount of storage space, video information is generally compressed. Consequently, to display compressed video information stored on, for example, a CD-ROM or DVD, the compressed video information must be decompressed to provide the decompressed video information. The decompressed video information is then presented to a display as a bitstream. The decompressed bitstream of video information is typically stored as a bit field, or "bitmap," in memory locations that correspond to pixel positions on a display. The video information required to form a single screen of information on a display is referred to as a block, or frame.A goal of many video systems is to decode compressed video information quickly and efficiently to provide moving images by displaying a sequence of image frames.
[0003] The standardization of recording media, devices, and various aspects of data handling, such as video compression, is highly desirable in view of the continued growth of this technology and its applications. A number of compression and decompression standards have been developed or are currently under development to compress and decompress video information, such as the Motion Picture Expert Group (MPEG) standards for video encoding and decoding (e.g., MPEG-1, MPEG-2, MPEG-3, MPEG-4, MPEG-7, MPEG-21) or the Windows Media video compression standards (e.g., WMV9). Each of these MPEG and WMF standards is hereby incorporated by reference in its entirety in the same manner as if described herein.
[0004] In general, video compression techniques include intra-frame compression and inter-frame compression, which operate to compress video information by reducing both spatial and temporal redundancies contained within video image frames and image blocks, respectively. Intra-frame compression methods use only information contained within the image frame to compress the image frame, which is referred to as an I-frame or image block. Inter-frame compression methods compress image frames with respect to preceding and / or subsequent image frames and are typically referred to as predicted image frames, P-frames, or B-frames. Intra-frame and inter-frame compression methods typically use spatial or block-based coding, whereby a video image frame is divided into blocks for encoding (also referred to as a block transformation process). For example,An I-picture frame is divided into 8x8 blocks. The blocks are encoded using a discrete cosine transform (DCT) coding scheme, where coefficients are encoded as the amplitude of a specific cosine basis function, or other transforms (such as the integer transform). The transformed coefficients are then quantized, producing coefficients with non-zero amplitude levels or sequences (or subsequences) of coefficients with zero amplitude levels. The quantized coefficients are then sequence-level encoded (or sequence-length encoded) to condense the long sequences or runs of zero-valued coefficients.The results are then entropy-encoded in a variable-length encoder (VLC), using statistical coding techniques that assign codewords to the values to be encoded, or using other suitable random coding techniques, such as context-based adaptive binary arithmetic coding (CABAC), adaptive context variable-length coding (CAVLC), and the like. Values that occur more frequently are assigned short codewords, and values that occur less frequently are assigned long codewords. On average, the more frequent, shorter codewords dominate, so the encoding string is shorter than the original data.Thus, spatial or block-based coding techniques compress the digital information associated with a single image. To compress the digital information, video compression techniques such as P-frames and / or B-frames are used to exploit the fact that there is a temporal correlation between consecutive frames. Interframe compression techniques detect the difference between different image frames and then spatially encode the difference information using DCT quantization, sequence length, and entropy coding techniques, although other implementations may employ different block configurations. For example, a P-frame is divided into 16×16 macroblocks (e.g., with four 8×8 luminance blocks and two 8×8 chrominance blocks), and the macroblocks are compressed.Regardless of whether intra-frame compression techniques or inter-frame compression techniques are employed, the use of spatial or block-based coding methods to encode the video data means that the compressed video data is variable-length encoded and otherwise compressed using block-based compression methods in the manner described above.
[0005] At the receiving device or the playback device, the compression steps are reversed to decode the video data processed with block transforms. Fig. 1 shows a conventional system 30 for decompressing video information, comprising an input stream decoding section 35, a motion decoder 38, an adder 39, a frame buffer 40, and a display 41. The input stream decoder 35 receives a stream of compressed video information in the input buffer 31, performs variable-length decoding in the VLC decoder 32, reverses the zigzag pattern and quantization in the inverse quantizer 33, inverts the DCT transformation in the IDCT 34, and provides blocks of statically decompressed video information to the adder 39. In the motion decoding section 38, the motion compensation unit 37 receives motion information from the VLC decoder 32 and a copy of the previous image data (contained in the previous image storage buffer 36) and provides motion-compensated pixels to the adder 39.The adder 39 receives the statically decompressed video information and the motion compensated pixels and supplies decompressed pixels to the frame buffer 40, which supplies the information to the display 41.
[0006] In conventional video coding and decoding structures, blocking artifacts (as noticeable discontinuities between blocks) can be introduced into an image frame from block-based transformation, motion compensation, quantization, and / or other lossy processing steps. Previous attempts to reduce blocking artifacts used overlap smoothing or unblock filtering (in-loop processing or post-processing) to process image frames by smoothing the boundaries between blocks. For example, the WMF9 standard specifies that overlap smoothing and unblocking be performed on the entire image to reduce blocking artifacts.When WMV9 decoding is enabled, overlap smoothing is performed only on the 8 × 8 block boundaries, starting with smoothing in the vertical direction for the entire image frame, and then smoothing is performed on the overlap area in the horizontal direction for the entire image frame. Next, in-loop block cancellation, if enabled, is performed in this order: (i) all 8 × 8 block horizontal boundary lines in the image frame are filtered, starting from the topmost length; (ii) all 8 × 4 sub-block horizontal boundary lines in the image frame are filtered, starting from the topmost line; (iii) all vertical boundary lines of the 8 × 8 blocks are filtered, starting from the leftmost line; and (iv) all vertical boundary lines of the 4 × 8 sub-blocks are filtered, starting from the leftmost line.In conventional solutions, two passes are applied to the entire image frame, with the first pass performing overlap smoothing and the second pass performing in-loop block cancellation. Although other requirements (e.g., according to a QUANT parameter and block types) must be met, which also apply when determining whether or not to perform processing for each step, the intent of these processes is to smooth across the edges of the 16 × 16 macroblocks, the 8 × 8 blocks, or the 4 × 4 subblocks, thereby removing the block processing artifacts introduced by the 2D transformation and quantization.
[0007] In processor-based approaches to handling video decompression, adding a smoothing function or a deblocking function is a computationally intensive filtering process. This sequence of processing can be performed in software if there is a large buffer to accommodate an image frame, for example, VGA-sized at 640 × 480 pixels, which corresponds to 307 kBytes. On the other hand, in hardware-based decoding solutions, smoothing and deblocking are not performed concurrently, and deblocking is performed on the image frame as a whole, requiring a large amount of local memory, which imposes corresponding bus bandwidth requirements and thus sacrifices memory access time.Consequently, there is a strong need to reduce the processing requirements associated with decompression techniques and to improve decompression operations, particularly considering overlap smoothing and / or deblocking filtering operations. Further limitations and disadvantages of conventional systems will become apparent to those skilled in the art after studying the remainder of this application, with reference to the drawings and the following detailed description.
[0008] The publication “A platform-based architecture of loop filter for AVS” in 7 th ICSP'04, 2004 by Sheng, Bin et al. discloses functional details of an AVS decoder, describing a platform-based loop filter structure for removing block artifacts from DCT-transformed data for the Chinese AVS audio and video coding standard.
[0009] The paper "An implemented architecture of deblocking filter for H.264 / AVC" in ICIP '04, 2004 by Sheng, Bin et al. discloses functional details of an AVC decoder, describing a processing sequence for a deblocking filter. Overview of the invention
[0010] A method and a device according to the invention are defined by the independent claims.
[0011] By applying a combination of software and hardware to perform video decompression, a flexible decompression system is provided that is configured to quickly and efficiently execute a variety of different video compression schemes. The flexible decompression system includes a processor for performing upstream decompression steps and a video accelerator for performing downstream decompression steps. In order to reduce the memory bandwidth requirements in the video accelerator for performing overlap smoothing and in-loop deblocking filtering operations on video image data, the in-loop filter is connected to a scratch buffer or storage device that enables piecewise processing in overlap smoothing and in-loop deblocking in a macroblock-based manner. The use of a scratch buffer to perform piecewise processing orPerforming the partial processing of filtering operations in a macroblock-based manner is significantly more efficient than the frame-based approach. Since the size of the scratch buffer is related to the width of the frame, the size of the on-chip memory can be reduced. The scratch buffer size is large enough to accommodate no more than a series of partially filtered blocks from a frame of video data.
[0012] According to one or more embodiments of the present invention, a video processing system, apparatus, and method are provided in which a processor and video decoding circuitry decode video data processed with block transforms into a plurality of macroblocks, each macroblock comprising 8x8 blocks. In conjunction with the decoding operations, an in-loop filter and scratch buffer provided in one or more integrated circuits are used to perform piecewise processing by smoothing and deblocking selected pixel data in a first macroblock to generate one or more completed blocks and one or more partially filtered blocks, wherein at least one of the partially filtered blocks includes control data and pixel data stored in a scratch buffer.As a result, a block adjacent to a previously processed macroblock may be fully filtered for overlap smoothing and deblocking and then output as a completed block during a first filtering operation, while a block adjacent to a subsequently processed macroblock may be partially filtered for overlap smoothing and deblocking and subsequently output as a partially filtered block to the scratch buffer. The scratch buffer is used by the in-loop filter to provide a partially filtered block used to smooth and deblock blocks of pixel data in the first macroblock, the retrieved partially filtered block having been generated during processing of a previous macroblock. By sequentially processing each row orSeries of macroblocks in a video frame for overlap smoothing and deblocking sequentially in each macroblock, the in-loop filter and work buffer can be used to sequentially perform smoothing and deblocking on macroblocks in a pipelined processing.
[0013] Those skilled in the art will appreciate the objects, advantages and other novel features of the present invention from the following detailed description when considered in conjunction with the appended claims and the accompanying drawings. Short description of the drawings Fig. Figure 1 shows a block diagram of a system for decompressing video information. Fig. Figure 2 shows a block diagram of an exemplary video decompression system constructed in accordance with the present invention. Fig. Figure 3 shows a simplified representation of an in-loop filtering process in which a work buffer is used to efficiently handle overlap smoothing and in-loop unblocking in hardware according to a selected embodiment of the present invention. Fig. Figure 4 shows an exemplary technology for reducing the features in a decoded image frame using a smoothing and deblocking filter in a video encoder or decoder. Fig. Figures 5a to k show how piecewise processing can be used to implement smoothing and deblocking procedures for luminance blocks. Fig. Figures 6a to f show how piecewise processing can be used to implement smoothing and deblocking procedures for chrominance blocks. Modes for carrying out the invention
[0014] Although illustrative embodiments of the present invention are described below, it should be understood that the present invention may be practiced without the specific details, and that numerous implementation-specific choices may be made based on the invention as described herein to achieve the particular goals of developers, such as compatibility with system-dependent and business-dependent frameworks that may vary from one implementation to another. Although such a development effort may be complex and time-consuming, it is nevertheless within the skill of the art given the present disclosure. For example, selected aspects are shown in block diagram form rather than detailed illustrations in order to provide a clearer illustration of the present invention.Such descriptions and illustrations are used by those skilled in the art to describe the content of their work and to convey it to others skilled in the art. The present invention will now be described below with reference to the drawings.
[0015] Fig. Figure 2 shows a block diagram of an exemplary video decompression system 100 according to the present invention. As shown, the video decompression system 100 can be implemented in any video display device, such as a desktop or portable computer, a wired or mobile device, a personal digital assistant, a mobile phone or cellular phone, or any other video display device that includes video image functions. As shown in Fig. 2, the video decompression system 100 is configured as a host computer or application processing unit including a bus 95 connected to one or more processors or processing units 50 and a video or media acceleration hardware unit 101. Furthermore, the video decompression system 100 includes a main memory system including a large DDR SDRAM 62, 64 accessed via a DDR controller 60. Additionally or alternatively, one or more memories (e.g., IDE 72, a flash memory unit 74, a ROM 76, etc.) are accessed via the static memory controller 70. The DDR SDRAM or other memories may be integrated into the video compression system 100 or may be external components. Of course, other peripheral devices or display devices (82, 84, 86, 92) may also be accessed via the corresponding controllers 80, 90.For simplicity and clarity, not all elements included in video decompression system 100 are described in detail. Such details are well known to those skilled in the art and may vary depending on the particular computer vendor or microprocessor type. Furthermore, video decompression system 100 may include other buses, devices, and / or subsystems depending on the desired implementation. For example, video decompression system 100 may include cache memories, modems, parallel or serial interfaces, SCSI interfaces, network interface cards, and the like. In the illustrated embodiment, CPU 50 executes the software stored in flash memory 74 and / or SDRAM 62, 64.
[0016] In the Fig. In the video decompression system 100 illustrated in Figure 2, the CPU 50 performs the initial variable-length decoding function, as indicated by the VLD block 52, while the media acceleration hardware unit 101 performs inverse quantization 104, inverse transform 106, motion compensation 108, in-loop filtering 110, color space conversion 112, scaling 114, and filtering 116 on the decoded data. The resulting decoded data may be temporarily stored in the output buffer 118 and / or frame buffer (not shown) before being presented on the display 92.By splitting the decoding processing functions between processor 50 and media acceleration hardware 101, the input decoding steps (e.g., variable-length decoding) can be implemented in software to accommodate a variety of compression schemes (e.g., MPEG-1, MPEG-2, MPEG-3, MPEG-4, MPEG-7, MPEG-21, WMV9, etc.). The decoded data generated at the input is provided to media acceleration hardware 101, which further decodes the decoded data to supply pixel values to output buffer 118 or the image buffer on a macroblock-by-macrblock basis until an image frame is complete.
[0017] During operation, the video decompression system 109 receives a compressed video signal from a video signal source, such as a CD-ROM, a DVD, or other storage device. The compressed video signal is provided as a stream of compressed video information to the processor 50, which executes instructions to decode the variable-length encoded portion of the compressed signal to provide a variable-length decoded data (VLD) signal. Once the software wizard is used to perform variable-length decoding, the VLD data (which includes a header, matrix weights, motion vectors, transformed residual coefficients, and even differential motion vectors) is passed to the media acceleration hardware unit 101, which may be done directly or using the data compression techniques described more fully in U.S. patent application Ser. No.No. 11 / 042,365 (title: "Lightweight Compression of Input Data"). In the medium acceleration hardware unit 101, once the VLD data is received, the data is provided to the inverse zigzag and quantization circuit 104, which decodes the VLD data signal to provide a zigzag decoded signal. The inverse zigzag processing and quantization compensates for the fact that while a compressed video signal is compressed in a zigzag sequence length encoding, the zigzag decoded signal is provided to the inverse DCT circuit 106 as sequential blocks of information. Thus, the zigzag decoded signal provides blocks that are in an order required for raster scanning across the display 92.The zigzag-decoded signal is fed to the inverse transform circuit 106 (e.g., IDCT or inverse integer transform), which performs an inverse discrete cosine transform on the zigzag-decoded video signal on a block-by-block basis to provide statically decompressed pixel values or decompressed error terms. The statically decompressed pixel values are processed on a block-by-block basis by the motion compensation unit 108, which provides intra-frame predicted and bidirectional motion compensation, supporting one, two, and four motion vectors (16x16, 16x8, and 8x8 blocks).The in-loop filter 110 performs overlap smoothing and / or deblocking to reduce or eliminate blocking artifacts according to the WMV9 compression standard by using the scratch buffer 111 to store partially completed macroblock filter data, as described in more detail below. The color space converter 112 converts one or more input data formats (e.g., YcbCr 4:2:0) to one or more output formats (e.g., RGB), and the result is formed and / or scaled in the filter 116.
[0018] As disclosed herein, the in-loop smoothing and deblocking filter 110 removes discontinuities at boundaries between adjacent blocks by partially filtering or processing each row of macroblocks during a first pass. Processing of the partially processed blocks is then completed during processing of the next row of macroblocks. With this technique, a small working buffer 111 is sufficient to store the partially processed blocks in a working buffer, as opposed to using a large memory to store the entire image content for filtering, as is the case in conventional deblocking processes.Since the processing of each block for overlap smoothing and block deblocking is performed on a row-by-row basis, the completed blocks may be output from the filter 110 to a FIFO buffer (not shown) before the blocks are transferred to the CSC 112.
[0019] Fig. Figure 3 shows a simplified representation of a macroblock-based in-loop filtering process using a scratch buffer to efficiently handle overlap smoothing and deblocking according to a selected embodiment of the present invention. In the filtering process, each pass of the in-loop filter for a series of macroblocks produces fully completed blocks (fully filtered to smooth and deblock the blocks) and partially completed blocks (stored in the scratch buffer for subsequent use in smoothing and deblocking the next series of macroblocks). As shown, each individual macroblock (e.g., macroblock 4 or "mb4," which contains four luminance blocks mb4y0, mb4y1, mb4y2, and mb4y3) undergoes the following sequence of processing steps in the in-loop filtering process: (i) fully completing the smoothing and demoblation of the 8x8 blocks (e.g., mb4y0, mb4y1) adjacent to the previous macroblock (e.g., macroblock 1) and partially completing the smoothing and demoblation of the 8x8 blocks (mb4y2, mb4y3) adjacent to the next macroblock (e.g., macroblock 7); (ii) outputting the completed 8x8 blocks (e.g., mb4y0, mb4y1) and storing the partially completed 8x8 blocks (e.g., mb4y2, mb4y3) in the work buffer; (iii) retrieving the partially completed 8x8 blocks (e.g., mb4y2, mb4y3) from the work buffer when the next macroblock (e.g., macroblock 7) is in processing and completing the processing of the retrieved 8x8 blocks (e.g., mb4y2, mb4y3); and (iv) Outputting the completed 8x8 blocks (e.g., mb4y2, mb4y3) together with the completed 8x8 blocks of the next macroblock (e.g., mb7y0, mb7y1).
[0020] Although the details of the implementation may vary depending on the application, Fig. 3 shows an illustrative embodiment, wherein the image frame 150 currently being processed in the in-loop filter 110 is constructed from macroblocks (e.g., macroblocks mb0, mb1, mb2, mb3, mb4, mb5, mb6, mb7, mb8, etc.) and arranged in multiple rows (e.g., a first row 151 formed from mb0, mb1, and mb2). As in Fig. 3a, the in-loop filter 110 has already performed a first pass through the first row of macroblocks 151. As a result of the first pass through the first row 151, the upper blocks (mb0y0, mb0y1, mb1y0, mb1y1, mb2y0, mb2y1) are fully processed with regard to overlap smoothing and block cancellation (as indicated by the hatching), while the lower blocks (mb0y2, mb0y3, mb1y2, mb1y3, mb2y2, mb2y3) are only partially processed with regard to overlap smoothing and block cancellation. For the purpose of completing the partially processed blocks from the first row 151 during the next pass of the macroblock processing, the partially processed blocks from the first row 151 are stored in a work buffer 111 (as indicated by the dot pattern). As additionally shown in Fig. 3a, the in-loop filter 110 has begun processing the second row of macroblocks 152 with a process that completes the partially processed blocks from the first row 151. As a result, blocks mb0y2 and mb3y0 are fully processed with regard to smoothing and block cancellation, and block mb3y2 is only partially processed (and stored in the scratch buffer). With regard to the blocks processed by the filter 110 in Fig. 3a are to be processed (as indicated by the diagonal hatching of the filtered blocks 154), the blocks mb3y1 and mb3y3 are partially processed with regard to smoothing and unblocking (and preserved in the filter 110), the block mb1y2 is a partially completed block retrieved from the scratch buffer, and the remaining blocks (mb4y0 and mb4y2) are obtained from the current macroblock (e.g., macroblock 4). When the in-loop filter 110 processes the filtered blocks 154, the smoothing and unblocking is completed in one or more of the partially processed blocks (e.g., mb0y3, mb3y1), while the remaining blocks mb1y2, mb4y0, mb4y2, and mb3y3) are only partially completed.
[0021] After processing the filtered blocks 156, the loop-internal filter 110 receives new data. This is indicated by the image frame 155, which is Fig. 3b, where it is shown that the filter 110 obtains filtered blocks 156 by outputting completed blocks (e.g., mb0y3, mb3y1), by storing one or more of the partially completed blocks (e.g., mb3y3) in the scratch buffer, by shifting the remaining partially completed blocks (e.g., mb1y2, mb4y0, mb4y2) in the filter by one block position, by retrieving a partially completed block (e.g., mb1y3) from the previous row of macroblocks, and by loading new blocks (e.g., mb4y1, mb4y3) from the current macroblock.When the in-loop filter 110 processes the filtered blocks 156, the smoothing and unblocking is completed in one or more of the partially processed blocks (e.g., mb1y2, mb4y0), while the remaining blocks (mb1y3, mb4y1, mb4y3, and mb4y2) are only partially completed.
[0022] After processing the filtered blocks 156, new data is again inserted into the loop-internal filter 110, as indicated by the image frame 157 shown in Fig. 3c. In particular, the filter 110 obtains filtered blocks 158 by outputting completed blocks (e.g., mb1y2, mb4y0), storing one or more of the partially completed blocks (e.g., mb4y2) in the scratch buffer, shifting the remaining partially completed blocks (e.g., mb1y3, mb4y1, mb4y3) in the filter by one block position, retrieving a partially completed block (e.g., mb2y2) from the previous row of macroblocks, and loading new blocks (e.g., mb5y0, mb5y2) from the current macroblock. When the in-loop filter 110 processes the filtered blocks 158, the smoothing and unblocking is completed in one or more of the partially processed blocks (e.g., mb1y3, mb4y1), while the remaining blocks (mb2y2, mb5y0, mb5y2, and mb4y3) are only partially completed.At this point, the smoothing and deblocking of the upper blocks (mb4y0, mb4y1) in macroblock 4 are completed, but the lower blocks (mb4y2, mb4y3) are only partially completed. By storing the partially stored lower blocks in the work buffer, the filtering operations can be completed when the filter 110 processes the next series of macroblocks.
[0023] Further details of an alternative embodiment of the present invention are set forth in Fig. 4, which illustrates a technique (200) for reducing the blocking effects in a decoded image frame using a smoothing and deblocking filter in a video encoder or decoder. It should be noted that the illustrated technique can be employed to process luminance or chrominance blocks, although there may be edge cases specifically at the edge of each image frame where one skilled in the art will adjust and apply the present invention as needed. For simplicity, the present disclosure focuses primarily on the in-loop filtering steps performed within macroblocks of each image frame.
[0024] According to Fig. 4, once the video encoder / decoder generates at least the first macroblock for the image frame (201), the in-loop filter processes the top row of macroblocks, processing one macroblock at a time to filter the boundaries of each block with its neighboring blocks. As can be seen, there are no partially completed blocks outside the image frame for use in the filtering process, since no smoothing or deblocking is performed at the edge of the image frame. However, once the first row of macroblocks is filtered, the scratch buffer is filled with partially completed blocks. Starting with the first macroblock (201), the encoder / decoder loads the required blocks and retrieves a previously partially completed neighboring block from the scratch buffer (excluding the first row of macroblocks).If a macroblock consists of four blocks of luminance (y0, y1, y2, y3) and two blocks of chroma (Cb, Cr), the blocks are fed into the encoder / decoder hardware in the following order: y0, y1, y2, y3, Cb, Cr.
[0025] The video encoder / decoder then filters predetermined boundaries of the blocks loaded into the filter with neighboring blocks or sub-blocks (210). In a selected embodiment, a piecewise processing technique is employed to partially process each block in the filter. For example, after decoding an 8x8 block in a luminance or chrominance plane, all or part of the left and / or right (vertical) edges are subjected to a smoothing filtering process (211). Additionally or alternatively, all or part of the top and / or bottom (horizontal) edges of the block are subjected to a smoothing filtering process (212). In addition to overlap smoothing, a deblocking filtering process is employed.a filtering process for reducing block-induced phenomena is applied to all or part of the selected horizontal boundary lines of the 8x8 blocks (213) and / or to all or part of the selected horizontal boundary lines of the 8x4 sub-blocks (214). Additionally or alternatively, the block de-filtering process is applied to all or part of the selected vertical boundary lines of the 8x8 blocks (215) and / or to all or part of the selected boundary lines of the 4x4 sub-blocks (216). Once the blocks in the filter have been piecemeal processed, the results are stored or advanced in the filter for further processing. In particular, completed blocks in the filter are fed out of the filter (217) to enable the processing of new data in the filter.In addition, partially completed blocks that are not processed with the new blocks are stored in a work buffer (219) for subsequent use and for further processing with the next row of macroblocks, unless the last row of macroblocks is processed (a negative result from decision 218), in which case the work buffer step (219) is omitted.
[0026] After making space in the filter by storing selected blocks (217, 219), the filter can now process new data. In particular, if there are additional blocks in the image frame (positive result in decision 220), the remaining partially filtered blocks in the filter are shifted to the left (222). For rows below the top row (negative result in decision (224)), the available space in the filter is filled by retrieving the next partially completed block from the work buffer (226) and filling any remaining space in the filter with new blocks (228). Once the filter has loaded the new data, the block filtering process 210 is executed on the new set of filter blocks.By repeating this sequence of operations, each macroblock in the image frame is sequentially filtered to retrieve partially completed blocks from the scratch buffer that were generated during processing of the previous row of macroblocks and to store partially filtered blocks in the scratch buffer memory for subsequent use during processing of the next row of macroblocks. On the other hand, if there are no remaining blocks to be filtered (negative result in decision 222), the smoothing and demolding process for the current image frame is terminated. At this time, the next image frame is fetched (230), and the filtering process is repeated, starting with the first macroblock in the new image frame.
[0027] The Fig. Figures 5a to 5k show an illustrative embodiment of the present invention, illustrating how the WMV9 smoothing and deblocking procedure for luminance blocks in macroblock 4 or ("mb4") can be implemented using a fractional processing technique. The starting point for the filtering process is shown in Fig. 5a, where the filter 320 is already loaded with blocks (e.g., mb0y3, mb3y1, and mb3y3) that have already been processed with regard to overlap smoothing (see, for example, the oval indication 322) and block cancellation (see, for example, the stroke indication 322).
[0028] The filter 320 is then filled with further blocks, as shown in Fig. 5b. In particular, a partially completed block (e.g., mb1y2) is retrieved from the scratchpad and loaded into the filter. Furthermore, selected blocks (e.g., mb4y9, mb4y2) from the current macroblock (e.g., mb4) are loaded into the encoder / decoder, although the loading sequence for the macroblocks (e.g., mb4y0, mb4y1, mb4y2, and mb4y3) requires that at least one of the blocks be loaded at this time but not moved in the filter 320, as indicated by 321.
[0029] When the filter blocks are loaded, the filter 320 performs the overlap smoothing piecewise, as shown in Fig. 5c. In particular, vertical overlap smoothing (V) is performed on selected inner vertical edges 301, 302. Next, horizontal overlap smoothing (H) is performed on selected inner horizontal edges (e.g., 303, 304, 305, 306).
[0030] After the filter blocks are partially smoothed, the filter 320 performs piecewise block cancellation as shown in Fig. 5d. First, a horizontal in-loop block deblocking (HD) is performed at selected 8x8 block boundaries (e.g., 307, 308, 309, 310), followed by a horizontal in-loop block deblocking at selected sub-block boundaries (HDH), e.g., 311, 312, 313, 314. Next, the filter performs a vertical in-loop block deblocking (VD) at selected 8x8 block boundaries (e.g., 315, 316), followed by a vertical in-loop block deblocking at selected sub-block boundaries (VDH) (e.g., 317, 318).
[0031] In each of the smoothing and deblocking steps described above, the order in which the boundary pieces are filtered does not matter, since there is no dependency between the boundary pieces. In addition to the special order of subprocessing described in the Fig. 5c and Fig. 5d, other sequences or filtering steps may also be implemented according to the present invention. For example, a different sequence of boundary edge pieces may be filtered. Furthermore, there may be more than one type of smoothing or deblocking mode applied, and the filtering operation may affect up to 3 or more pixels on each side of the boundary, depending on the filtering scheme. For example, in the MPEG standard, two deblocking modes are used to apply a short filter to 1 pixel on each side of the block edge in one mode, and to apply a longer filter to 2 pixels on each side in the second mode. In other embodiments, the filter definitions, the number of different filters, and / or adaptive filtering conditions may be adapted to meet specific requirements.
[0032] Once the smoothing and unblocking filter operations are finished, the processed filter blocks are saved and moved as shown in Fig. 5e. In particular, completed blocks (e.g., mb0y3, mb3y1) are output as they are completed. Furthermore, one or more partially completed blocks (e.g., mb3y3) may be moved to the work buffer for subsequent use when the processing of the lower neighboring macroblock (see, for example, macroblock 6 in relation to mb3y3 in Fig. 3) is processed. Remaining partially completed blocks in the filter (for example, mb1y2, mb4y4, mb4y2) are then moved within the filter to make room for new data. The result of the output, storage, and move steps is shown in Fig. 5f shown.
[0033] The filter 320 can now be filled with new data blocks, as in Fig. 5g. In particular, a partially completed block (e.g., mb1y3) is retrieved from the work buffer and loaded into the filter. Furthermore, the remaining blocks (e.g., mb4y1, mb4y3) from the current macroblock (e.g., mb4) are loaded into the filter 320.
[0034] Once the filter blocks are loaded, the filter 320 performs piecewise overlap smoothing as in Fig. 5h. In particular, vertical overlap smoothing (V1, V2) is performed on selected inner vertical edges.
[0035] After the filter blocks are partially smoothed, the filter 320 performs piecewise block cancellation as shown in the Fig. 5i. First, a horizontal in-loop block cancellation (HD1, HD2, HD3, HD4) is performed at selected 8x8 block boundaries, followed by an in-loop block cancellation at selected sub-block boundaries (HDH1, HDH2, HDH3, HDH4). Next, the filter performs a vertical in-loop block cancellation (VD1, VD2) at selected 8x8 block boundaries, followed by a vertical in-loop block cancellation at selected sub-block boundaries (VDH1, VDH2).
[0036] Once the smoothing and unblocking filter operations are finished, the processed filter blocks are saved and moved as shown in Fig. 5. In particular, completed blocks (e.g., mb1y2, mb4y0) are output because they are completed. Furthermore, one or more partially completed blocks (e.g., mb4y2) are moved to the scratch buffer for subsequent use when the lower neighboring macroblock is processed. Remaining partially completed blocks in the filter (e.g., mb1y3, mb4y1, mb4y3) are then moved within the filter to make room for new data. The result of the output, storage, and move steps is shown in Fig. 5k, which corresponds exactly to the initial filter state shown in Fig. 5a. Consequently, the sequence of steps described in the Fig. 5a to k, are repeated to continue filtering the filter blocks, including the next macroblock (for example, macroblock 5 or mb5).
[0037] According to the Fig. 6a to f, an illustrative embodiment of the present invention will now be described to show how the WMV9 smoothing and deblocking procedures can be implemented on a Cb or Cr block in a macroblock using a partial processing technique. Since the Cb and Cr macroblocks are similar, the example will be given with respect to an actual Cb macroblock labeled with the index Cb(x,y). The starting point for the filtering process is in Fig. 6a, where the filter 420 is already loaded with blocks (e.g., Cb(x-1, y-1) and Cb(x-1,y)) that are already partially processed with regard to overlap smoothing (see, for example, the oval label 422) and block cancellation (see, for example, the line label 423).
[0038] The filter 420 is then filled with further blocks as in Fig. 6b. In total, a partially completed block (e.g., Cb(x,y-1)) is retrieved from the working buffer and loaded into the filter. Furthermore, the block (e.g., Cb(x,y)) from the current macroblock is loaded into the filter 420. Once the filter blocks are loaded, the filter 420 performs piecewise overlap smoothing, as shown in the Fig. 6c. In particular, vertical overlap smoothing (V) is performed on selected inner vertical edges. Next, horizontal overlap smoothing (H1, H2) is performed on selected inner horizontal edges.
[0039] After the filter blocks are partially smoothed, the filter 420 performs piecewise block cancellation as shown in the Fig. 6d. First, a horizontal in-loop block cancellation (HD1, HD2) is performed at selected 8x8 block boundaries, followed by a horizontal in-loop block cancellation of selected sub-block boundaries (HDH1, HDH2). The filter then performs a vertical in-loop block cancellation (VD) at selected 8x8 block boundaries, followed by a vertical in-loop block cancellation at selected sub-block boundaries (VDH).
[0040] Once the smoothing and unblocking filter operations are finished, the processed filter blocks are saved and moved as shown in Fig. 6e. In particular, the completed block (e.g., Cb(x-1, y-1)) is output since it is completed. Furthermore, the partially completed block (e.g., Cb(x-1, y)) is moved to the scratch buffer for subsequent use when the lower neighboring macroblock is processed. Remaining partially completed blocks in the filter (e.g., Cb(x, y-1), Cb(x,y)) are then moved within the filter to make room for new data. The result of the output, store, and move steps is shown in Fig. 6f, which corresponds exactly to the initial filter state shown in Fig. 6a. Consequently, this sequence can be compared to the Fig. The steps shown in Figures 6a to 6f are repeated to filter the next macroblock.
[0041] As can be seen from the foregoing, by providing a small working buffer in the hardware decoding unit, the in-loop filter can store partially completed filter results from the luminance and chrominance blocks of the current macroblock (denoted as MB(x,y)) in the working buffer. The stored filter results can then be used in processing the blocks adjacent to the macroblock in the row below. In particular, when the filter processes the directly underlying macroblock, i.e., MB(x,y+1), the stored data for MB(x,y) is retrieved from the working buffer and used for processing MB(x,y+1).
[0042] Although the partially completed filter results stored in the scratch buffer should contain at least the 8x8 pixel data, in a selected embodiment, the scratch buffer also stores the control data for determining whether boundary filtering is required for the block. For example, the control data for each block may include the set of headers for the 6 blocks in the current macroblock, including 1 mv or 4 mv selection, block addresses, the block's position in the image frame, the mb mode, transform size, coefficient (0 or non-0), and motion vectors (two forward in the x and y directions and two backward in the x and y directions). The data may be packed to allow efficient application of sequence size.
[0043] Due to the small size of the scratch buffer, the memory can be located on the same chip as the video accelerator, although for typical frame sizes, the scratch buffer can be integrated on a different chip, such as DDR memory or other external memory. By providing the scratch buffer, improved memory access behavior is achieved. By minimizing the size of the scratch buffer, the manufacturing costs for the media acceleration hardware unit can be reduced compared to providing a large buffer to store data blocks for the entire frame. For example, a scratch buffer used to store partially completed filter results, including control data and pixel data, can be calculated as follows: Size of the working buffer approximately (576 bytes) x (number of macroblocks in the horizontal direction in the image frame).
[0044] As can be seen from the foregoing, the size of the scratch buffer is relatively small when the size of the image frame is large in the vertical direction. In other words, the size of the scratch buffer depends on the size of the image frame in the horizontal direction. Note that the fractional processing techniques described herein can be advantageously used in conjunction with a large memory for storing the entire image frame to increase the speed of filtering operations, since filtering can begin at the first macroblock before the entire image frame is decoded. However, the use of a scratch buffer offers advantages in terms of cost and speed compared to using a large memory to store the entire image frame, which entails higher costs and requires longer access times.The use of a work buffer can also be advantageously integrated into pipeline processing in the filter algorithm.
[0045] The specific embodiments disclosed are merely illustrative in nature and should not be construed as limitations on the present invention, for the invention may be modified and practiced in different but equivalent ways as will become apparent to one skilled in the art having the benefit of the present teachings. Therefore, the foregoing description is not intended to limit the invention to the specific form disclosed, but rather is intended to cover such alternatives, modifications, and equivalents as may fall within the spirit and scope of the invention as defined by the appended claims, so that those skilled in the art will recognize that they can make various changes, substitutions, and modifications without departing from the spirit and scope of the invention in its broadest form.
Claims
[1] A method (200) for decoding video data processed by block transformations into a plurality of macroblocks, each macroblock comprising blocks, comprising: performing (210) smoothing and deblocking on selected pixel data in a first macroblock with an in-loop filter to produce at least a first partially filtered block and at least a first completed block; and storing (219) the first partially filtered block in a working buffer for use in smoothing and unblocking selected pixel data in a second macroblock; wherein the in-loop filter sequentially processes each row of macroblocks in a video frame for overlap smoothing and block deblocking, one macroblock at a time, where the size of the working buffer is sufficiently large to accommodate no more than a series of partially filtered blocks from an image frame of video data. [2] The method of claim 1, further comprising: retrieving a second partially filtered block from the working buffer for use in the smoothing and deblocking step, the second partially filtered block being generated during processing of a previous macroblock. [3] The method of claim 1, wherein the in-loop filter sequentially performs smoothing and deblocking on a plurality of macroblocks in a pipelined processing manner. [4] A method according to claim 1, comprising: performing smoothing and deblocking on selected pixel data in a first row of macroblocks to produce a plurality of partially filtered blocks; and Storing the plurality of partially filtered blocks in the working buffer (111). [5] The method of claim 4, further comprising: retrieving a first partially filtered block from the working buffer (111) when selected pixel data in a second row of macroblocks is subjected to smoothing and deblocking; and Performing smoothing and deblocking on selected pixel data in the first partially filtered block to complete the smoothing and deblocking processing on the first partially filtered block to thereby generate a completed block. [6] The method of claim 1, wherein the smoothing and deblocking step (210) comprises: performing the overlap smoothing and the block cancellation piecewise on at least a first block within each macroblock such that the first block is partially processed in a first filtering operation, stored in the working buffer (111) and subsequently completely processed in a second filtering operation. [7] Apparatus in a video processing system (100) for decoding video information from a compressed video data stream, the apparatus comprising: a processor (50) that partially decodes the compressed video data stream to generate partially decoded video data; and a video decoding circuit (101) that decodes the partially decoded video data to generate video frames, the video decoding circuit comprising a work buffer (111) and an in-loop filter (110) for performing partial processing of overlap smoothing and in-loop deblocking sequentially on each row of a plurality of macroblocks in the video frame, each macroblock comprising blocks, and wherein the working buffer (111) stores partially filtered blocks from a first series of macroblocks, so that each partially filtered block can be retrieved by the in-loop filter during overlap smoothing and deblocking of a second series of macroblocks, wherein the size of the working buffer (111) is sufficiently large to accommodate no more than a series of partially filtered blocks from an image frame of video data.
Citation Information
Patent Citations
11/042,365