Method and apparatus for neighboring block availability in video coding
Patent Information
- Application Number
- JP2023084821
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-04-25
- Filing Date
- 2023-05-23
- Publication Date
- 2025-07-25
AI Technical Summary
Existing video encoding and decoding technologies face challenges in efficiently utilizing neighboring block information for intra-block copy prediction, leading to suboptimal compression ratios and increased data requirements.
The method involves determining the prediction mode of neighboring blocks and inserting relevant prediction information into a predictor list for the current block, considering factors like slice, tile, and overlap to enhance intra-block copy prediction efficiency.
This approach improves video encoding efficiency by optimizing the use of neighboring block data, reducing redundancy and enhancing compression performance.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Incorporation by Reference This disclosure claims the benefit of priority of U.S. Patent Application No. 16 / 394,071, filed Apr. 25, 2019, entitled “METHODS FOR NEIGHBORING BLOCK AVAILABILITY OF INTRA BLOCK COPY PREDICTION MODE,” which claims the benefit of priority of U.S. Provisional Application No. 62 / 704,053, filed Feb. 6, 2019, entitled “METHODS FOR NEIGHBORING BLOCK AVAILABILITY OF INTRA BLOCK COPY PREDICTION MODE.” The entire disclosure of the above applications is hereby incorporated by reference in its entirety.
[0002] This disclosure generally describes embodiments related to video encoding.
Background Art
[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. Work of the inventors named herein, in so far as it is described in this background section, may not constitute prior art at the time of filing with respect to aspects of the description that may not qualify as prior art. Nothing in this disclosure is to be taken as an admission, whether express or implied, that the described matter is prior art to the present disclosure.
[0004] Video encoding and decoding can be performed using interpicture prediction with motion compensation. Uncompressed digital video can contain a series of pictures, each picture having spatial dimensions of, for example, 1920x1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (informally also called frame rate), for example, 60 frames per second or 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920x1080 luminance sample resolution at a 60 Hz frame rate) requires a bandwidth of nearly 1.5 Gbit / s. Using such video for one hour requires more than 600 GB of storage space.
[0005] One of the purposes of video encoding and decoding is to reduce the redundancy of the input video signal through compression. Compression helps to reduce the aforementioned bandwidth or storage space requirements, sometimes by more than two orders of magnitude. Both lossless and lossy compression, and combinations thereof, can be used. Lossless compression refers to a technique that allows an exact copy of the original signal to be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough that the reconstructed signal is still useful for its intended purpose. For video, lossy compression is widely adopted. The amount of distortion that is acceptable varies depending on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that the greater the acceptable / acceptable distortion, the higher the compression ratio can be.
[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy coding.
[0007] Video codec technology can include a technique called intra coding. In intra coding, sample values are represented without referencing samples or other data from a previously reconstructed reference picture. In some video codecs, the picture is spatially divided into blocks of samples. A picture can be an intra picture if all blocks of samples are coded in intra mode. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and therefore can be used as the first picture in an encoded video bitstream and video session, or as a still image. Samples in intra blocks can be transformed, and the transformation coefficients can be quantized before entropy coding. Intra prediction is a technique that minimizes the sample values of the domain before transformation. In some cases, the smaller the DC value and the smaller the AC coefficient after transformation, the fewer bits are needed at a particular quantization step size to represent the block after entropy coding.
[0008] For example, traditional intra-encoding, as known from MPEG-2 generation encoding techniques, does not use intra-prediction. However, some newer video compression techniques include those that attempt to do so from metadata obtained during encoding / decoding of surrounding sample data and / or blocks of data that are spatially adjacent and preceding in the decoding order. Such techniques will hereafter be referred to as “intra-prediction” techniques. It should be noted that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not from the reference picture itself.
[0009] Intra-prediction can take various forms. If two or more of these techniques can be used in a given video coding technique, the techniques used can be coded in intra-prediction mode. In some cases, a mode may have submodes and / or parameters, which can be coded individually or included in a mode codeword. The codeword used for a particular combination of mode / submode / parameters can affect the coding efficiency of intra-prediction, as can the entropy coding technique used to convert the codeword into a bitstream.
[0010] A specific mode of intra-prediction was introduced in H.264, improved in H.265, and further refined with new coding techniques such as the Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). A predictor block can be formed using adjacent sample values belonging to already available samples. The sample values of adjacent samples are copied to the predictor block according to direction. The reference to direction in use may be encoded in the bitstream or may be predicted on its own.
[0011] Referring to Figure 1A, the lower right shows a subset of nine predictor directions known from the 33 possible predictor directions in H.265 (corresponding to the 33 angular modes of the 35 intra-modes). The point where the arrows converge (101) represents the predicted sample. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.
[0012] Continuing to refer to Figure 1A, a 4x4 sample square block (104) is depicted in the upper left (indicated by a dashed bold line). The square block (104) contains 16 samples, each labeled "S", with its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in block (104) in both the Y and X dimensions. Since the block size is 4x4 samples, S44 is in the lower right. Further reference samples are shown following a similar numbering scheme. The reference samples are labeled with R relative to block (104), its Y position (e.g., row index), and its X position (column index). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed, and therefore negative values do not need to be used.
[0013] Intra-picture prediction can work by copying a reference sample value from an adjacent sample, depending on the prediction direction of the signal. For example, suppose the encoded video bitstream contains signaling for this block indicating the prediction direction, which corresponds to arrow (102). That is, the sample is predicted from one or more prediction samples to the upper right, at an angle of 45 degrees from the horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.
[0014] In certain cases, to calculate a reference sample, especially when the direction is not evenly divisible by 45 degrees, the values of multiple reference samples can be combined, for example, by interpolation.
[0015] As video coding techniques have advanced, the number of possible directions has increased. H.264 (2003) could represent nine different directions. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at the time of disclosure. Experiments have been conducted to identify the most likely directions, using specific techniques of entropy coding to represent those possible directions with a small number of bits, accepting a specific penalty for less likely directions. Furthermore, the direction itself can sometimes be predicted from adjacent directions used in adjacent, already decoded blocks.
[0016] Figure 1B shows a schematic diagram (105) illustrating the 65 intra-prediction directions by JEM to illustrate the number of prediction directions that increase over time.
[0017] The mapping of intra-predicted direction bits within an encoded video bitstream representing direction can vary from video coding technique to video coding technique, ranging from complex adaptive schemes involving intra-predicted modes, codewords, and most likely modes from predicted direction to simple direct mappings to similar techniques. However, in all cases, there can be certain directions that are statistically less likely to occur in video content than certain other directions. Since the goal of video compression is to reduce redundancy, these less likely directions will be represented by more bits than the more likely directions in a well-functioning video coding technique.
[0018] Motion compensation is a potentially lossy compression technique in which blocks of sample data from a previously reconstructed picture or a portion of it (a reference picture) are spatially shifted in the direction indicated by a motion vector (MV), and then used to predict a newly reconstructed picture or portion of a picture. In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, the third of which indicates the reference picture in use (the latter can indirectly be the time dimension).
[0019] In some video compression techniques, the motion vector (MV) applicable to a specific region of sample data can be predicted from other MVs, for example, from MVs related to another region of sample data that is spatially adjacent to the region being reconstructed and precedes that MV in the order of decoding. This significantly reduces the amount of data required to encode the MV, thus eliminating redundancy and improving compression. For example, when encoding an input video signal derived from a camera (called natural video), MV prediction can be effective because regions larger than the region to which a single MV applies are statistically more likely to move in the same direction, and therefore, in some cases, can be predicted using similar motion vectors derived from MVs of adjacent regions. As a result, the MV detected in a particular region is similar to or identical to the MV predicted from the surrounding MVs and can be represented with fewer bits after entropy coding than would be used if the MV were encoded directly. In some cases, MV prediction can be an example of lossless compression of the signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself may suffer losses, for example, due to rounding errors when calculating the predictor from several surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec.H.265, "High Efficiency Video Coding," December 2016). Of the many MV prediction mechanisms provided by H.265, the one described here is a technique hereafter referred to as "spatial merging."
[0021] Referring to Figure 1C, the current block (111) can be predictable from a previous block of the same size that has been spatially shifted, and may include samples found by the encoder during the motion search process. Instead of directly encoding its MV, the MV can be derived from the most recent (in decoding order) reference picture using the MV associated with one or more reference pictures, for example, A0, A1, and B0, B1, B2 (112-116, respectively) that are associated with one of the five surrounding samples. In H.265, MV prediction can use predictors from the same reference pictures used by the adjacent blocks. [Overview of the project] [Means for solving the problem]
[0022] Aspects of this disclosure provide methods and apparatus for video coding / decoding. In some examples, the apparatus for video decoding includes a receiving circuit and a processing circuit.
[0023] The processing circuit is configured to decode the prediction information for the current block in the current encoded picture, which is part of the encoded video sequence. The prediction information indicates a first prediction mode to be used for the current block. The processing circuit determines whether adjacent blocks adjacent to the current block and reconstructed before the current block use the first prediction mode. Next, in response to the decision that adjacent blocks use the first prediction mode, the processing circuit inserts the prediction information from adjacent blocks into the predictor list for the first prediction mode. Finally, the processing circuit reconstructs the current block according to the predictor list for the first prediction mode.
[0024] According to one aspect of the present disclosure, in response to a determination that an adjacent block uses a second prediction mode different from the first prediction mode, the processing circuit is further configured to determine whether the second prediction mode is an intra prediction mode. Next, in response to a determination that the second prediction mode is not an intra prediction mode, the processing circuit inserts prediction information from the adjacent block into the predictor list of the first prediction mode.
[0025] In one embodiment, the processing circuit determines whether an adjacent block is in the same slice as the current block. A slice is a group of blocks in raster scan order, and the group of blocks within a slice uses the same prediction mode. In response to a determination that an adjacent block is in the same slice as the current block, the processing circuit inserts prediction information from the adjacent block into the predictor list of the first prediction mode.
[0026] In one embodiment, the processing circuit determines whether an adjacent block is in the same tile as the current block or in the same tile group. A tile is a region of a picture and is processed in parallel and independently. A tile group is a group of tiles, and the same header can be shared among the groups of tiles. In response to a determination that an adjacent block is in the same tile or the same tile group as the current block, the processing circuit inserts prediction information from the adjacent block into the predictor list of the first prediction mode.
[0027] In one embodiment, the processing circuit determines whether an adjacent block overlaps with the current block. In response to a determination that an adjacent block does not overlap with the current block, the processing circuit inserts prediction information from the adjacent block into the predictor list of the first prediction mode.
[0028] In one embodiment, the first prediction mode includes at least one of an intra-block copy prediction mode and an inter prediction mode.
[0029] In one embodiment, the prediction information from adjacent blocks includes at least one of a block vector and a motion vector. The block vector indicates an offset between an adjacent block and the current block and is used to predict the current block when the adjacent block is encoded in an intra-block copy prediction mode. The motion vector is used to predict the current block when the adjacent block is encoded in an inter prediction mode.
[0030] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform a method for video decoding.
[0031] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
Brief Description of the Drawings
[0032] [Figure 1A] It is a schematic diagram of an exemplary subset of intra prediction modes. [Figure 1B] It is a diagram of an exemplary intra prediction direction. [Figure 1C] It is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Figure 2] It is a schematic diagram of a simplified block diagram of a communication system according to one embodiment. [Figure 3] It is a schematic diagram of a simplified block diagram of a communication system according to one embodiment. [Figure 4] It is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 5] It is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 6] It is a block diagram of an encoder according to another embodiment. [Figure 7] It is a block diagram of a decoder according to another embodiment. [Figure 8] This is an illustrative diagram of an intrablock copy prediction mode according to one embodiment. [Figure 9A] This figure shows an exemplary update process in the effective search range of the intrablock copy prediction mode according to one embodiment. [Figure 9B] This figure shows an exemplary update process in the effective search range of the intrablock copy prediction mode according to one embodiment. [Figure 9C] This figure shows an exemplary update process in the effective search range of the intrablock copy prediction mode according to one embodiment. [Figure 9D] This figure shows an exemplary update process in the effective search range of the intrablock copy prediction mode according to one embodiment. [Figure 10] This is a flowchart outlining an exemplary process according to one embodiment of the present disclosure. [Figure 11] This is a schematic diagram of a computer system according to one embodiment. [Modes for carrying out the invention]
[0033] Figure 2 shows a simplified block diagram of a communication system (200) according to one embodiment of the present disclosure. The communication system (200) includes, for example, a plurality of terminal devices that can communicate with each other over a network (250). For example, the communication system (200) includes a first pair of terminal devices (210) and (220) interconnected over the network (250). In the example of Figure 2, the first pair of terminal devices (210) and (220) perform one-way transmission of data. For example, terminal device (210) can encode video data (e.g., a stream of video pictures captured by terminal device (210)) for transmission to another terminal device (220) over the network (250). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. Terminal device (220) can receive the encoded video data from the network (250), decode the encoded video data to recover the video pictures, and display the video pictures according to the recovered video data. One-way data transmission can be common in applications such as media serving.
[0034] In another example, the communication system (200) includes a second pair of terminal devices (230) and (240) that perform bidirectional transmission of encoded video data, for example, during a video conference. For bidirectional transmission of data, in one example, each terminal device of terminal devices (230) and (240) can encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other terminal devices of terminal devices (230) and (240) via the network (250). Each terminal device of terminal devices (230) and (240) can also receive encoded video data transmitted by the other terminal devices of terminal devices (230) and (240), decode the encoded video data to recover the video pictures, and display the video pictures on an accessible display device according to the recovered video data.
[0035] In the example in Figure 2, terminal devices (210), (220), (230), and (240) may be represented as a server, a personal computer, and a smartphone, but the principles of this disclosure are not limited thereto. Embodiments of this disclosure find applications using laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (250) represents any number of networks that transmit encoded video data between terminal devices (210), (220), (230), and (240), including, for example, wired and / or wireless communication networks. Communication network (250) can exchange data over circuit-switched and / or packet-switched channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of network (250) may not be important to the operation of this disclosure unless described below herein.
[0036] Figure 3 shows the arrangement of a video encoder and video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter can be equally applied to other video-enabled applications, such as video conferencing, digital television, and storing compressed video on digital media including CDs, DVDs, and memory sticks.
[0037] The streaming system may include a capture subsystem (313) which may include a video source (301), such as a digital camera, to create, for example, a stream (302) of uncompressed video pictures. In one example, the stream (302) of video pictures includes a sample captured by the digital camera. The stream (302) of video pictures, which is drawn with a thick line to highlight a higher data volume compared to encoded video data (304) (or encoded video bitstream), can be processed by an electronic device (320) which includes a video encoder (303) coupled to the video source (301). The video encoder (303) may include hardware, software, or a combination thereof, enabling or implementing aspects of the disclosed subject as will be described in more detail below. The encoded video data (304) (or encoded video bitstream (304)), which is drawn with a thin line to highlight a lower data volume compared to the stream (302) of video pictures, can be stored in a streaming server (305) for future use. One or more streaming client subsystems, such as client subsystems (306) and (308) in Figure 3, can access a streaming server (305) to retrieve copies (307) and (309) of encoded video data (304). Client subsystem (306) may include, for example, a video decoder (310) within an electronic device (330). The video decoder (310) decodes the incoming copy (307) of the encoded video data and creates an outgoing stream of a video picture (311) that can be rendered on a display (312) (e.g., a display screen) or another rendering device (not shown). In some streaming systems, the encoded video data (304), (307), and (309) (e.g., video bitstreams) can be encoded according to specific video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265.For example, a video coding standard under development is informally known as Multipurpose Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0038] It should be noted that electronic devices (320) and (330) may include other components (not shown). For example, electronic device (320) may include a video decoder (not shown), and electronic device (330) may also include a video encoder (not shown).
[0039] Figure 4 shows a block diagram of a video decoder (410) according to one embodiment of the present disclosure. The video decoder (410) may be included in an electronic device (430). The electronic device (430) may include a receiver (431) (e.g., a receiving circuit). The video decoder (410) can be used in place of the video decoder (310) in the example of Figure 3.
[0040] The receiver (431) can receive one or more encoded video sequences to be decoded by the video decoder (410), which in the same or different embodiments may be one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel (401), which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (431) may receive the encoded video data together with other data, e.g., encoded audio data and / or auxiliary data streams, which may be forwarded to their respective usage entities (not shown). The receiver (431) can isolate the encoded video sequences from other data. To counteract network jitter, a buffer memory (415) may be coupled between the receiver (431) and the entropy decoder / parser (420) (hereinafter, "Parser (420)"). In certain applications, the buffer memory (415) is part of the video decoder (410). In other configurations, it may be located outside the video decoder (410) (not shown). Also, outside the video decoder (410), there may be a buffer memory (not shown) to counteract network jitter, and further inside the video decoder (410), there may be another buffer memory (415) to accommodate playout timing. When the receiver (431) is receiving data from a store / forward device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (415) may not be necessary or may be small. For use in best-effort packet networks such as the Internet, the buffer memory (415) may be necessary, and can be relatively large, advantageously adaptively sized, and may be implemented at least partially in an operating system or similar element (not shown) outside the video decoder (410).
[0041] The video decoder (410) may include a parser (420) for reconstructing symbols (421) from the encoded video sequence. The categories of these symbols may include information used to manage the operation of the video decoder (410), and potentially information for controlling rendering devices, such as a rendering device (412) (e.g., a display screen) which, as shown in Figure 4, is not an integral part of the electronic device (430) but can be coupled to the electronic device (430). The rendering device control information may be in the form of supplemental extension information (SEI messages) or video usability information (VUI) parameter set fragments (not shown). The parser (420) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may follow video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, and arithmetic coding with or without context sensitivity. The parser(420) can extract from the encoded video sequence a set of at least one subgroup parameters for a subgroup of pixels in the video decoder, based on at least one parameter corresponding to a group. Subgroups can include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), and predictive units (PU). The parser(420) can also extract from the encoded video sequence information such as transform coefficients, quantization parameter values, and motion vectors.
[0042] The parser (420) can perform entropy decoding / parsing operations on the video sequence received from buffer memory (415) in order to create symbols (421).
[0043] The reconstruction of the symbol (421) may involve multiple different units, depending on the type of the encoded video picture or part thereof (e.g., inter and intra picture, inter and intra block), and other factors. Which units are involved and how can be controlled by subgroup control information parsed from the video sequence encoded by the parser (420). The flow of such subgroup control information between the parser (420) and the following multiple units is not illustrated for clarity.
[0044] Beyond the functional blocks already described, the video decoder (410) can be conceptually subdivided into several functional units, as described below. In actual embodiments operating under commercial constraints, many of these units can interact closely with each other and, at least partially, integrate with each other. However, for the sake of illustrating the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0045] The first unit is the scaler / inverse unit (451). The scaler / inverse unit (451) receives control information from the parser (420) as symbols (421), including the quantized transformation coefficients, as well as the transformation to be used, block size, quantization coefficients, and quantization scaling matrix. The scaler / inverse unit (451) can output a block containing sample values that can be input to the aggregator (455).
[0046] In some cases, the output samples of the scaler / inverse transform (451) may relate to intra-encoded blocks; that is, blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from previously reconstructed portions of the current picture. Such prediction information can be provided by the intra-picture prediction unit (452). In some cases, the intra-picture prediction unit (452) generates a block of the same size and shape as the block being reconstructed, using surrounding already reconstructed information fetched from the current picture buffer (458). The current picture buffer (458) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (455) may, sample by sample, add the prediction information generated by the intra-prediction unit (452) to the output sample information provided by the scaler / inverse transform unit (451).
[0047] In other cases, the output samples of the scaler / inverse unit (451) may relate to intercoded and potentially motion-compensated blocks. In such cases, the motion-compensated prediction unit (453) may access the reference picture memory (457) to fetch samples to be used for prediction. After motion-compensating the fetched samples according to the symbols (421) related to the blocks, these samples can be added by the aggregator (455) to the output of the scaler / inverse unit (451) (in this case, called residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (457) from which the motion-compensated prediction unit (453) fetches prediction samples can be controlled by motion vectors available to the motion-compensated prediction unit (453) in the form of symbols (421) which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory (457) when the exact motion vectors of subsamples are used, and motion vector prediction mechanisms, etc.
[0048] The output samples of the aggregator (455) may be subjected to various loop filtering techniques in the loop filter unit (456). Video compression techniques may include in-loop filtering techniques controlled by parameters contained in the encoded video sequence (also called the encoded video bitstream) and made available to the loop filter unit (456) as symbols (421) from the parser (420), but may also respond to metadata obtained during the decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and previously reconstructed and loop-filtered sample values.
[0049] The output of the loop filter unit (456) may be a sample stream that can be output to the rendering device (412) or stored in reference picture memory (457) for use in future interpicture prediction.
[0050] A particular encoded picture, once fully reconstructed, can be used as a reference picture for future predictions. For example, when the encoded picture corresponding to the current picture is fully reconstructed and that encoded picture is identified as a reference picture (e.g., by the parser (420)), the current picture buffer (458) becomes part of the reference picture memory (457), and a new current picture buffer can be reallocated before starting the reconstruction of the next encoded picture.
[0051] The video decoder (410) can perform decoding operations according to a specified video compression technique in a standard such as ITU-T Rec.H.265. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that the encoded video sequence conforms to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, a profile can select a particular tool from all the tools available in the video compression technique or standard as the only tool that can be used in that profile. Also required for compliance is that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further limited by the virtual reference decoder (HRD) specification and the HRD buffer management metadata notified in the encoded video sequence.
[0052] In one embodiment, the receiver (431) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (410) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, or forward error correction code.
[0053] Figure 5 shows a block diagram of a video encoder (503) according to one embodiment of the present disclosure. The video encoder (503) is contained within an electronic device (520). The electronic device (520) includes a transmitter (540) (e.g., a transmitting circuit). The video encoder (503) can be used in place of the video encoder (303) in the example of Figure 3.
[0054] The video encoder (503) can receive video samples from a video source (501) (not part of the electronic device (520) in the example in Figure 5) which can capture video images encoded by the video encoder (503). In another example, the video source (501) is part of the electronic device (520).
[0055] The video source (501) can provide a source video sequence encoded by the video encoder (503) in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (501) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (501) may be a camera that captures local image information as a video sequence. The video data can be provided as a series of separate pictures that give motion when viewed sequentially. The pictures themselves can be organized as a spatial array of pixels, and each pixel may contain one or more samples, depending on the sampling structure, color space, etc., at the time of use. Those skilled in the art will readily understand the relationship between pixels and samples. The following description focuses on samples.
[0056] According to one embodiment, the video encoder (503) can encode pictures of a source video sequence in real time or under any other time constraints required by the application, and compress them into an encoded video sequence (543). Forcing an appropriate encoding rate is one function of the controller (550). In some embodiments, the controller (550) controls and is functionally coupled to other functional units, as described below. For clarity, couplings are not shown. Parameters set by the controller (550) may include rate control-related parameters (such as picture skip, quantizer, lambda value for rate distortion optimization technique), picture size, picture group (GOP) layout, and maximum motion vector search range. The controller (550) may be configured to have other appropriate functions related to the video encoder (503) optimized for a particular system design.
[0057] In some embodiments, the video encoder (503) is configured to operate in an encoding loop. For an oversimplified explanation, in one example, the encoding loop may include a source coder (530) (responsible for creating symbols, such as a symbol stream, based on, for example, the input picture to be encoded and a reference picture) and a (local) decoder (533) embedded in the video encoder (503). The decoder (533) reconstructs the symbols to create sample data in a similar manner to how a (remote) decoder also creates (since the compression between the symbols and the encoded video bitstream is lossless in the video compression techniques considered in the disclosed subject). The reconstructed sample stream (sample data) is input to a reference picture memory (534). Since the decoding of the symbol stream yields bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference picture memory (534) are also bit-accurate between the local encoder and the remote encoder. In other words, the predictive part of the encoder "sees" the exact same sample values as the reference picture samples that the decoder "sees" when using the predictions during decoding. This fundamental principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained due to, for example, channel errors) is also used in some related technologies.
[0058] The operation of the “local” decoder (533) may be the same as that of a “remote” decoder, such as the video decoder (410) which has already been described in detail above in relation to Figure 4. However, also briefly referring to Figure 4, since symbols are available and the encoding / decoding of symbols to the encoded video sequence by the entropy coder (545) and parser (420) may be lossless, the entropy decoding portion of the video decoder (410), including the buffer memory (415) and parser (420), does not have to be fully implemented in the local decoder (533).
[0059] An observation that can be made at this point is that decoder techniques other than parsing / entropy decoding present in the decoder must also be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject focuses on decoder operation. The description of encoder techniques can be omitted as it is the inverse of the comprehensively described decoder techniques. More detailed explanations are necessary only in specific areas and are provided below.
[0060] During operation, in some examples, the source coder (530) may perform motion-compensated predictive coding, predictively coding the input picture by referencing one or more previously coded pictures from a video sequence designated as “reference pictures”. In this way, the coding engine (532) codes the difference between the pixel blocks of the input picture and the pixel blocks of the reference picture that may be selected as a predictive reference to the input picture.
[0061] The local video decoder (533) can decode the encoded video data of a picture that may be designated as a reference picture based on symbols created by the source coder (530). The operation of the encoding engine (532) may, advantageously, be a loss process. If the encoded video data can be decoded by a video decoder (not shown in Figure 5), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (533) can replicate the decoding process that may be performed by the video decoder on the reference picture and store the reconstructed reference picture in the reference picture cache (534). In this way, the video encoder (503) can locally store a copy of the reconstructed reference picture that has common content as the reconstructed reference picture acquired by the far-end video decoder (without transmission errors).
[0062] The predictor (535) can perform predictive searches of the encoding engine (532). That is, for a new picture to be encoded, the predictor (535) can search the reference picture memory (534) for sample data (as candidate reference pixel blocks) or specific metadata such as motion vectors and block shapes of reference pictures that could serve as appropriate predictive criteria for the new picture. The predictor (535) can operate on a sample block and pixel block basis to find appropriate predictive references. In some cases, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (534), as determined by the search results obtained by the predictor (535).
[0063] The controller (550) can manage the encoding operations of the source coder (530), including, for example, setting parameters and subgroup parameters used to encode video data.
[0064] The outputs of all the aforementioned functional units can be entropically coded by the entropy coder (545). The entropy coder (545) converts the symbols generated by the various functional units into coded video sequences by losslessly compressing the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.
[0065] The transmitter (540) can buffer the encoded video sequence created by the entropy coder (545) and prepare it for transmission over a communication channel (560), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (540) can merge the encoded video data from the video coder (503) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).
[0066] The controller (550) can manage the operation of the video encoder (503). During encoding, the controller (550) can assign a specific encoding picture type to each encoded picture, which can influence the encoding technique that may be applied to each picture. For example, a picture is often assigned as one of the following picture types:
[0067] An intra-picture (I-picture) may be one that can be encoded and decoded without using other pictures in the sequence as a source of prediction. Some video codecs consider various types of intra-pictures, such as independent decoder refresh ("IDR") pictures. Those skilled in the art are familiar with the variations of I-pictures and their respective uses and characteristics.
[0068] A predictive picture (P-picture) may be one that can be encoded and decoded using intra-prediction or inter-prediction, which uses up to one motion vector and reference index to predict the sample value of each block.
[0069] A bidirectional predictive picture (B-picture) may be one that can be encoded and decoded using intra-prediction or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0070] A source picture is typically divided spatially into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each), and each block can be encoded. Blocks can be predictively encoded by referencing other (already encoded) blocks, as determined by the encoding assignment applied to each picture in the block. For example, blocks in picture I can be non-predictively encoded, or they can be predictively encoded by referencing already encoded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks in picture P can be predictively encoded via spatial prediction or via temporal prediction referencing one previously encoded reference picture. Blocks in picture B can be predictively encoded via spatial prediction or via temporal prediction referencing one or two previously encoded reference pictures.
[0071] The video encoder (503) can perform encoding operations in accordance with a given video encoding technique or standard, such as ITU-T Rec.H.265. In these operations, the video encoder (503) can perform various compression operations, including predictive encoding operations that utilize temporal and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard being used.
[0072] In one embodiment, the transmitter (540) may transmit additional data along with the encoded video. The source coder (530) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR extension layers, redundant data in other forms such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.
[0073] Video can be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated as intra-prediction) utilizes the spatial correlation of a particular picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If the blocks of the current picture resemble the reference blocks of a previously encoded and still buffering video reference picture, then the blocks of the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture and may have a third dimension to identify the reference picture if multiple reference pictures are used.
[0074] In some embodiments, bidirectional prediction techniques can be used for interpicture prediction. According to the bidirectional prediction technique, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are decoded earlier than the current picture in the video (but may be past and future in display order, respectively). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. A block can be predicted by a combination of the first and second reference blocks.
[0075] Furthermore, merge mode techniques can be used for interpicture prediction to improve coding efficiency.
[0076] According to some embodiments of this disclosure, predictions such as interpicture prediction and intrapicture prediction are performed in units of blocks. For example, according to the HEVC standard, the pictures of a series of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs of the pictures are of the same size (e.g., 64x64 pixels, 32x32 pixels, 16x16 pixels). Generally, a CTU contains three coding tree blocks (CTBs). These are one luminance CTB and two saturation CTBs. Each CTU may be recursively quad-tree divided into one or more coding units (CUs). For example, a 64x64 pixel CTU may be divided into one 64x64 pixel CU, four 32x32 pixel CUs, or sixteen 16x16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU (e.g., inter-prediction type or intra-prediction type). The CU is divided into one or more prediction units (PUs) depending on its temporal and / or spatial predictability. Generally, each PU includes a luminance prediction block (PB) and two saturation PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luminance prediction block as an example of a prediction block, the prediction block contains a matrix of values (e.g., luminance values) for pixels such as 8x8 pixels, 16x16 pixels, 8x16 pixels, and 16x8 pixels.
[0077] Figure 6 shows a video encoder (603) according to another embodiment of the present disclosure. The video encoder (603) is configured to receive a processing block (e.g., a prediction block) of sample values in the current video picture within a series of video pictures, and to encode the processing block into an encoded picture which is part of an encoded video sequence. In one example, the video encoder (603) is used instead of the video encoder (303) in the example of Figure 3.
[0078] In the HEVC example, the video encoder (603) receives a matrix of sample values for a processing block, such as an 8x8 sample prediction block. The video encoder (603) determines whether the processing block is best encoded using intra-mode, inter-mode, or bidirectional prediction mode, for example, using rate-distortion optimization. If the processing block is encoded in intra-mode, the video encoder (603) can encode the processing block into an encoded picture using intra-prediction techniques; if the processing block is encoded in inter-mode or bidirectional prediction mode, the video encoder (603) can encode the processing block into an encoded picture using inter-prediction or bidirectional prediction techniques, respectively. In certain video encoding techniques, merge mode is an interpicture prediction submode in which the motion vector is derived from one or more motion vector predictors without the benefit of an encoded motion vector component outside the predictor. In certain other video encoding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (603) includes other components, such as a mode determination module (not shown) for determining the mode of the processing block.
[0079] In the example in Figure 6, the video encoder (603) includes an interencoder (630), an intraencoder (622), a residual calculator (623), a switch (626), a residual encoder (624), a general-purpose controller (621), and an entropy encoder (625), all coupled together as shown in Figure 6.
[0080] The interencoder (630) is configured to receive a sample of the current block (e.g., a processing block), compare the block to one or more reference blocks in the reference picture (e.g., blocks from the previous and subsequent pictures), generate interprediction information (e.g., a description of redundant information by the intercoding technique, motion vectors, merge mode information), and compute an interprediction result (e.g., a predicted block) based on the interprediction information using any appropriate technique. In some examples, the reference picture is a decoded reference picture that is decoded based on encoded video information.
[0081] The intra encoder (622) is configured to receive a sample of the current block (e.g., a processing block), and optionally compare the block to a block already encoded in the same picture, generate quantized coefficients after transformation, and optionally also generate intra prediction information (e.g., intra prediction direction information using one or more intra coding techniques). In one example, the intra encoder (622) also computes an intra prediction result (e.g., a prediction block) based on the intra prediction information and reference block in the same picture.
[0082] The general-purpose controller (621) is configured to determine general-purpose control data and to control other components of the video encoder (603) based on the general-purpose control data. For example, the general-purpose controller (621) determines the mode of a block and provides control signals to the switch (626) based on the mode. For example, if the mode is intra-mode, the general-purpose controller (621) controls the switch (626) to select intra-mode results for use by the residual calculator (623), controls the entropy encoder (625) to select intra-prediction information, and includes the intra-prediction information in the bitstream. If the mode is inter-mode, the general-purpose controller (621) controls the switch (626) to select inter-prediction results for use by the residual calculator (623), controls the entropy encoder (625) to select inter-prediction information, and includes the inter-prediction information in the bitstream.
[0083] The residual calculator (623) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra-encoder (622) or inter-encoder (630). The residual encoder (624) operates on the residual data and is configured to encode the residual data to generate conversion coefficients. In one example, the residual encoder (624) is configured to convert the residual data from the spatial domain to the frequency domain and generate conversion coefficients. The conversion coefficients are then subjected to a quantization process to obtain quantized conversion coefficients. In various embodiments, the video encoder (603) also includes a residual decoder (628). The residual decoder (628) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra-encoder (622) and inter-encoder (630). For example, an interencoder (630) can generate a decoded block based on decoded residual data and interprediction information, and an intraencoder (622) can generate a decoded block based on decoded residual data and intraprediction information. The decoded block is appropriately processed to generate a decoded picture, which is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.
[0084] The entropy encoder (625) is configured to format the bitstream to include the encoded blocks. The entropy encoder (625) is configured to include various information according to an appropriate standard such as the HEVC standard. For example, the entropy encoder (625) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Note that, according to the disclosed subject, when encoding blocks in either inter-mode or bidirectional prediction mode merge submode, there is no residual information.
[0085] Figure 7 shows a video decoder (710) according to another embodiment of the present disclosure. The video decoder (710) is configured to receive an encoded picture which is part of an encoded video sequence, and to decode the encoded picture to produce a reconstructed picture. In one example, the video decoder (710) is used instead of the video decoder (310) in the example of Figure 3.
[0086] In the example shown in Figure 7, the video decoder (710) includes an entropy decoder (771), an interdecoder (780), a residual decoder (773), a reconstruction module (774), and an intradecoder (772), which are coupled together as shown in Figure 7.
[0087] The entropy decoder (771) may be configured to reconstruct specific symbols from the encoded picture that represent the syntactic elements constituting the encoded picture. Such symbols may include, for example, the mode in which the block is encoded (e.g., intra-mode, inter-mode, bidirectional prediction mode, the latter two being merge sub-mode or another sub-mode), prediction information that can identify specific samples or metadata used for prediction by the intra-decoder (772) or inter-decoder (780), respectively (e.g., intra-prediction information or inter-prediction information), and residual information in the form of, for example, quantized transformation coefficients. In one example, if the prediction mode is inter-prediction mode or bidirectional prediction mode, inter-prediction information is provided to the inter-decoder (780), and if the prediction type is intra-prediction type, intra-prediction information is provided to the intra-decoder (772). The residual information may be subject to inverse quantization and provided to the residual decoder (773).
[0088] The interdecoder (780) is configured to receive interprediction information and generate interprediction results based on the interprediction information.
[0089] The intra decoder (772) is configured to receive intra prediction information and generate prediction results based on the intra prediction information.
[0090] The residual decoder (773) is configured to perform inverse quantization to extract the dequantized conversion coefficients and process the dequantized conversion coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (773) may also require certain control information (including quantization parameters (QP)), which may be provided by the entropy decoder (771) (the data path is not shown because this may only require a small amount of control information).
[0091] The reconstruction module (774) is configured to combine the residuals output by the residual decoder (773) and the prediction results (which may be output by the inter-prediction module or intra-prediction module) in the spatial domain to form a reconstructed block, which may be part of a reconstructed picture, and by extension, part of a reconstructed video. Note that other appropriate operations, such as deblocking operations, can be performed to improve visual quality.
[0092] It should be noted that the video encoders (303), (503), and (603), and the video decoders (310), (410), and (710) can be implemented using any suitable technique. In one embodiment, the video encoders (303), (503), and (603), and the video decoders (310), (410), and (710) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (303), (503), and (503), and the video decoders (310), (410), and (710) can be implemented using one or more processors that execute software instructions.
[0093] Generally, block-based compensation is based on different pictures. Such block-based compensation is sometimes called motion compensation. However, block compensation can be performed from a previously reconstructed region within the same picture. Such block compensation may be called intra-picture block compensation, current picture reference (CPR), or intra-block copy (IBC).
[0094] In IBC prediction mode, in some embodiments, the displacement vector indicating the offset between the current block and the reference block within the same picture is called the block vector (BV). Note that the reference block has already been reconstructed before the current block. Furthermore, in parallel processing, reference regions located at tile / slice boundaries or wavefront ladder shape boundaries may be excluded from use as available reference blocks. Due to these constraints, the block vector may differ from the motion vector, which can take any value (positive or negative in either the x or y direction) in motion compensation.
[0095] Figure 8 shows an embodiment of the intrablock copy prediction mode according to one embodiment of the present disclosure. In Figure 8, gray blocks indicate that a block has already been decoded, and white blocks indicate that a block has not been decoded or is in the process of being decoded. Thus, in the current picture (800), the block vector (802) points from the current block (801) to the reference block (803). The current block (801) is being reconstructed, and the reference block (803) has already been reconstructed.
[0096] According to some embodiments, the encoding of block vectors can be either explicit or implicit. In explicit mode, the difference between the block vector and its predictor is communicated. In implicit mode, the block vector is reconstructed from its predictor in a manner similar to motion vector prediction in merge mode. The resolution of the block vector may be limited to integer positions in one embodiment, but may point to fractional positions in another embodiment.
[0097] According to some embodiments, the use of block-level IBC prediction mode can be indicated using either a block-level flag (called an IBC flag) or a reference index. When using the reference index method, the picture currently being decoded is treated as the reference picture that is placed at the last position in the reference picture list. This reference picture can also be managed together with other time reference pictures in the decoded picture buffer (DPB).
[0098] Furthermore, there are several variations of IBC prediction modes. In some examples, the reference block is flipped horizontally or vertically before being used to predict the current block, which is sometimes called a flipped IBC prediction mode. In other examples, each compensation unit within an MxN coded block is an Mx1 or 1xN line, which is sometimes called a line-based IBC prediction mode.
[0099] In current VVCs, the search range for IBC prediction mode is limited to the current CTU. In some embodiments, the memory for storing reference samples in IBC prediction mode is 1 CTU size (e.g., four 64x64 regions). For example, the memory stores four 64x64 regions of samples, one of which may be the currently reconstructed sample, and the other three 64x64 regions may be reference samples.
[0100] In some embodiments, the effective search range of the IBC prediction mode can be extended to several parts of the CTU to the left of the current CTU without modifying the memory (i.e., 1 CTU size for a 64x64 area). For example, an update process can be used within such an effective search range.
[0101] Figures 9A to 9D illustrate embodiments of an update process using an effective search range (i.e., intra-picture block compensation) of the IBC prediction mode according to embodiments of the present disclosure. The update process can be performed on a 64x64 luminance sample basis for each of the four 64x64 block regions in the CTU-sized memory, and reference samples from the same region from the left CTU can be used to predict the encoded blocks of the current CTU until any of the blocks in the same region of the current CTU are being encoded or have been encoded.
[0102] During this process, in some embodiments, the stored reference sample from the left CTU is updated with the reconstructed sample from the current CTU. In Figures 9A to 9D, gray blocks indicate blocks that have already been reconstructed, white blocks indicate blocks that have not been reconstructed, and blocks with vertical stripes and the text "Curr" indicate the current encoded / decoded block. Also, in each figure, the left four blocks (911) to (914) belong to the left CTU (910), and the right four blocks (901) to (904) belong to the current CTU (900).
[0103] Note that all four blocks (911) to (914) of the left CTU(910) have already been rebuilt. Therefore, memory first stores all four of these blocks of reference samples from the left CTU(910), and then updates the blocks of reference samples from the left CTU(910) with the current blocks in the same area from the current CTU(900).
[0104] For example, in Figure 9A, the current block (901) in the current CTU(900) is being rebuilt, and the block in the same location to the left of the current block (901) in CTU(910) is block (911). Block (911) is in the same location in CTU(910) as the area in CTU(900) where the current block (901) is located. Thus, the memory area that stores the reference sample of block (911) is updated to store the rebuilt sample of the current block (901), and an "X" is marked on block (911) in Figure 9A to indicate that the reference sample of block (911) is no longer stored in memory.
[0105] Similarly, in Figure 9B, the current block (902) in the current CTU(900) is being rebuilt, and the block in the same location to the left of the current block (902) in CTU(910) is block (912). Block (912) is in the same location in CTU(910) to the left of the area where the current block (902) is located in the current CTU(900). Thus, the memory area that stores the reference sample of block (912) is updated to store the rebuilt sample of the current block (902), and a "X" is marked on block (912) in Figure 9B to indicate that the reference sample of block (912) is no longer stored in memory.
[0106] In Figure 9C, the current block (903) in the current CTU(900) is being rebuilt, and the block in the same location to the left of the current block (903) in CTU(910) is block (913). Block (913) is in the same location in CTU(910) to the left of the area where the current block (903) is located in the current CTU(900). Thus, the memory area that stores the reference sample of block (913) is updated to store the rebuilt sample of the current block (903), and an "X" is marked on block (913) in Figure 9C to indicate that the reference sample of block (913) is no longer stored in memory.
[0107] In Figure 9D, the current block (904) in the current CTU(900) is being rebuilt, and the block in the same location to the left of the current block (904) in CTU(910) is block (914). Block (914) is in the same location in CTU(910) to the left of the area where the current block (904) is located in the current CTU(900). Thus, the memory area that stores the reference sample of block (914) is updated to store the rebuilt sample of the current block (904), and an "X" is marked on block (914) in Figure 9C to indicate that the reference sample of block (914) is no longer stored in memory.
[0108] According to some embodiments, in intra-prediction mode, if an adjacent block exists and has been reconstructed before the current coded block, the adjacent block can be used as the predictor for the current coded block. In some embodiments, in inter-prediction mode, the adjacent block can be used as the predictor for the current coded block if the adjacent block has not been coded in intra-prediction mode, except that the adjacent block exists and has already been reconstructed before the current coded block.
[0109] However, if an intra-block copy prediction mode is included and is considered a separate mode from the intra-prediction mode or inter-prediction mode, the availability of adjacent blocks becomes more complex, requiring the development of a suitable method for checking the availability of adjacent blocks. In this regard, embodiments of the present disclosure present a method for efficiently checking the availability of adjacent blocks for the current encoded block.
[0110] When the block vector predictor list for IBC prediction mode (whether a merge list or block vector differential coding) is constructed separately from the motion vector predictor list (whether a merge list or motion vector differential coding), in some embodiments, a uniform adjacent availability check condition is used to determine whether the relevant information of adjacent blocks can be used as predictors for the current coded block.
[0111] According to some embodiments, if adjacent blocks are encoded in the same prediction mode as the current block, the adjacent blocks can be used to predict the current block.
[0112] In one embodiment, if the current block is encoded in IBC prediction mode, the current block can be predicted using adjacent blocks encoded in IBC prediction mode. If it is determined that adjacent blocks are available (as further specified below), prediction information from the adjacent blocks (such as block vectors) is added to the predictor list in IBC prediction mode.
[0113] In another embodiment, if the current block is encoded in interprediction mode, the current block can be predicted using adjacent blocks encoded in interprediction mode. When it is determined that adjacent blocks are available (as further specified below), prediction information from the adjacent blocks (such as motion vectors, reference indexes, and prediction directions) is added to the interprediction mode predictor list.
[0114] In another embodiment, if the current block is encoded in intra-predictive mode, the current block can be predicted using adjacent blocks encoded in intra-predictive mode. When it is determined that adjacent blocks are available (as further specified below), prediction information from adjacent blocks is added to the intra-predictive predictor list.
[0115] In some embodiments, determining the availability of an adjacent block can be divided into two steps. The first step includes checking the decoding order. In one example, if an adjacent block is in a different slice, tile, or tile group compared to the current block, it is determined that the adjacent block is unavailable to predict the current block. A slice may be a group of blocks in raster scan order, and groups of blocks within a slice may use the same prediction mode. Tiles may be regions of a picture and may be processed independently in parallel. A tile group may be a group of tiles and may share the same header among groups of tiles. In another example, if an adjacent block is not reconstructed before the current block, it is determined that the adjacent block is unavailable to predict the current block.
[0116] After the first step is completed (for example, after it has been determined that the adjacent block is available in terms of the decoding order), the second step is applied to the adjacent block to check its predictive availability. For example, if the adjacent block overlaps with the current block (for example, if the adjacent block is not fully constructed before the current block, or if at least one sample of the adjacent block is not constructed before the current block), the adjacent block is determined to be unavailable for predicting the current block. In another example, if the adjacent block is encoded in a different predictive mode than the current block, the adjacent block is determined to be unavailable for predicting the current block.
[0117] According to some embodiments, if an adjacent block is encoded in IBC prediction mode and the current block is encoded in interprediction mode, the adjacent block can be used to predict the current block.
[0118] According to some embodiments, if adjacent blocks are encoded in interprediction mode and the current block is encoded in IBC prediction mode, the adjacent blocks can be used to predict the current block.
[0119] In both of the above decisions, according to some embodiments, an adjacent block is checked to see if it is encoded in intra-predictive mode. If an adjacent block is encoded in intra-predictive mode, it is determined that the adjacent block is unavailable to predict the current block. In addition, several other conditions of the adjacent block are also checked (e.g., the decoding order of the adjacent block). In one example, if an adjacent block is in a different slice, tile, or tile group compared to the current block, it is determined that the adjacent block is unavailable to predict the current block. A slice may be a group of blocks in raster scan order, and groups of blocks within a slice may use the same prediction mode. A tile may be a region of a picture and may be processed independently in parallel. A tile group may be a group of tiles and may share the same header among groups of tiles. In another example, if an adjacent block is not reconstructed before the current block, it is determined that the adjacent block is unavailable to predict the current block. In yet another example, if an adjacent block overlaps with the current block, it is determined that the adjacent block is unavailable to predict the current block. In another example, if an adjacent block is encoded in a different prediction mode than the current block, it is determined that the adjacent block cannot be used to predict the current block.
[0120] In addition to a method for checking the availability of adjacent blocks as predictors for the current block, embodiments of the present disclosure include a method for determining the boundary intensity of a deblocking filter used to reconstruct the current block in IBC prediction mode. While the current block is being reconstructed, a deblocking filter can be used to improve visual quality and prediction performance by smoothing sharp edges at the boundaries between decoded blocks, and the boundary intensity is used to indicate the intensity level used by the deblocking filter. In one embodiment, an intensity level of "0" may indicate that the deblocking filter is not performed, and an intensity level of "1" may indicate that the deblocking filter is performed. In another embodiment, three or more intensity levels, for example, a low intensity level, a medium intensity level, and a high intensity level, may be used. By performing the deblocking filter at different intensity levels, the current block may have different visual qualities.
[0121] In IBC prediction mode, the luminance and chroma components are encoded in separate coding tree structures. The chroma CU can operate in subblock mode. For example, subblocks of the chroma CU can be derived from luminance positions at the same location in the subblock, and thus different subblocks can have different block vectors. Accordingly, according to embodiments of the present disclosure, the boundary strength of the deblocking filter when the deblocking filter is applied to the chroma CU can be determined using several methods.
[0122] According to some embodiments, the deblocking filter is performed at the subblock boundary of the saturation subblock, and the boundary strength (BS) at the subblock boundary is determined by evaluating the difference between the respective saturation block vectors of the two encoded subblocks associated with the subblock boundary. In one embodiment, the deblocking filter is performed if the absolute value of the difference between the horizontal components (or vertical components) of the respective block vectors is greater than or equal to 1 in integer luminance samples (or, in other words, 4 in 1 / 4 luminance samples). Otherwise, the deblocking filter is not performed.
[0123] According to some embodiments, the deblocking filter is performed at the subblock boundaries of the saturation CU or at the 4x4 block boundaries, whichever is greater. In one embodiment, when the subblock size of the saturation CU is 2x2, the deblocking filter is not performed at some 2x2 subblock boundaries, but instead at each of the 4x4 grid boundaries. For these 4x4 boundaries, the boundary strength (BS) at the boundary can be determined by evaluating the difference between the respective block vectors of the two encoded subblocks associated with the subblock boundary. If the absolute difference between the horizontal (or vertical) components of each block vector is 1 or greater in units of integer luminance samples (or another expression with the same meaning: 4 in units of 1 / 4 luminance samples), the deblocking filter is performed. Otherwise, the deblocking filter is not performed.
[0124] Figure 10 shows a flowchart outlining an exemplary process (1000) according to one embodiment of the present disclosure. Process (1000) can be used to reconstruct blocks encoded in predictive mode (inter / intra / IBC, etc.) to generate predictive blocks of the blocks being reconstructed. In various embodiments, process (1000) is performed by processing circuits such as processing circuits for terminal devices (210), (220), (230) and (240), processing circuits performing the functions of a video encoder (303), processing circuits performing the functions of a video decoder (310), processing circuits performing the functions of a video decoder (410), processing circuits performing the functions of an intra predictive module (452), processing circuits performing the functions of a video encoder (503), processing circuits performing the functions of a predictor (535), processing circuits performing the functions of an intra encoder (622), and processing circuits performing the functions of an intra decoder (772). In some embodiments, process (1000) is implemented by software instructions, and therefore, when the processing circuit executes the software instructions, the processing circuit executes process (1000).
[0125] Process (1000) can generally be started in step (S1010), in which process (1000) decodes prediction information for the current block in the current encoded picture, which is part of the encoded video sequence. The prediction information indicates that a first prediction mode is being used for the current block. The first prediction mode may be one of intra, inter, or IBC prediction modes.
[0126] Process (1000) proceeds to step (S1020), where process (1000) determines whether adjacent blocks adjacent to the current block and rebuilt before the current block use the first prediction mode.
[0127] Process (1000) proceeds to step (S1030), in which process (1000) inserts prediction information from the adjacent block into the predictor list of the first prediction mode in response to the decision that the adjacent block will use the first prediction mode.
[0128] Process (1000) proceeds to step (S1040), where process (1000) reconstructs the current block according to the predictor list of the first prediction mode.
[0129] After rebuilding the current block, process (1000) will terminate.
[0130] The above techniques may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media. For example, Figure 11 shows a computer system (1100) suitable for carrying out a particular embodiment of the disclosed subject matter.
[0131] Computer software can be coded using any suitable machine code or computer language, which is subject to mechanisms such as assembly, compilation, and linking, and can create code containing instructions that can be executed directly or through interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.
[0132] Instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, and Internet of Things devices.
[0133] The components shown in Figure 11 for the computer system (1100) are essentially illustrative and are not intended to imply any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. Furthermore, the configuration of the components should not be construed as having any dependencies or requirements relating to any one or combination of components shown in the exemplary embodiments of the computer system (1100).
[0134] The computer system (1100) may include certain human interface input devices. Such human interface input devices may respond to input from one or more human users through, for example, haptic input (keystrokes, swipes, data glove movements, etc.), audio input (voice, applause, etc.), visual input (gestures, etc.), and olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as voice (speech, music, ambient sounds, etc.), images (scanned images, still images taken with a camera, etc.), and video (2D video, 3D video including stereoscopic video, etc.).
[0135] Input human interface devices may include one or more of the following (only one of each is depicted): keyboard (1101), mouse (1102), trackpad (1103), touchscreen (1110), data glove (not shown), joystick (1105), microphone (1106), scanner (1107), and camera (1108).
[0136] The computer system (1100) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (1110), data glove (not shown), or joystick (1105), although there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers (1109), headphones (not shown)), visual output devices (e.g., screens (1110) including CRT screens, LCD screens, plasma screens, OLED screens, etc., some of which may be capable of outputting three or more dimensions by means such as two-dimensional visual output or stereoscopic output, with or without touchscreen input functionality and with or without tactile feedback functionality, virtual reality glasses (not shown), holographic displays, smoke tanks (not shown)), and printers (not shown).
[0137] The computer system (1100) may also include human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW (1120) with media such as CD / DVD (1121), thumb drives (1122), removable hard drives or solid-state drives (1123), legacy magnetic media such as tapes and floppy disks (not shown), and special ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0138] Those skilled in the art should also understand that the term “computer-readable medium” as used in relation to the subject matter currently disclosed does not include a transmission medium, carrier wave, or other transient signal.
[0139] The computer system (1100) may also include interfaces to one or more communication networks. These networks may be, for example, wireless, wired, or optical. Networks may further be local, wide-area, metropolitan, vehicle and industrial, real-time, or latency-tolerant. Examples of networks include local area networks such as Ethernet, cellular networks including Wi-Fi, GSM, 3G, 4G, 5G, LTE, etc., wide-area digital networks for wired or wireless TV including cable TV, satellite TV, and terrestrial broadcast TV, and vehicle and industrial networks including CANBus. Certain networks typically require an external network interface adapter connected to a specific general-purpose data port or peripheral bus (1149) (e.g., the computer system's USB port (1100)), while others are generally integrated into the core of the computer system (1100) by connecting to a system bus, as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1100) may communicate with other entities. Such communications can be unidirectional, receive only (e.g., television broadcasting), transmit only in one direction (e.g., from CANbus to a specific CANbus device), or bidirectional, for example, to other computer systems using a local or wide-area digital network. As described above, specific protocols and protocol stacks can be used for each of these networks and network interfaces.
[0140] The aforementioned human interface device, human-accessible storage device, and network interface can be mounted on the core (1140) of the computer system (1100).
[0141] A core (1140) may include one or more central processing units (CPUs) (1141), graphics processing units (GPUs) (1142), specialized programmable processing units in the form of field-programmable gate areas (FPGAs) (1143), and hardware accelerators for specific tasks (1144). These devices may be connected via a system bus (1148) along with read-only memory (ROM) (1145), random access memory (1146), internal mass storage devices such as internal hard drives and SSDs (1147) that are not accessible to the user. In some computer systems, the system bus (1148) can be accessed in the form of one or more physical plugs to enable expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (1148) or via a peripheral bus (1149). Peripheral bus architectures include PCI, USB, etc.
[0142] The CPU (1141), GPU (1142), FPGA (1143), and accelerator (1144) can execute certain instructions that, when combined, can constitute the aforementioned computer code. This computer code can be stored in ROM (1145) or RAM (1146). Transition data can also be stored in RAM (1146), while persistent data can be stored, for example, in internal mass storage (1147). High-speed storage and retrieval to any of the memory devices can be made possible by using cache memory closely associated with one or more CPUs (1141), GPUs (1142), mass storage (1147), ROM (1145), RAM (1146), etc.
[0143] Computer-readable media may have computer code on them for performing operations performed on various computers. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a kind that is well known and available to persons skilled in computer software technology.
[0144] As an example, but not limited to, a computer system having architecture (1100), specifically a core (1140), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be user-accessible mass storage devices as described above, as well as media related to specific storage devices of the core (1140) of a non-transient nature, such as core internal mass storage (1147) or ROM (1145). Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core (1140). The computer-readable media may include one or more memory devices or chips, depending on the specific needs. The software may cause the core (1140) and specifically the processor (including a CPU, GPU, FPGA, etc.) within it to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (1146) and modifying such data structures according to processes defined by the software. In addition, or as an alternative, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in a circuit (e.g., an accelerator (1144)), which may operate in place of or with software to perform a particular process or a particular part of a particular process described herein. References to software may include logic as needed, and vice versa. References to computer-readable media may include, as needed, a circuit (such as an integrated circuit (IC)) that stores software for execution, a circuit that embodies logic for execution, or both. This disclosure encompasses any suitable combination of hardware and software.
[0145] While this disclosure describes several exemplary embodiments, there are many variations, substitutions, and alternative equivalents that fall within the scope of this disclosure. Those skilled in the art will therefore understand that numerous systems and methods not expressly shown or described herein can be devised that embody the principles of the disclosure and thus fall within its spirit and scope.
[0146] (1) A method for video decoding in a decoder includes the steps of: decoding prediction information for a current block in a current encoded picture which is part of an encoded video sequence, wherein the prediction information indicates a first prediction mode to be used for the current block; determining whether an adjacent block adjacent to the current block and reconstructed before the current block uses the first prediction mode; inserting the prediction information from the adjacent block into a predictor list for the first prediction mode in response to the determination that the adjacent block uses the first prediction mode; and reconstructing the current block according to the predictor list for the first prediction mode.
[0147] (2) The method of feature (1) further includes the steps of: determining whether the second prediction mode is an intra-prediction mode in response to a decision that an adjacent block uses a second prediction mode different from the first prediction mode; and inserting prediction information from the adjacent block into the predictor list of the first prediction mode in response to a decision that the second prediction mode is not an intra-prediction mode.
[0148] (3) The method of feature (1) further includes the steps of determining whether an adjacent block is in the same slice as the current block, where the slice is a group of blocks in raster scan order and the group of blocks in the slice uses the same prediction mode, and in response to the determination that the adjacent block is in the same slice as the current block, inserting prediction information from the adjacent block into the predictor list of the first prediction mode.
[0149] (4) The method of feature (1) includes the steps of determining whether an adjacent block is on the same tile or in the same tile group as the current block, where a tile is a region of a picture and is processed in parallel and independently, and a tile group is a group of tiles and shares the same header among tile groups; and in response to the determination that the adjacent block is on the same tile or in the same tile group as the current block, inserting prediction information from the adjacent block into the predictor list of the first prediction mode.
[0150] (5) The method of feature (1) further includes the steps of determining whether an adjacent block overlaps with the current block, and, in response to the determination that the adjacent block does not overlap with the current block, inserting prediction information from the adjacent block into the predictor list of the first prediction mode.
[0151] (6) The method of feature (1), wherein the first prediction mode includes at least one of an intrablock copy prediction mode and an interpretation mode.
[0152] (7) The method of feature (6), wherein the prediction information from an adjacent block includes at least one of a block vector and a motion vector, the block vector indicating an offset between the adjacent block and the current block and used to predict the current block when the adjacent block is encoded in intra-block copy prediction mode, and the motion vector used to predict the current block when the adjacent block is encoded in inter-prediction mode.
[0153] (8) An apparatus including a processing circuit, the processing circuit is configured to decode prediction information for the current block in the current encoded picture which is part of an encoded video sequence, the prediction information indicates a first prediction mode to be used for the current block, determine whether an adjacent block adjacent to the current block and reconstructed before the current block uses the first prediction mode, and in response to the determination that the adjacent block uses the first prediction mode, insert the prediction information from the adjacent block into the predictor list of the first prediction mode, and reconstruct the current block according to the predictor list of the first prediction mode.
[0154] (9) The apparatus of feature (8), wherein the processing circuit is further configured to determine whether the second prediction mode is an intra-prediction mode in response to a decision that an adjacent block uses a second prediction mode different from the first prediction mode, and to insert prediction information from the adjacent block into the predictor list of the first prediction mode in response to a decision that the second prediction mode is not an intra-prediction mode.
[0155] (10) The apparatus of feature (8), wherein the processing circuit is configured to determine whether an adjacent block is in the same slice as the current block, where a slice is a group of blocks in raster scan order, and the group of blocks in a slice uses the same prediction mode, and in response to the determination that an adjacent block is in the same slice as the current block, it inserts prediction information from the adjacent block into a predictor list of the first prediction mode.
[0156] (11) The apparatus of feature (8), wherein the processing circuit determines whether an adjacent block is on the same tile or in the same tile group as the current block, the tile being a region of a picture and processed independently in parallel, the tile group being a group of tiles and sharing the same header among tile groups, and in response to the determination that the adjacent block is on the same tile or in the same tile group as the current block, the circuit is further configured to insert predictive information from the adjacent block into a predictor list of a first predictive mode.
[0157] (12) The apparatus of feature (8), wherein the processing circuit is further configured to determine whether an adjacent block overlaps with the current block, and in response to the determination that the adjacent block does not overlap with the current block, to insert prediction information from the adjacent block into a predictor list of a first prediction mode.
[0158] (13) The apparatus of feature (8), wherein the first prediction mode includes at least one of an intrablock copy prediction mode and an interpretation mode.
[0159] (14) The apparatus of feature (13), wherein prediction information from an adjacent block includes at least one of a block vector and a motion vector, the block vector indicating an offset between the adjacent block and the current block and used to predict the current block when the adjacent block is encoded in intra-block copy prediction mode, and the motion vector used to predict the current block when the adjacent block is encoded in inter-prediction mode.
[0160] (15) A non-temporary computer-readable storage medium storing a program executable by at least one processor, the program comprising: decoding predictive information for a current block in a current encoded picture which is part of an encoded video sequence, wherein the predictive information indicates a first predictive mode to be used for the current block; determining whether an adjacent block adjacent to the current block and reconstructed before the current block uses the first predictive mode; inserting predictive information from the adjacent block into a predictor list for the first predictive mode in response to the determination that the adjacent block uses the first predictive mode; and reconstructing the current block according to the predictor list for the first predictive mode.
[0161] (16) A non-temporary computer-readable storage medium of feature (15), wherein the program to be stored further performs the steps of: determining whether the second prediction mode is an intra-prediction mode in response to a determination that an adjacent block uses a second prediction mode different from a first prediction mode; and inserting prediction information from the adjacent block into the predictor list of the first prediction mode in response to a determination that the second prediction mode is not an intra-prediction mode.
[0162] (17) A non-temporary computer-readable storage medium of feature (15), wherein the program to be stored comprises the steps of determining whether an adjacent block is in the same slice as the current block, where the slice is a group of blocks in raster scan order and the group of blocks in the slice uses the same prediction mode, and in response to the determination that the adjacent block is in the same slice as the current block, inserting prediction information from the adjacent block into a predictor list of the first prediction mode.
[0163] (18) A non-temporary computer-readable storage medium of feature (15), wherein the program to be stored comprises the steps of determining whether an adjacent block is on the same tile or in the same tile group as the current block, where a tile is a picture area and is processed concurrently and independently, and a tile group is a group of tiles and shares the same header, and in response to the determination that the adjacent block is on the same tile or in the same tile group as the current block, inserting predictive information from the adjacent block into a predictor list of a first predictive mode.
[0164] (19) A non-temporary computer-readable storage medium of feature (15), wherein the program to be stored further performs the steps of determining whether an adjacent block overlaps with the current block, and, in response to the determination that the adjacent block does not overlap with the current block, inserting predictive information from the adjacent block into a predictor list of a first predictive mode.
[0165] (20) A non-temporary computer-readable storage medium of feature (15), wherein the first prediction mode includes at least one of an intrablock copy prediction mode and an interpretation mode. Appendix A: Abbreviations AMVP: Advanced Motion Vector Prediction ASIC: Application-Specific Integrated Circuit BMS: Benchmark Set BV: Block Vector CANBus: Controller Area Network Bus CD: Compact Disc CPR: Current Picture Reference CPU: Central Processing Unit CRT: Cathode Ray Tube CTB: Encoded Tree Block CTU: Encoding Tree Unit CU: Encoding Unit DPB: Decoder Picture Buffer DVD: Digital Video Disc FPGA: Field-Programmable Gate Area GOP: Picture Group GPU: Graphics Processing Unit GSM: Global System for Mobile Communications HEVC: High Efficiency Video Coding HRD: Virtual Reference Decoder IBC: Intrablock Copy IC: Integrated Circuit JEM: Collaborative Search Model LAN: Local Area Network LCD: Liquid crystal display LTE: Long term evolution MV: Motion Vector OLED: Organic Light-Emitting Diode PB: Prediction Block PCI: Peripheral device interconnection PLD: Programmable Logic Device PU: Prediction Unit RAM: Random Access Memory ROM: Read-only memory SCC: Screen Content Coding SEI: Supplementary and Extended Information SNR: Signal-to-noise ratio SSD: Solid State Drive TU: Conversion Unit USB: Universal Serial Bus VUI: Video Usability Information VVC: Versatile Video Coding [Explanation of Symbols]
[0166] 101 samples 102 Arrow 103 Arrow 104 blocks 105 Schematic Diagram 111 blocks 200 Communication Systems 210 Terminal device 220 Terminal devices 230 Terminal devices 250 Networks 301 Video Sources 302 Video picture stream 303 Video Encoder 304 encoded video data 305 Streaming Server 306 Client Subsystem 307 Incoming call copy 310 Video Decoder 311 Video Picture 312 displays 313 Capture Subsystem 320 Electronic Devices 330 Electronic Devices 401 Channel 410 Video Decoder 412 rendering devices 415 buffer memory 420 Entropy Decoder / Parser 421 Symbols 430 Electronic Devices 431 Receiver 451 Scaler / Inverse Unit 452 IntraPicture Prediction Units 453 Motion Compensation Prediction Unit 455 Aggregator 456 Loop Filter Unit 457 Reference Picture Memory 458 picture buffer 501 Video Sources 503 Video Encoder 520 Electronic Devices 530 Source Coder 532 coding engine 533 Local Decoder 534 Reference Picture Memory 535 Predictor 540 Transmitter 543 Encoded Video Sequence 545 Entropy Coder 550 Controller 560 communication channels 603 Video Encoder 621 General-purpose controller 622 Intra Encoders 623 Residual Calculator 624 residual encoder 625 Entropy Encoder 626 switches 628 Residual Decoder 630 Interencoders 710 Video Decoder 771 Entropy Decoder 772 Intra Decoder 773 Residual Decoder 774 Reconstruction Module 780 Interdecoder 800 Current Picture 801 Current Block 802 Block Vectors 803 Reference Block 900 Encoding Tree Units (CTUs) 901 Current Block 902 Current Block 903 Current Block 904 Current Block 910 Encoding Tree Unit (CTU) 911 Blocks in the same location 912 Blocks in the same location 913 Blocks in the same location 914 Blocks in the same location 1000 processes 1100 Computer System 1101 Keyboard 1102 Mouse 1103 Trackpad 1105 Joystick 1106 Mike 1107 Scanner 1108 Camera 1109 Speaker 1110 Touchscreen 1121 Medium 1122 Thumb Drive 1123 Solid State Drive 1140 cores 1143 Field-Programmable Gate Area (FPGA) 1144 Accelerator 1145 Read-only memory (ROM) 1146 Random Access Memory 1147 Mass storage 1148 System Bus 1149 Peripheral bus
Claims
1. A method for video decoding performed by a decoder, comprising: decoding first prediction information indicating a first prediction mode used for a current block in a current encoded picture that is part of an encoded video sequence, wherein the first prediction mode is an intra block copy prediction mode; determining whether an adjacent block adjacent to the current block and reconstructed before the current block uses an intra prediction mode; inserting second prediction information from the adjacent block into a predictor list of the first prediction mode in response to (i) a determination that the adjacent block does not use the intra prediction mode, then (ii) a determination that the adjacent block uses the same first prediction mode as the current block, and (iii) a determination that the adjacent block is in at least one of the same slice, the same tile, or the same tile group as the current block; reconstructing the current block according to the predictor list of the first prediction mode; wherein the second prediction information decoded from the adjacent block includes a block vector; the method.
2. The method according to claim 1, further comprising determining whether the adjacent block is in the same slice as the current block, wherein the slice is a group of blocks in raster scan order and the blocks in the slice use the same prediction mode. The method according to claim 1, further comprising:
3. The method according to claim 1 or 2, further comprising determining whether the adjacent block is in the same tile or the same tile group as the current block, wherein the tile is a region of the picture and is processed in parallel and independently, and the tile group is a group of tiles that share the same header among the groups of tiles. The method according to claim 1 or 2, further comprising:
4. The method according to any one of claims 1 to 3, further comprising determining whether the adjacent block overlaps with the current block, and inserting the second prediction information into the predictor list of the first prediction mode in response to a determination that the adjacent block does not overlap with the current block. The method according to any one of claims 1 to 3, further comprising: inserting the second prediction information into the predictor list of the first prediction mode in response to a determination that the adjacent block does not overlap with the current block.
5. Decoding first prediction information indicating a first prediction mode used for a current block in a current encoded picture that is part of an encoded video sequence, wherein the first prediction mode is an intra block copy prediction mode, determining whether an adjacent block adjacent to the current block and reconstructed before the current block uses an intra prediction mode, inserting second prediction information from the adjacent block into a predictor list of the first prediction mode in response to (i) a determination that the adjacent block does not use the intra prediction mode, then (ii) a determination that the adjacent block uses the same first prediction mode as the current block, and (iii) a determination that the adjacent block is in at least one of the same slice, the same tile, or the same tile group as the current block, including a processing circuit configured to reconstruct the current block according to the predictor list of the first prediction mode, wherein the second prediction information decoded from the adjacent block includes a block vector, An apparatus.
6. The processing circuit is configured to determine whether the adjacent block is in the same slice as the current block, wherein the slice is a group of blocks in raster scan order and the group of blocks within the slice is further configured to use the same prediction mode. The apparatus according to claim 5.
7. The processing circuit is configured to determine whether the adjacent block is in the same tile or the same tile group as the current block, wherein the tile is a region of the picture and is processed in parallel and independently, and the tile group is a group of tiles and is further configured to share the same header among the groups of tiles. The apparatus according to claim 5 or 6.
8. The processing circuit is configured to determine whether the adjacent block overlaps with the current block, further configured to insert the second prediction information into the predictor list of the first prediction mode in response to a determination that the adjacent block does not overlap with the current block. The apparatus according to any one of claims 5 to 7.
9. A program for causing at least one processor to execute the method according to any one of claims 1 to 4.