Video decoder, method, non-transitory computer-readable medium, and method of video encoder
The RR-IBC mode optimizes Intra Block Copy prediction in video coding standards by determining flip modes and patterns, addressing memory and implementation challenges, enhancing decoding efficiency in VVC.
Patent Information
- Application Number
- JP2024518164
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-07
- Filing Date
- 2022-11-08
- Publication Date
- 2025-07-25
AI Technical Summary
Existing video coding standards, such as HEVC, face challenges in efficiently utilizing Intra Block Copy (IBC) modes for prediction, particularly in terms of memory management and hardware implementation complexity, especially with the introduction of Versatile Video Coding (VVC).
The implementation of reconstruction-reordered Intra Block Copy (RR-IBC) mode in video decoders, which includes determining the flip mode and predicting a flip pattern based on neighboring reconstruction samples and reference blocks, optimizing memory usage and decoding processes.
Enhances the efficiency of Intra Block Copy modes by reducing memory requirements and simplifying hardware implementation, improving decoding performance in video coding standards like VVC.
Smart Images

Figure 2025523729000003 
Figure 2025523729000004 
Figure 2025523729000005
Abstract
Description
Technical Field
[0001] [Related Applications] This application claims priority to U.S. Provisional Application No. 63 / 388,527, filed Jul. 12, 2022, and U.S. Patent Application No. 17 / 982,126, filed Nov. 7, 2022, which are hereby incorporated by reference in their entireties.
[0002] [Technical Field] This disclosure relates to Intra BC flip type prediction and signaling methods.
Background Art
[0003] The H.265 / HEVC (High Efficiency Video Coding) standards issued by ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, these two standardization organizations formed the JVET (Joint Video Exploration Team) together to explore the possibility of developing the next-generation video coding standard after HEVC. In October 2017, they issued the Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, a total of 22 CfP responses regarding standard dynamic range (SDR), 12 CfP responses regarding high dynamic range (HDR), and 12 CfP responses regarding 360 video categories were submitted respectively. In April 2018, all the received CfP responses were evaluated at the 122nd MPEG / 10th JVET meeting. As a result of this meeting, JVET officially started the standardization process for next-generation video coding beyond HEVC. The new standard was named Versatile Video Coding (VVC), and JVET was renamed the Joint Video Experts Team. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC29 / WG11) published the VVC video coding standard (version 1).
Summary of the Invention
[0004] The following presents a simplified overview of one or more embodiments of the present disclosure to provide a basic understanding of such embodiments. This overview is not an extensive overview of all intended embodiments, nor is it intended to identify key elements of all embodiments or to describe the scope of any or all embodiments. Its sole purpose is to present some concepts of one or more embodiments of the present disclosure in a simplified form as a prelude to the more detailed description presented later.
[0005] The present disclosure provides a method for Intra BC (Intra BC) flip type prediction and signaling.
[0006] According to an exemplary embodiment, a method executed by at least one processor in a video decoder includes receiving a coded video bitstream including a current picture including at least one block. The method includes determining that the at least one block is predicted in a RR-IBC (reconstruction-reordered intra block copy) mode. The method includes obtaining a syntax element from the at least one block, the syntax element indicating a flip mode. The method includes determining whether a reconstruction flip is applied to the at least one block. The method includes predicting a flip pattern of the at least one block based on neighboring reconstruction samples of the at least one block and a corresponding reference block of the at least one block in response to determining that the reconstruction flip is applied to the at least one block. The method further includes decoding the at least one block based on the flip mode and the predicted flip pattern.
[0007] According to an exemplary embodiment, a video decoder includes at least one memory configured to store computer program code, and at least one processor configured to access the computer program code and operate as instructed by the computer program code. The computer program code includes reception code configured to cause the at least one processor to receive a coded video bitstream including a current picture including at least one block. The computer program code includes first determination code configured to cause the at least one processor to determine that the at least one block should be predicted in a RR-IBC (reconstruction-reordered intra block copy) mode. The computer program code includes acquisition code configured to cause the at least one processor to obtain a syntax element indicating a flip mode from the at least one block. The computer program code includes second determination code configured to cause the at least one processor to determine whether a reconstruction flip is to be applied to the at least one block. The computer program code includes prediction code configured to cause the at least one processor to determine a flip pattern of the at least one block based on neighboring reconstruction samples of the at least one block and corresponding reference blocks of the at least one block in response to determining that the reconstruction flip is to be applied to the at least one block. The computer program code further includes decoding code configured to cause the at least one processor to decode the at least one block based on the flip mode and the predicted flip pattern.
[0008] According to an exemplary embodiment, there is provided a non-transitory computer-readable medium storing instructions that, when executed by a processor in a video decoder, cause the processor to execute a method, the method including receiving a coded video bitstream including a current picture. The method includes determining that at least one block is predicted in a RR-IBC (reconstruction-reordered intra block copy) mode. The method includes obtaining a syntax element from the at least one block, the syntax element indicating a flip mode. The method includes determining whether a reconstruction flip is applied to the at least one block. The method includes predicting a flip pattern of the at least one block based on neighboring reconstruction samples of the at least one block and corresponding reference blocks of the at least one block in response to determining that the reconstruction flip is applied to the at least one block. The method further includes decoding the at least one block based on the flip mode and the predicted flip pattern.
[0009] Additional embodiments will be described in the following description, and will be apparent in part from the description, and may also be learned by practice of the presented embodiments of the disclosure.
Brief Description of the Drawings
[0010] The above and other features and aspects of embodiments of the present disclosure will be apparent from the following description in connection with the accompanying drawings.
[0011]
Figure 1
[0012]
Figure 2
[0013]
Figure 3
[0014]
Figure 4
[0015]
Figure 5
[0016]
Figure 6
[0017]
Figure 7
[0018]
Figure 8
[0019]
Figure 9
[0020]
Figure 10
[0021]
Figure 11
[0022]
Figure 12
[0023]
Figure 13
[0024]
Figure 14A
Figure 14B
Figure 14C
[0025]
Figure 15
[0026]
Figure 16
[0027]
Figure 17
MODE FOR CARRYING OUT THE INVENTION
[0028] The following detailed description of the exemplary embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.
[0029] The foregoing disclosure has provided illustrations and descriptions, but is not intended to be exhaustive or to limit the implementation to the disclosed detailed forms. In light of the foregoing disclosure, changes and modifications are possible or may be obtained from experience in implementation. Further, one or more features or components of one embodiment can be incorporated into or combined with another embodiment (or one or more features of another embodiment). Further, in the flowcharts and descriptions of operations provided below, it is understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be executed (at least partially) simultaneously, and the order of one or more operations may be switched.
[0030] It will be apparent that the systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting. Accordingly, the operations and behaviors of the systems and / or methods are described herein without reference to specific software code. It is understood that software and hardware may be designed to implement the systems and / or methods based on the description herein.
[0031] Even if a particular combination of features is recited in the claims and / or disclosed in the specification, these combinations do not limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not recited in the claims and / or disclosed in the specification. Each of the dependent claims listed below can directly depend on only one claim, but the disclosure of possible implementations includes each dependent claim in combination with all other claims in the set of claims.
[0032] Any element, act, or instruction used in this specification should not be construed as important or essential unless explicitly described. Also, as used in this application specification, the articles "a" and "an" are intended to include one or more items and may be used synonymously with "one or more." When only one item is intended, the term "one" or a similar word is used. Also, the terms "comprising," "having" (has, have, having, include, including), etc., as used in this specification are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless explicitly stated otherwise. Further, an expression such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.
[0033] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) may include at least two terminals (110, 120) interconnected via a network (150). In one-way data transmission, the first terminal (110) may code video data at a local location for transmission to another terminal (120) via the network (150). The second terminal (120) may receive the coded video data of another terminal from the network (150), decode the coded data, and display the restored video data. One-way data transmission may be common in media serving applications and the like.
[0034] FIG. 1 shows a second terminal pair (130, 140) applied to support two-way transmission of a coding video, which may occur, for example, during a video conference. In two-way transmission of data, each terminal (130, 140) may code video data captured locally in order to transmit it to other terminals via a network (150). Each terminal (130, 140) may also receive the coded video data transmitted by other terminals, may decode the coded data, and may display the restored video data on a local display device.
[0035] In FIG. 1, the terminals (110-140) may be shown as a server, a personal computer, and a smartphone, and / or any other type of terminal. For example, the terminals (110-140) may be a laptop computer, a tablet computer, a media player, and / or dedicated video conferencing equipment. The network (150) represents any number of networks that transmit coded video data among the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data on a circuit-switched and / or packet-switched channel. Representative networks include an electronic communication network, a local area network, a wide area network, and / or the Internet. For the purposes of the discussion of the present invention, the architecture and topology of the network (150) may not be important for the operation of the present disclosure, unless otherwise specifically stated hereinafter.
[0036] FIG. 2 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example of the application of the disclosed subject matter. The disclosed subject matter is equally applicable to, for example, video conferencing, digital TV, CD, DVD, memory stick, etc., storage of compressed video on digital media, other video-enabled applications, etc.
[0037] As shown in FIG. 2, the streaming system (200) can include a capture subsystem (213) that can include a video source (201) and an encoder (203). The video source (201) can be, for example, a digital camera and can be configured to generate an uncompressed video sample stream (202). The uncompressed video sample stream (202) can provide a high data volume when compared to an encoded video bitstream and can be processed by an encoder (203) coupled to the video source (201) which can be, for example, a camera. The encoder (203) can include hardware, software, or a combination thereof and can enable or implement aspects of the disclosed subject matter as detailed below. The encoded video bitstream (204) can include a low data volume when compared to the sample stream and can be stored at a streaming server (205) for future use. One or more streaming clients (206) can access the streaming server (205) to read a video bitstream (209) that can be a copy of the encoded video bitstream (204).
[0038] In an embodiment, the streaming server (205) can also function as a MANE (Media-Aware Network Element). For example, the streaming server (205) can be configured to prune the encoded video bitstream (204) to adapt potentially different bitstreams to one or more streaming clients (206). In an embodiment, the MANE can be provided separately from the streaming server (205) within the streaming system (200).
[0039] The streaming client (206) can include a video decoder (210) and a display (212). The video decoder (210) can decode, for example, a video bitstream (209) that is an incoming copy of an encoded video bitstream (204) and generate an output video sample stream (211) that can be rendered on the display (212) or another rendering device (not shown). In some streaming systems, the video bitstreams (204, 209) can be encoded according to a particular video coding / compression standard. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. A video coding standard under development is known informally as VVC (Versatile Video Coding). Embodiments of the present disclosure may be used in the context of VVC.
[0040] FIG. 3 shows an exemplary functional block diagram of a video decoder (210) attached to a display (212) according to an embodiment of the present disclosure. The video decoder (210) can include a channel (312), a receiver (310), a buffer (315) which may be, for example, a buffer memory, an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra prediction unit (352), a motion compensation prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (358). In at least one embodiment, the video decoder (210) can include an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The video decoder (210) may be implemented partially or entirely in software executed on one or more CPUs with associated memory.
[0041] In this and other embodiments, receiver (310) may receive one or more coded video sequences to be decoded by decoder (210), one coded video sequence at a time. Here, the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequences may be received from a channel (312) that may be a hardware / software link to a storage device storing the encoded video data. Receiver (310) may receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams that may be transferred to respective usage entities (not shown). Receiver (310) may separate the coded video sequences from the other data. To remove network jitter, buffer (315) may be coupled between receiver (310) and entropy decoder / parser (320) (hereinafter, "parser"). Buffer (315) may not be used or may be made small when receiver (310) is receiving data controllably from a storage / transfer device with sufficient bandwidth or from an isosynchronous network. When used in a best-effort packet network such as the Internet, buffer (315) may be necessary and can be made relatively large and advantageously can be sized adaptively.
[0042] Video decoder (210) may include a parser (320) to reconstruct symbols (321) from an entropy-coded video sequence. The categories of these symbols include, for example, information used to manage the operation of decoder (210) and, in some cases, information for controlling a rendering device such as a display (212) that may be coupled to the decoder as shown in FIG. 2. The control information for the rendering device may be in the form of, for example, an SEI (Supplementary Enhancement Information) message or a VUI (Video Usability Information) parameter set fragment (not shown). The parser (320) may parse / entropy-decode the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser 320 may extract a set of subgroup parameters from the coded video sequence based on at least one parameter corresponding to at least one of the subgroups of pixels in the video decoder. Subgroups may include GOP (Groups of Picture), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (320) may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.
[0043] The parser (320) may perform an entropy decoding / parsing operation on the video sequence received from the buffer (315) to generate symbols (321). The reconstruction of the symbols (321) may include a plurality of different units depending on the type of the coded video picture or a part thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. Which units are included and how they are included can be controlled by subgroup control information parsed from the video sequence coded by the parser (320). Such a flow of subgroup control information between the parser (320) and the following plurality of units is not shown for clarity.
[0044] Beyond the functional blocks already mentioned, the decoder (210) may be conceptually subdivided into a number of functional units as will be described hereinafter. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0045] One unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive, as symbols (321) from the parser (320), quantized transform coefficients and control information including which transform should be used, block size, quantization coefficients, quantization scaling matrix, etc. The scaler / inverse transform unit (351) can output a block including sample values that can be input to the aggregator (355).
[0046] In some examples, the output samples of the scaler / inverse transform (351) may be related to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) generates a block of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information fetched from the current (partially reconstructed) picture in the picture memory (358). The aggregator 355, in some cases, adds, for each sample, the prediction information generated by the intra prediction unit 352 to the output sample information provided by the scaler / inverse transform unit 351.
[0047] In other cases, the output samples of the scaler / inverse transform unit (351) may be related to inter-coded, and possibly motion-compensated, blocks. In such cases, the motion-compensation prediction unit (353) can access the reference picture memory (357) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbols (321) associated with the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) to generate the output sample information (in this case, called residual samples or a residual signal). The address in the reference picture memory (357) from which the motion-compensation prediction unit (353) fetches the prediction samples may be controlled by a motion vector. The motion vector may be available to the motion-compensation prediction unit (353) in the form of, for example, a symbol (321) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of the sample values fetched from the reference picture memory (357) when an exact motion vector for sub-sampling is in use, a motion vector prediction mechanism, etc.
[0048] The output samples of the aggregator (355) may undergo various loop filtering techniques in the loop filter unit (356). The video compression technique is controlled by parameters included in the coded video bitstream and made available to the loop filter unit (356) as symbols (321) from the parser (320), but also responds to meta information obtained during the decoding of the previous part (in decoding order) of the coded picture or coded video sequence, and may include in-loop filtering techniques that can also respond to previously reconstructed and loop-filtered sample values.
[0049] The output of the loop filter unit (356) can be an output to a rendering device such as a display (212) and a sample stream that can be stored in the reference picture memory (357) for use in future inter-picture prediction.
[0050] Once a particular coded picture is completely reconstructed, it can be used as a reference picture for future prediction. When the coded picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (320)), the current reference picture can become part of the reference picture memory (357), and the fresh current picture memory can be reallocated before starting the reconstruction of subsequent coded pictures.
[0051] The video decoder (210) may perform a decoding operation in accordance with a predetermined video compression technique established by a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax of the video compression technique or standard, in the sense that the video compression technique or standard, and in its profile document therein, specifies that the coded video sequence conforms to the syntax specified by the video compression technique or standard in use. Also, in order to comply with some video compression techniques or standards, the complexity of the coded video sequence may be within the limits determined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured, for example, in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted through the HRD (Hypothetical Reference Decoder) specification and metadata for HRD buffer management signaled in the coded video sequence.
[0052] In an embodiment, the receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to correctly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR extension layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0053] FIG. 4 shows an exemplary functional block diagram of a video encoder (203) associated with a video source (201) according to an embodiment of the present disclosure. The video encoder (203) can include, for example, an encoder that is a source coder (430), a coding engine (432), a (local) decoder (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy coder (445), a control unit (450), and a channel (460).
[0054] The encoder (203) may receive video samples from a video source (201) (not part of the encoder) that can capture a video image to be coded by the encoder (203). The video source (201) may provide a source video sequence to be coded by the encoder (203) in the form of a digital video sample stream of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCb, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media providing system, the video source (201) may be a storage device storing pre-prepared video. In a video conferencing system, the video source (201) may be a camera that captures local image information as a video sequence. The video data may subsequently be provided as a plurality of individual pictures that give motion when viewed in succession. The pictures themselves may be organized as a spatial array of pixels. Each pixel may include one or more samples depending on the sampling structure, color space, etc. in use. One of ordinary skill in the art can immediately understand the relationship between a pixel and a sample. The following description focuses on samples.
[0055] According to one embodiment, the encoder (203) may code and compress pictures of the source video sequence into the coded video sequence (443) in real time or under any other time constraints required by the application. Implementing an appropriate coding speed is a function of the control unit (450). The control unit (450) may control other functional units as described hereinafter and may be functionally coupled to other functional units. The coupling is not shown for clarity. Parameters set by the control unit (450), rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique,...), picture size, GOP (group of pictures) layout, maximum motion vector search range, etc. may be included. Those skilled in the art can immediately identify other functions of the control unit (450) when related to a video encoder (203) optimized for a specific system design.
[0056] Some video encoders operate in what those skilled in the art would immediately recognize as a "coding loop." As a very simplified explanation, the coding loop can include an encoding portion of a source coder (430) that generates symbols based on an input picture and reference pictures to be coded, and a (local) decoder (433) incorporated within the encoder (203) that reconstructs the symbols to generate sample data that a (remote) decoder could generate when the compression between the symbols and the coded video bitstream is lossless in a particular video compression technique. This reconstructed sample stream may be input into a reference picture memory (434). When the decoding of the symbol stream yields bit-exact results independent of the decoder location (local or remote), the contents of the reference picture memory are also bit-exact between the local encoder and the remote encoder. In other words, the prediction portion of the encoder "sees" exactly the same sample values as the decoder would "see" as reference picture samples when prediction is used during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example due to channel errors) is known to those skilled in the art.
[0057] The operation of the "local" decoder (433) can be the same as that of the "remote" decoder (210) detailed above in connection with FIG. 3. However, since the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder (445) and the parser (320) can be lossless, the entropy decoding portion of the decoder (210) including the channel (312), the receiver (310), the buffer (315), and the parser (320) need not be fully implemented in the local decoder (433).
[0058] The consideration made in this regard is that any decoder technology other than the parse / entropy decoding existing in the decoder needs to exist in substantially the same functional form as that in the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technologies can be omitted since they may be the reverse of the decoder technologies that are comprehensively described. More detailed descriptions are necessary only in specific areas and are provided below.
[0059] During operation, in some examples, the source coder (430) may perform motion-compensated predictive coding. This predictively codes the input frame by referring to one or more previously coded frames from the video sequence designated as the "reference frame". In this method, the coding engine (432) codes the difference between a pixel block of the input frame and a pixel block of the reference frame that may be selected as a prediction reference for the input frame.
[0060] The local video decoder (433) may decode the coded video data of the frame that may be designated as the reference frame based on the symbols generated by the source coder (430). The operation of the coding engine (432) may advantageously be a lossy process. When the coded video data can be decoded in a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (433) may replicate the decoding process that may be performed by the video decoder for the reference frame and produce a reconstructed reference frame to be stored in the reference picture memory (434). In this way, the encoder (203) may locally store a copy of the reconstructed reference frame having the same content as the reconstructed reference frame obtained by the remote video decoder (if there are no transmission errors).
[0061] Predictor (435) may perform prediction search for the coding engine (432). That is, for a new frame to be coded, predictor (435) may search the reference picture memory (434) for sample data (such as a candidate reference pixel block) or specific metadata such as reference picture motion vectors, block shapes, etc. that can function as appropriate prediction criteria for the new picture. Predictor (435) may operate on a sample block - pixel block basis to find appropriate prediction criteria. In some examples, the input picture may have prediction criteria drawn from a plurality of reference pictures stored in the reference picture memory (434) as determined by the search results obtained by predictor (435).
[0062] Control unit (450) may manage the coding operations of video coder (430), including, for example, setting parameters and subgroup parameters used for the coding of video data. The outputs of all the aforementioned functional units may undergo entropy coding in entropy coder (445). The entropy coder converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques well known to those skilled in the art, such as Huffman coding, variable - length coding, arithmetic coding, etc.
[0063] The transmitter (440) may buffer the coded video sequence generated by the entropy coder (445) for transmission via a communication channel (460) that may be a hardware / software link to a storage device capable of storing the coded video data. The transmitter (440) may merge the coded video data from the video coder (430) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown). The control unit (450) may manage the operation of the encoder (203). During coding, the control unit (450) may assign to each coded picture a type of the particular coded picture that may affect the coding technique applicable to each picture. For example, a picture may often be assigned as an intra picture (I picture), a predicted picture (P picture), or a bi-directionally predicted picture (B picture).
[0064] An intra picture (I picture) may be a picture that can be coded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, IDR (Independent Decoder Refresh) pictures. Those skilled in the art will recognize variations of I pictures, and their individual applications and characteristics.
[0065] A predicted picture (P picture) may be a picture that can be coded and decoded using intra prediction or inter prediction, typically using one motion vector and a reference index to predict the sample values of each block.
[0066] A bi-directionally predictive picture (B picture) may be a picture that can be coded and decoded using intra prediction or inter prediction using up to two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predictive picture can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0067] The source picture may generally be spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and coded block by block. The blocks may be coded predictively by reference to other (already coded) blocks determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively, or they may be coded predictively by reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be coded predictively via spatial prediction or via temporal prediction by reference to one previously coded reference picture. Blocks of a B picture may be coded non-predictively via spatial prediction or via temporal prediction by reference to one or two previously coded reference pictures.
[0068] The video coder (203) may perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec. H.265. In that operation, the video coder (203) may perform various compression operations including predictive coding operations that utilize the temporal and spatial redundancy in the input video sequence. The coded video data may thus conform to the syntax specified by the video coding technology or standard being used.
[0069] In one embodiment, the transmitter (440) may transmit additional data along with the encoded video. The video coder (430) may include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI (Supplementary Enhancement Information) messages, VUI (Visual Usability Information) parameter set fragments, and the like.
[0070] In some embodiments, the IBC coding tool appears in the HEVC SCC extension as current picture referencing (CPR). The concept of IBC requires only the current frame, but follows the coding path for inter prediction. The main motivation behind the concept was the reference structure in which the addressing mechanism to the reference samples was a two-dimensional spatial vector. Another advantage of such an architecture is that the specification changes required for the integration of IBC are minimal, and assuming that the manufacturer already implements HEVC version 1, the implementation burden can be reduced. For the reasons described above, CPR in the HEVC SCC extension is a special inter prediction mode, and as a result, it brings about almost the same decoding process as the same syntax structure.
[0071] In some embodiments, since the IBC (or CPR) is in the mutual prediction mode, the intra-only prediction slice needs to be a prediction slice to enable the use of IBC. When IBC is applicable, the coder needs to extend the reference picture list by one entry of the pointer to the current picture. That is, the current picture takes a buffer of at most one picture size in the decoded picture buffer (DPB). The IBC mode signaling is implicit, that is, when the selected reference picture points to the current picture, the coding unit adopts IBC. In contrast to normal inter prediction, the reference samples in IBC processing are not filtered and the corresponding reference picture is a long-term reference. To minimize memory requirements, the coder can release the buffer immediately after reconstructing the current picture
[32] . If the reconstructed picture is a reference picture, the coder can return the filtered version of the reconstructed picture to the DPB as a short-term reference.
[0072] Figure 5 shows the concept of IBC in HEVC and VVC. Each square represents a coding tree unit (501). The gray shaded area (504) represents the already coded area, and the white shaded area (505) represents the next coding area. IBC in HEVC allows the use of the gray shaded area (504) except for the two CTUs at the upper right of the current CTU (501) to enable Wavefront Parallel Processing (WPP). On the other hand, IBC in VVC allows only the CTU on the left side of the current CTU, indicated by the dotted frame, as the reference area.
[0073] As shown in Fig. 5, the reference to the reconstruction area is performed via a two-dimensional so-called block vector (BV), similar to inter prediction, and its prediction and coding reuse the motion vector (MV) prediction and coding of the inter prediction process. However, the luma BV has an integer resolution instead of a 1 / 4 precision with respect to the possible BV (502) of a regular inter CTU in HEVC and the possible BV (503) of the current coding unit in VVC. The figure summarizes the concept of IBC in HEVC and VVC, and each square represents a coding tree unit (CTU) (501). The gray shaded area (504) indicates the already coded area, and the white shaded area (505) indicates the next coding area. IBC in HEVC allows the use of the gray shaded area (504) except for the two CTUs to the upper right of the current CTU in order to enable WPP (Wavefront Parallel Processing). On the other hand, IBC in VVC allows only the CTU to the left of the current CTU, indicated by the dotted frame prediction mode, as the reference area. As a result, the decoded motion vector difference (MVD) of the BV needs to be shifted two positions to the left before being added to the predictor for the final BV reconstruction.
[0074] Special processing is required for implementation and performance reasons, resulting in differences from the normal inter prediction mode, which are as follows. IBC reference samples are not filtered. That is, they are samples reconstructed before the in-loop filtering process including DBF and Sample Adaptive Offset (SAO). On the other hand, for other inter prediction modes of HEVC, filtered samples are used. There is no luma sample interpolation for IBC, and chroma sample interpolation is only required when the chroma BV is non-integer when derived from the luma BV. A special case occurs when the chroma BV is non-integer and the reference block is near the boundary of the available area. The surrounding reconstructed samples may be outside the boundary to perform chroma interpolation. The BV pointing to a single line adjacent to the boundary cannot avoid such cases.
[0075] Figure 6 shows the update process of the reference sample memory (RSM) at four intermediate times (601 to 604) during the reconstruction process. The lightly shaded area (605) indicates the reference samples of the left neighboring CTU, the darkly shaded area (606) indicates the reference samples of the current CTU, and the white shaded area (607) indicates the next coding area.
[0076] In some embodiments, the valid reference region for IBC in the HEVC SCC extension is, except for some exceptions for the purpose of parallel processing, the region where almost the entire current picture has already been reconstructed. FIG. 6 shows the reference region for IBC in HEVC and the configuration in VVC, where only the coding tree unit to the left of the current CTU (608) functions as the reference sample region at the start of the reconstruction process of the current CTU. A drawback of the concept in HEVC is the need for additional memory in the DPB where the hardware implementation typically uses external memory. The additional access to external memory is accompanied by an increase in memory bandwidth, making the concept using the DPB less attractive. VVC uses on-chip fixed memory that can implement IBC, significantly reducing the complexity of implementing IBC in the hardware architecture. Another important change is to address the signaling concept that is separated from the integration within the inter-prediction process as in the HEVC SCC extension.
[0077] In some embodiments, the IBC architecture in VVC forms a dedicated coding mode, and the IBC mode is a third prediction mode other than the intra and inter prediction modes. The bitstream transmits an IBC syntax element indicating the IBC mode of the coding unit when the block size is 64×64 or less. As a result, the maximum CU size that can utilize IBC is 64×64 to implement the continuous memory update mechanism of the reference sample memory (RSM). However, the reference sample addressing mechanism remains the same as in the HEVC SCC extension by showing a two-dimensional offset and reusing the vector coding process of inter prediction. Another special case occurs when the CST is active, and the coder cannot derive the chroma BV from the luma BV, and as a result, IBC is used only for the luma coding block.
[0078] The IBC design in VVC adopts a fixed memory size of 128×128 for each color component to store reference samples, enabling the possibility of on-chip placement in hardware implementation. The maximum CTU size in VVC is also 128×128, that is, the reference sample memory (RSM) can hold the samples of a single CTU when the maximum CTU size configuration is equal to 128×128. A special feature of the RSM is the continuous update mechanism that replaces the reconstructed samples of the left neighboring CTU with the reconstructed samples of the current CTU. Figure 6 shows an example of a simplified RSM for the update mechanism at four intermediate times (601 - 604) during the reconstruction process. The lightly shaded area (605) represents the reference samples of the left neighboring CTU, and the darkly shaded area (606) represents the reference samples of the current CTU (608). At the first intermediate time (601) representing the start of the reconstruction of the current CTU (608), the RSM is composed only of the reference samples of the left neighboring CTU. At the other three intermediate times (602 - 604), the reconstruction process replaces the samples of the left neighboring CTU with the transformation of the current CTU. The RSM is implicitly divided into four mutually independent 64×64 regions. The reset of this region occurs when the coder processes the first coding unit that would exist in the corresponding region when mapping the RSM to the CTU, facilitating the hardware implementation work. Figure 7 shows the concept of the spatial continuous update of the RSM, that is, the current CTU having the left neighboring CTU and the current coding unit. At the reconstruction time of this example, the process replaces the samples covered by the white shaded area (702) of the left neighboring CTU with the gray shaded area (701) of the current CTU. When the maximum CTU size is less than 128×128, the RSM can include multiple left neighboring CTUs, and as a result, multiple left neighboring CTUs are used. In some embodiments, when the maximum CTU size is equal to 32×32, the RSM can hold the samples of 15 left neighboring CTUs.
[0079] FIG. 7 shows the left neighboring CTU and the current CTU, and shows the effective reference region by the RSM design and its continuous update mechanism. The gray shaded area (701) covers the samples stored in the RSM, and the white shaded area (702) covers the replaced samples or the samples that have not been reconstructed.
[0080] In some embodiments, BV coding uses the processing specified for inter prediction, but uses simpler rules for candidate list construction. The candidate list construction for inter prediction can include candidates based on five spatial candidates, one temporal candidate, and six history-based candidates. The comparison of multiple candidates is necessary for history-based candidates to avoid duplicate entries in the final candidate list. Further, the list construction can include pairwise average candidates. In contrast, in the IBC list construction process, only two spatial neighboring BVs and five history-based BVs (HBVPs) are considered, and only the first HBVP is compared with the spatial candidates when added to the candidate list. In normal inter prediction, two different candidate lists are used, one for merge mode and the other for normal mode, while the IBC candidate list is for both cases. However, up to six candidate lists can be used in merge mode, while only the first twelve candidates are used in normal mode. Block vector difference (BVD) coding employs motion vector difference (MVD) processing to make the final BV of any size. Also, this means that the reconstructed BV may point to a region outside the reference sample region, and correction is required by using modulo operations on the width and height of the RSM to remove the absolute offset in each direction.
[0081] In some embodiments, in AV1, the Intra Block Copy (IntraBC) mode uses a vector to place a prediction block within the same picture of the current block. The vector is called a block vector (BV), is signaled in the bitstream, and the precision representing the BV is in integer points. The prediction process in the Intra BC mode is similar to inter-picture prediction, and the main difference is that in Intra BC, the predictor block is formed from the reconstructed samples of the current picture (before applying loop filtering). Thus, Intra BC can be regarded as "motion compensation" within the current picture using a block vector as a motion vector. For the current block, a flag used to indicate whether Intra BC is valid for the current block is first transmitted in the bitstream. Next, if the current block is in the Intra BC mode, a prediction BV is subtracted from the current BV to derive a BV difference, and the BV difference is classified into four types according to the horizontal and vertical components of the BV difference value. The type information needs to be signaled within the bitstream, and then the BV difference values of the two (horizontal and vertical) components may be signaled.
[0082] Intra BC is very effective for the coding of screen content, but it poses problems for hardware design. To facilitate the hardware design, the following changes have been adopted. When Intra BC is permitted, the loop filter including the Deblocking filter, CDEF, and LR is disabled. This allows avoiding a dedicated second picture buffer for enabling Intra BC. To facilitate parallel decoding, prediction cannot exceed the restricted area. In one superblock, if the coordinates of its upper left position are (x0, y0), prediction at position (x, y) can be accessed from Intra BC only when the vertical coordinate is less than y0 and the horizontal coordinate is less than x0 + 2(y0 - y). To allow for the write-back delay of the hardware, Intra BC prediction cannot access the immediate reconstruction area. The restricted immediate reconstruction area may be 1 to n superblocks. Therefore, in addition to Revision 2, when the coordinates of the upper left position of one superblock are (x0, y0), Intra BC can access the prediction at position (x, y) when the vertical coordinate is less than y0 and the horizontal coordinate is less than x0 + 2(y0 - y) - D. Here, D represents the immediate reconstruction area restricted for Intra BC. When D is two superblocks specified in the AVM. Figure 8 shows the prediction area of the Intra BC mode in one superblock prediction. The gray shaded area (801) represents the permitted search area. The black striped area (802) represents the non-permitted search area, and the white striped area (803) represents the current block.
[0083] In some embodiments, a redesigned Intra BC mode with local reference ranges is disclosed in addition to the AV1 codec. The disclosed Intra BC assumes that a 1SB-sized "on-chip" memory (referred to as reference sample memory or RSM) is allocated to store reference samples and a 64x64-based memory reuse mechanism is applied. In AV1, the following changes have been made to the design of Intra BC. The maximum block size in the Intra BC mode is limited to 64x64. The reference block and the current block are assumed to be in the same SB row. The reference block can only be located in the current SB or one SB to the left of the current SB. When any of the 64x64 unit reference sample memories starts to be updated with reconstructed samples from the current SB, the previously stored (from the left SB) reference samples in the entire 64x64 unit are marked as unusable for generating Intra BC prediction samples.
[0084] Figure 9 shows an example of such a memory reuse mechanism using the disclosed method. When starting to code each SB, the RSM stores the samples of the previously coded SB state (0). When the current block is located in any of the four 64x64 regions within the current SB, the corresponding region in the RSM is emptied and used to store the samples of the current 64x64 coding region. In this way, the samples in the RSM are gradually updated by the samples within the current SB. When the current SB is fully coded, the entire RSM is filled with all the samples of the current SB state (4). In this example, the current B is first split using a quadtree split (900). The coding order of the four 64x64 regions is top left, top right, bottom left, bottom right. For other block splitting decisions, the RSM update process is the same, i.e., each region in the RSM is replaced with the reconstructed samples within the current SB. The upper row (902) shows the perspective of the RSM. The lower row (903) shows the perspective of the picture.
[0085] FIG. 10 shows the memory update processing in the RSM during the decoding of the SB with horizontal split 1001 or vertical split 1002 in the SB route. Depending on the position of the current coding block with respect to the current SB, the following is applied. When the current block corresponds to the upper left 64x64 block of the current SB, in addition to the samples already reconstructed in the current SB, reference samples in the lower right, lower left, and upper right 64x64 blocks of the left SB can also be referred to. When the current block corresponds to the upper right 64x64 block of the current SB, in addition to the samples already reconstructed in the current SB, if the luma sample located at (0,64) with respect to the current SB has not been reconstructed yet, the current block can also refer to the reference samples in the lower left 64x64 block and the lower right 64x64 block of the left SB. In other cases, the current block can also refer to the reference samples in the lower right 64x64 block of the left SB. When the current block corresponds to the lower left 64x64 block of the current SB, in addition to the samples already reconstructed in the current SB, if the luma position (64,0) with respect to the current SB has not been reconstructed yet, the current block can also refer to the reference samples in the upper right 64x64 block and the lower right 64x64 block of the left SB. In other cases, the current block can also refer to the reference samples in the lower right 64x64 block of the left SB. When the current block corresponds to the lower right 64x64 block of the current SB, only the samples already reconstructed in the current SB can be referred to.
[0086] In some embodiments, screen content coding tools such as Intra Block Copy (IBC) generate prediction blocks by directly copying previously coded reference regions within the same picture. As shown in FIG. 11, symmetry is often observed in video content, particularly in text character regions and computer-generated graphics within a screen content sequence. Therefore, specific screen content coding tools that take symmetry into account would be efficient in compressing this type of video content. In FIG. 11, horizontal symmetry is illustrated between contents 1101A and 1101B, and between contents 1102A and 1102B. Vertical symmetry is illustrated between contents 1103A and 1103B, and between contents 1104A and 1104B.
[0087] In JVET-Z0159, the RR-IBC (Reconstruction-Reordered IBC) mode is disclosed for screen content coding. When this is applied, samples within the reconstruction block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, but the prediction block is derived without being flipped. On the decoder side, the reconstruction block is flipped back to restore the original block.
[0088] In the RR-IBC coded block, two flipping methods, horizontal flip and vertical flip, are supported. First, for the IBC AMVP coded block, a syntax flag indicating whether the reconstruction is flipped, i.e., ibc_flip_flag, is signaled. If it is flipped, another flag specifying the flip type, i.e., ibc_flip_type, is further signaled. In IBC merge, the flip type is inherited from the neighboring block without syntax signaling. Considering horizontal or vertical symmetry, the current block and the reference block are usually aligned horizontally or vertically. Therefore, when the horizontal flip is applied, the vertical component of BV is not signaled and is assumed to be equal to 0. Similarly, when the vertical flip is applied, the horizontal component of BV is not signaled and is assumed to be equal to 0.
[0089] Figure 12 shows the BV adjustment for horizontal flip 1201 and vertical flip 1202. To better utilize the symmetry characteristics, a BV adjustment approach considering the flip is applied to refine the block vector candidates. For example, as shown in Figure 12, (x nbr ,y nbr ) and (x cur ,y cur ) represent the coordinates of the central samples of the neighboring block and the current block respectively, and BV nbr and BV cur represent the BV of the neighboring block and the current block respectively. Instead of directly inheriting the BV from the neighboring block, when the neighboring block is coded with a horizontal flip, the horizontal component of BV nbr (shown as BV nbr h ) is added with a motion shift to calculate the horizontal component of BV cur . That is, BV cur h = 2(x nbr - x cur ) + BV nbr h . Similarly, when the neighboring block is coded with a vertical flip, the vertical component of BV nbr (BVnbr v By adding a motion shift to the one shown as), BV cur The vertical component of is calculated. That is, BV cur v = 2(y nbr - y cur ) + BV nbr v is.
[0090] In a reconstruction reordered IBC design, for the Intra BC AMVP mode, first, a flag indicating whether the flip method is applied is signaled. If the flip method is applied, another flag indicating whether it is a horizontal flip or a vertical flip is further signaled. In the case of SCC, since the residual is small, the signaling of these modes is expensive.
[0091] Embodiments of the present disclosure may be used separately or combined in any order. Further, embodiments of the present disclosure may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non - transitory computer - readable medium.
[0092] In some embodiments, by way of example, the reconstruction of the current block is flipped according to a specific flip type (e.g., flip pattern). However, without limitation, the same operation may be applied to other transformations of the reconstructed block, including rotation and zoom.
[0093] Hereinafter, when referring to a template, it can refer to the neighboring samples above, to the left, to the right, and below the block. FIG. 13 is an exemplary diagram of a block template. The gray area including the upper and left neighboring reconstruction samples shows the template (1302) of the current block (1301).
[0094] In some embodiments, the selection of whether to apply a reconstruction flip and / or the selection of a flip pattern (e.g., horizontal flip or vertical flip) can be predicted using neighboring reconstruction samples of the current block and the reference block. This selection can be represented by using neighboring reconstruction samples as template matching or boundary smoothness checking according to embodiments of the present disclosure.
[0095] In some embodiments, different candidate flip patterns (including no flip) can be applied to the template of the current block, then the distortion between the template of the reference block 1403 and the flipped template 1402 of the current block 1401 is calculated, and then the selection that provides the minimum distortion can be derived as the predicted flip pattern. FIGS. 14A - C show examples for calculating the template distortion 1404 for three different flip patterns (horizontal flip, vertical flip, no flip). Here, the distortions associated with each flip pattern (e.g., 1404A, 1404B, 1404C) are compared, and the predicted mode with the minimum distortion is derived as the predicted flip pattern. The distortion includes, but is not limited to, SAD, SSE, SATD. In FIG. 14A, the template is flipped horizontally (1402A), and in FIG. 14B, the template is flipped vertically (1402B).
[0096] In some embodiments, the reconstruction residual of the current block may be flipped using candidate flip types (including no flip), the boundary samples of the reconstructed block are used as inputs for calculating a smoothness score together with neighboring reconstruction samples, and the flip type that provides the highest smoothness score is derived as the predicted IBC flip type.
[0097] In some embodiments, smoothness can be calculated using spatial neighborhood reconstruction samples and boundary samples of the current block, using a specific flip type, no flip 1501, vertical flip 1502, or horizontal flip 1503, as shown in FIG. 15. Smoothness can be measured using any of the following equations.
[0098]
Number
[0099]
Number
[0100] Here, r represents the neighborhood reconstruction sample array, and p represents the reconstructed block of the current block.
[0101] In some embodiments, instead of explicitly signaling the flag ibc_flip_flag, a flag is signaled to indicate whether the flip type used by the IBC matches the predicted selection. For example, a flag can be used to indicate that a block is predicted in flip mode (e.g., a flip operation is applied to the block).
[0102] In some embodiments, instead of explicitly signaling the flag ibc_flip_type, a flag is signaled to indicate whether the selected IBC flip type matches the predicted IBC flip type. In some embodiments, when it is signaled that this flag is false, all flip types except the predicted flip type are sorted based on the distortion or smoothness metric calculated as described above, and the sorted index is binarized and signaled.
[0103] In some embodiments, the flag ibc_flip_flag can be signaled to indicate whether a horizontal flip or a vertical flip is used. If this flag is true, the selected IBC flip type is the same as the predicted IBC flip type. Otherwise, the flip operation of the IBC block is not performed.
[0104] In some embodiments, the flag ibc_flip_flag can be signaled to indicate whether a horizontal flip or a vertical flip is used. If this flag is true, another ibc_flip_pred_flag may be signaled to indicate whether the predicted IBC flip type is being used as the IBC flip type. Otherwise, the flip operation of the IBC block is not performed. If ibc_flip_pred_flag is true, the IBC flip type is the same as the predicted IBC flip type, and otherwise, another flip type is applied to the IBC block.
[0105] In some embodiments, each flip type may be associated with an index value, and all flip types are sorted using the distortion or smoothness metric calculated as described above, where the sorted index may be binarized and signaled.
[0106] In some embodiments, the absolute value of the block vector difference is used to derive a context used to entropy code the above-mentioned flag indicating whether the predicted value and the selected value are the same.
[0107] In some embodiments, whether a flip is used and the flip type information may be derived by a combination of template matching and checking the parity of the coefficient block of the current coding block. For example, the number of even coefficients (including 0) that are even within a block may be used to infer that a flip is used. By using template matching, which flipping direction can be derived. Other derivation methods based on the parity of the coefficients can be derived in a similar manner. In some embodiments, it is not necessary to combine the coefficient parity check with template matching. It may be used alone to derive information related to flipping.
[0108] FIG. 16 shows an embodiment of a process (1600) for decoding a block based on a flip mode. The process (1600) may be performed by a decoder such as decoder (210). The process may start at operation (1602), where a coded video bitstream is received. The coded video bitstream may include a current picture that includes at least one block. The process proceeds to operation (1604), where it is determined that at least one block is predicted in the RR-IBC mode. The process proceeds to operation (1606), where syntax elements are obtained from at least one block. The syntax elements may indicate a flip mode. The process proceeds to operation (1608), where it is determined whether a reconstruction flip is to be applied to at least one block. In response to determining that a reconstruction flip is to be applied to at least one block, the process proceeds to operation (1610), where the flip pattern of at least one block is predicted based on at least one block and a corresponding reference block. For example, the predicted flip pattern may be one of a horizontal flip pattern, a vertical flip pattern, or a no-flip pattern. The process proceeds to operation (1612), where at least one block is decoded based on the flip mode and the predicted flip pattern.
[0109] The process proceeds to operation (1606), and a flip mode is selected based on the predicted flip mode and the signaling information included in the coded video bitstream. The process proceeds to operation (1608), and at least one block is decoded based on the selected flip mode.
[0110] The techniques of the embodiments of the present disclosure described above can be implemented as computer software using computer-readable instructions and are physically stored on one or more computer-readable media. For example, FIG. 17 shows a computer system (1700) suitable for implementing embodiments of the subject matter of the present disclosure.
[0111] The computer software can be processed by mechanisms such as assembly, compilation, linking, etc. to generate code including instructions executable directly or through interpretation, microcode execution, etc. by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., and can be coded using any suitable machine code or computer language.
[0112] The instructions can be executed on various computers or components thereof including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0113] The components shown in FIG. 17 of the computer system (1700) are exemplary in nature and do not imply any limitation as to the use or functionality of the computer software implementing the embodiments of the present disclosure. Further, the configuration of the components should not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system (1700).
[0114] The computer system (1700) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users through, for example, sensory input (e.g., keystrokes, swipes, data grab operations), voice input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human interface device can also be used to capture specific media that is not necessarily directly related to conscious input by a human, such as voice (e.g., conversation, music, ambient sound), images (e.g., scanned images, photographic images obtained from a digital camera), and video (e.g., including 2D video, 3D video, stereoscopic video).
[0115] The input human interface device may include one or more of a keyboard (1701), a mouse (1702), a trackpad (1703), a touch screen (1710), a data grab, a joystick (1705), a microphone (1706), a scanner (1707), and a camera (1708) (only one of which is shown).
[0116] The computer system (1700) may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through sensory output, sound, light, and smell / taste. Such human interface output devices can include tactile output devices (e.g., a touch screen (1710), a data glove, or a joystick (1705) with tactile feedback, although there are also tactile feedback devices that do not function as input devices). For example, such devices may include audio output devices (such as speakers (1709), headphones (not shown), etc.), visual output devices (such as screens (1710) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without a touch screen input function and with or without a tactile feedback function, some of which can output two-dimensional visual output or output beyond three dimensions through means such as stereo output), virtual reality glasses (not shown), holographic displays, smoke tanks (not shown), and printers (not shown).
[0117] The computer system (1700) may also include a human-accessible memory device and associated media such as optical media like a CD / DVDROM / RW (1720) with media (1721) such as CDs / DVDs, a thumb drive (1722), a removable hard drive or a solid-state drive (1723), legacy magnetic media such as tapes and floppy disks (not shown), and devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown).
[0118] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.
[0119] The computer system (1700) may also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan area, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, vehicle and industrial including CANBus, etc. Certain networks generally require an external network interface attached to a specific general-purpose data port or peripheral device bus (1749) (e.g., a USB port of the computer system (1700)). Others are generally integrated into the core of the computer system 1700 by attachment to a system bus as described later (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using these networks, the computer system (1700) can communicate with other entities. Such communication can be only unidirectional reception (e.g., broadcast TV), only unidirectional transmission (e.g., CANBus to a specific CANBus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. Such communication can include communication to a cloud computing environment (1755). Specific protocols and protocol stacks may be used for each of the above-described networks and network interfaces.
[0120] The aforementioned human interface device, human-accessible storage device, and network interface (1754) may be attached to the core (1740) of the computer system (1700).
[0121] The core (1740) may include one or more central processing units (CPUs) (1741), a graphics processing unit (GPU) (1742), a dedicated programmable processing unit in the form of an FPGA (1743), a hardware accelerator for specific tasks (1744), and the like. These devices may be connected through a system bus (1748) together with a built-in mass storage device (1747) such as a read-only memory (ROM) (1745), a random access memory (1746), an internal hard drive that is not accessible to users, an SSD, and the like. In some computer systems, the system bus (1748) is accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, and the like. Peripheral devices can be attached directly to the core's system bus (1748) or through a peripheral device bus (1749). The architecture of the peripheral device bus includes PCI, USB, and the like. A graphics adapter (1750) may be included in the core.
[0122] The CPU (1741), GPU (1742), FPGA (1743), and accelerator (1744) can be combined to execute specific instructions that can generate the aforementioned computer code. The computer code can be stored in the ROM (1745) or RAM (1746). Temporary data can also be stored in the RAM (1746), while permanent data can be stored, for example, in the built-in mass storage device (1747). Fast storage and reading to any of the memory devices can be enabled through the use of a cache memory that can be closely associated with one or more of the CPU (1741), GPU (1742), mass storage device (1747), ROM (1745), RAM (1746), and the like.
[0123] A computer-readable medium may have computer code for performing operations implemented by various computers. The medium and the computer code may be specially designed and configured for the purposes of the present disclosure, or may be of the kind well known and available to those skilled in the field of computer software.
[0124] By way of example and not limitation, a computer system (1700) having an architecture, and in particular a core (1740), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be a specific storage device of the core (1740) with non-transitory characteristics such as an on-core mass storage device (1747) or ROM (1745), and media associated with a user-accessible mass storage device as described above. The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1740). The computer-readable media can include one or more memory devices or chips as required by a particular need. The software can cause the core (1740) and in particular the processor (including a CPU, GPU, FPGA, etc.) therein to execute a particular process or a particular part of a particular process described herein, including the definition of a data structure stored in a RAM (1746) and the modification of the data structure according to the process defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of an implementation (e.g., an accelerator (1744)) within logic hardwired or other circuitry operable with or instead of the software to execute a particular process or a particular part of a particular process described herein. References to software include logic and, where appropriate, vice versa. References to computer-readable media can, where appropriate, include circuitry (such as an integrated circuit (IC)) for storing software for execution, circuitry for implementing logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0125] The foregoing disclosure has provided illustration and description, but is not intended to be exhaustive or to limit implementation to the disclosed detailed forms. Modifications and variations are possible in light of the above disclosure, or may be acquired from experience in implementation.
[0126] It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed in this specification is an illustration of an exemplary approach. It is understood that the specific order or hierarchy of blocks in a process / flowchart may be rearranged based on design preferences. Additionally, some blocks may be combined or omitted. The appended method claims do not mean to present the elements of the various blocks in an exemplary order and are not limited to the specific order or hierarchy presented.
[0127] Some embodiments may relate to systems, methods, and / or computer-readable media in any possible integration of technical detail levels. Further, one or more of the above components may be stored on a computer-readable medium and implemented as instructions executable by at least one processor (and / or can include at least one processor). A computer-readable medium can include a computer-readable non-transitory storage medium having computer-readable program instructions thereon for causing a processor to execute operations.
[0128] A computer-readable storage medium may be a tangible device that holds and stores instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, static random access memory, portable compact disk read-only memory, digital versatile disks, memory sticks, floppy disks, mechanically-coded devices such as punch cards, or raised structures in grooves having instructions recorded thereon, and any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as, for example, radio waves or other freely-propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses passing through an optical fiber cable), or electrical signals transmitted through a wire.
[0129] The computer-readable program instructions described herein can be downloaded to each computer / processing device from a computer-readable storage medium or can be downloaded from an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.
[0130] The computer-readable program code / instructions for performing the calculations may be in any combination of one or more programming languages, including assembly instructions, instruction set architecture instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in an object-oriented programming language such as Smalltalk, C++, and a procedural programming language such as the "C" programming language or a similar programming language. The computer-readable program instructions can be executed entirely on the user's computer, partly on the user's computer, partly on the user's computer, partly on the user's computer, partly on a remote computer, partly on a remote computer or a remote computer or server, as a stand-alone software package. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network or a wide area network, or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to perform an aspect or operation.
[0131] These computer-readable program instructions are provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in the flowchart and / or block diagram block(s). These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement the function / act manner specified in the flowchart and / or block diagram block(s).
[0132] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block(s).
[0133] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible embodiments of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of one or more executable instructions for implementing a particular logical function. These methods, computer systems, and computer-readable media may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those shown in the drawings. In some alternative implementations, the functions described in the blocks may occur out of the order described in the figures. For example, two blocks shown in succession may in fact be executed simultaneously or substantially simultaneously, or the blocks may be executed in the reverse order, depending on the related functionality. Each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a special-purpose hardware-based system that performs the specified functions or operations, or a combination of special-purpose hardware and computer instructions.
[0134] It will be apparent that the systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.
[0135] The above disclosure also encompasses the embodiments listed below.
[0136] (Feature 1) A method executed by at least one processor in a video decoder, the method comprising: Receiving a coded video bitstream including a current picture including at least one block; Determining that the at least one block should be predicted in a RR-IBC (reconstruction-reordered intra block copy) mode; Obtaining a syntax element from the at least one block, the syntax element indicating a flip mode; Determining whether a reconstruction flip is to be applied to the at least one block; In response to determining that the reconstruction flip is to be applied to the at least one block, predicting a flip pattern of the at least one block based on neighboring reconstruction samples of the at least one block and corresponding reference blocks of the at least one block; Decoding the at least one block based on the flip mode and the predicted flip pattern; A method comprising.
[0137] (Feature 2) The method according to (Feature 1), wherein the predicted flip pattern is one of a horizontal flip pattern, a vertical flip pattern, and a no-flip pattern.
[0138] (Feature 3) The step of predicting the flip pattern includes calculating a template distortion of each flip pattern, The method according to (Feature 2), wherein the predicted flip pattern is the flip pattern having the minimum distortion.
[0139] (Feature 4) The at least one block includes a template including neighboring reconstruction samples on at least two sides of the at least one block, and determining the template distortion of the horizontal flip pattern includes determining the distortion between the template of the at least one block and the template applied to a corresponding reference block with the template flipped horizontally, and determining the template distortion of the vertical flip pattern includes determining the distortion between the template of the at least one block and the template applied to a corresponding reference block with the template flipped vertically, and determining the template distortion of the no-flip pattern includes determining the distortion between the template of the at least one block and the template applied to a corresponding reference block with the template not flipped, the method according to (Feature 3).
[0140] (Feature 5) The step of predicting the flip pattern includes calculating a smoothness score for each flip pattern, The predicted flip pattern is the flip pattern having the highest smoothness score, the method according to (Feature 2).
[0141] (Feature 6) The smoothness score of the no-flip pattern is based on the neighboring reconstruction samples of the at least one block and the boundary samples of the at least one block, the smoothness score of the horizontal flip pattern is based on the neighboring reconstruction samples of the at least one block flipped horizontally and the boundary samples of the at least one block flipped horizontally, and the smoothness score of the vertical flip pattern is based on the neighboring reconstruction samples of the at least one block flipped vertically and the boundary samples of the at least one block flipped vertically, the method according to (Feature 5).
[0142] (Feature 7) The selected flip pattern is the predicted flip pattern based on the determination that the coded video bitstream includes information indicating that the predicted flip pattern is to be used, according to the method described in (Feature 2).
[0143] (Feature 8) The selected flip pattern is the no-flip pattern based on the determination that the coded video bitstream includes information indicating that the predicted flip pattern is not to be used, according to the method described in (Feature 2).
[0144] (Feature 9) A video decoder, at least one memory configured to store computer program code, at least one processor configured to access the computer program code and operate as instructed by the computer program code, comprising, wherein the computer program code, reception code configured to cause the at least one processor to receive a coded video bitstream including a current picture including at least one block, first determination code configured to cause the at least one processor to determine that the at least one block is to be predicted in RR-IBC (reconstruction-reordered intra block copy) mode, acquisition code configured to cause the at least one processor to acquire syntax elements from the at least one block, the syntax elements indicating a flip mode, second determination code configured to cause the at least one processor to determine whether a reconstruction flip is to be applied to the at least one block, configured to cause the at least one processor to predict a flip pattern of the at least one block based on neighboring reconstruction samples of the at least one block and corresponding reference blocks of the at least one block in response to determining that the reconstruction flip is to be applied to the at least one block; configured to cause the at least one processor to decode the at least one block based on the flip mode and the predicted flip pattern; A video decoder comprising:
[0145] (Feature 10) The method according to (Feature 9), wherein the predicted flip pattern is one of a horizontal flip pattern, a vertical flip pattern, and a no-flip pattern.
[0146] (Feature 11) The prediction code is further configured to cause the processor to calculate a template distortion of each flip pattern. The video decoder according to (Feature 10), wherein the predicted flip pattern is a flip pattern having a minimum distortion.
[0147] (Feature 12) The at least one block includes a template including neighboring reconstruction samples on at least two sides of the at least one block. The prediction code is further configured to cause the at least one processor to determine the template distortion of the horizontal flip pattern by determining a distortion between the template of the at least one block and a template applied to a corresponding reference block obtained by flipping the template horizontally. The prediction code is further configured to cause the at least one processor to determine the template distortion of the vertical flip pattern by determining a distortion between the template of the at least one block and a template applied to a corresponding reference block obtained by flipping the template vertically. The prediction code is further configured to cause the at least one processor to determine the template distortion of the no-flip pattern by determining a distortion between a template of the at least one block and a template applied to a corresponding reference block where the template is not flipped, (Feature 11) the video decoder described.
[0148] (Feature 13) The prediction code is further configured to cause the at least one processor to predict the flip pattern by calculating a smoothness score for each flip pattern, The predicted flip pattern is a flip pattern having the highest smoothness score, (Feature 9) the video decoder described.
[0149] (Feature 14) The smoothness score of the no-flip pattern is based on neighboring reconstruction samples of the at least one block and boundary samples of the at least one block, and the smoothness score of the horizontal flip pattern is based on neighboring reconstruction samples of the at least one block flipped in the horizontal direction and boundary samples of the at least one block flipped in the horizontal direction, and the smoothness score of the vertical flip pattern is based on neighboring reconstruction samples of the at least one block flipped in the vertical direction and boundary samples of the at least one block flipped in the vertical direction, (Feature 13) the video decoder described.
[0150] (Feature 15) The selected flip pattern is the predicted flip pattern based on a determination that the coded video bitstream includes information indicating that the predicted flip pattern is to be used, (Feature 9) the video decoder described.
[0151] (Feature 16) The selected flip pattern is the no-flip pattern according to (Feature 9), based on the determination that the coded video bitstream contains information indicating that the predicted flip pattern is not used.
[0152] (Feature 17) A non-transitory computer-readable medium storing instructions that, when executed by a processor in a video decoder, cause the processor to perform a method, the method comprising: Receiving a coded video bitstream including a current picture including at least one block; Determining that the at least one block is to be predicted in RR-IBC (reconstruction-reordered intra block copy) mode; Obtaining a syntax element from the at least one block, the syntax element indicating a flip mode; Determining whether a reconstruction flip is to be applied to the at least one block; In response to determining that the reconstruction flip is to be applied to the at least one block, predicting a flip pattern of the at least one block based on neighboring reconstruction samples of the at least one block and corresponding reference blocks of the at least one block; Decoding the at least one block based on the flip mode and the predicted flip pattern; A non-transitory computer-readable medium including the above.
[0153] (Feature 18) The non-transitory computer-readable medium according to (Feature 17), wherein the predicted flip pattern is one of a horizontal flip pattern, a vertical flip pattern, and a no-flip pattern.
[0154] (Feature 19) The step of predicting the flip pattern includes the step of calculating the template distortion of each flip pattern, The predicted flip pattern is the flip pattern having the minimum distortion, the non-transitory computer-readable medium according to (Feature 17).
[0155] (Feature 20) The at least one block includes a template including neighboring reconstruction samples on at least two sides of the at least one block, and determining the template distortion of the horizontal flip pattern includes determining the distortion between the template of the at least one block and the template applied to the corresponding reference block with the template flipped horizontally, and determining the template distortion of the vertical flip pattern includes determining the distortion between the template of the at least one block and the template applied to the corresponding reference block with the template flipped vertically, and determining the template distortion of the no-flip pattern includes determining the distortion between the template of the at least one block and the template applied to the corresponding reference block with the template not flipped, the non-transitory computer-readable medium according to (Feature 19).
Claims
Claim 1 A method executed by at least one processor in a video decoder, the method comprising: receiving a coded video bitstream including a current picture including at least one block; determining that the at least one block is to be predicted in RR-IBC (reconstruction-reordered intra block copy) mode; obtaining a syntax element from the at least one block, the syntax element indicating a flip mode; determining whether a reconstruction flip is to be applied to the at least one block; in response to determining that the reconstruction flip is to be applied to the at least one block, predicting a flip pattern of the at least one block based on neighboring reconstruction samples of the at least one block and corresponding reference blocks of the at least one block; decoding the at least one block based on the flip mode and the predicted flip pattern; A method comprising the above steps. Claim 2 The method according to claim 1, wherein the predicted flip pattern is one of a horizontal flip pattern, a vertical flip pattern, and a no-flip pattern. Claim 3 The step of predicting the flip pattern includes calculating a template distortion of each flip pattern, The method according to claim 2, wherein the predicted flip pattern is the flip pattern having the minimum distortion. Claim 4 The at least one block includes a template including neighboring reconstruction samples on at least two sides of the at least one block, Determining the template distortion of the horizontal flip pattern includes determining the distortion between the template of the at least one block and the template applied to the corresponding reference block with the template flipped horizontally, Determining the template distortion of the vertical flip pattern includes determining the distortion between the template of the at least one block and the template applied to the corresponding reference block with the template flipped vertically, Determining the template distortion of the no-flip pattern includes determining the distortion between the template of the at least one block and the template applied to the corresponding reference block where the template is not flipped, according to the method of claim 3.
5. The step of predicting the flip pattern includes the step of calculating a smoothness score for each flip pattern, The predicted flip pattern is the flip pattern having the highest smoothness score, according to the method of claim 2.
6. The smoothness score of the no-flip pattern is based on the neighboring reconstruction samples of the at least one block and the boundary samples of the at least one block, The smoothness score of the horizontal flip pattern is based on the neighboring reconstruction samples of the at least one block flipped in the horizontal direction and the boundary samples of the at least one block flipped in the horizontal direction, The smoothness score of the vertical flip pattern is based on the neighboring reconstruction samples of the at least one block flipped in the vertical direction and the boundary samples of the at least one block flipped in the vertical direction, according to the method of claim 5.
7. The selected flip pattern is the predicted flip pattern based on the determination that the coded video bitstream includes information indicating that the predicted flip pattern is used, according to the method of claim 2.
8. The selected flip pattern is the no-flip pattern based on the determination that the coded video bitstream includes information indicating that the predicted flip pattern is not used, according to the method of claim 2.
9. A video decoder, At least one memory configured to store computer program code, At least one processor configured to access the computer program code and operate as directed by the computer program code, Including, The computer program code, Receiving code configured to cause the at least one processor to receive a coded video bitstream including a current picture including at least one block, a first determination code configured to cause the at least one processor to determine that the at least one block should be predicted in a RR-IBC (reconstruction-reordered intra block copy) mode; an acquisition code configured to cause the at least one processor to acquire a syntax element from the at least one block, the syntax element indicating a flip mode; a second determination code configured to cause the at least one processor to determine whether a reconstruction flip is to be applied to the at least one block; a prediction code configured to cause the at least one processor to predict a flip pattern of the at least one block based on neighboring reconstruction samples of the at least one block and a corresponding reference block of the at least one block in response to determining that the reconstruction flip is to be applied to the at least one block; a decoding code configured to cause the at least one processor to decode the at least one block based on the flip mode and the predicted flip pattern; A video decoder comprising.
10. The video decoder according to claim 9, wherein the predicted flip pattern is one of a horizontal flip pattern, a vertical flip pattern, and a no-flip pattern.
11. The prediction code is further configured to cause the processor to calculate a template distortion of each flip pattern. The video decoder according to claim 10, wherein the predicted flip pattern is a flip pattern having a minimum distortion.
12. The at least one block includes a template including neighboring reconstruction samples on at least two sides of the at least one block. The prediction code is further configured to cause the at least one processor to determine the template distortion of the horizontal flip pattern by determining a distortion between the template of the at least one block and a template applied to a corresponding reference block flipped in the horizontal direction. The prediction code is further configured to cause the at least one processor to determine the template distortion of the vertical flip pattern by determining a distortion between a template of the at least one block and a template applied to a corresponding reference block in which the template is vertically flipped. The video decoder according to claim 11, wherein the prediction code is further configured to cause the at least one processor to determine the template distortion of the no-flip pattern by determining a distortion between a template of the at least one block and a template applied to a corresponding reference block in which the template is not flipped. **Claim 13** The prediction code is further configured to cause the at least one processor to predict the flip pattern by calculating a smoothness score for each flip pattern. The video decoder according to claim 10, wherein the predicted flip pattern is a flip pattern having the highest smoothness score. **Claim 14** The smoothness score of the no-flip pattern is based on neighboring reconstruction samples of the at least one block and boundary samples of the at least one block. The smoothness score of the horizontal flip pattern is based on neighboring reconstruction samples of the at least one block flipped in the horizontal direction and boundary samples of the at least one block flipped in the horizontal direction. The video decoder according to claim 13, wherein the smoothness score of the vertical flip pattern is based on neighboring reconstruction samples of the at least one block flipped in the vertical direction and boundary samples of the at least one block flipped in the vertical direction. **Claim 15** The video decoder according to claim 9, wherein the selected flip pattern is the predicted flip pattern based on a determination that the coded video bitstream includes information indicating that the predicted flip pattern is to be used. **Claim 16** The video decoder according to claim 10, wherein the selected flip pattern is the no-flip pattern based on a determination that the coded video bitstream includes information indicating that the predicted flip pattern is not to be used. **Claim 17** A non-transitory computer-readable medium storing instructions, which, when executed by a processor in a video decoder, cause the processor to execute a method, the method comprising: Receiving a coded video bitstream including a current picture including at least one block; Determining that the at least one block is to be predicted in a RR-IBC (reconstruction-reordered intra block copy) mode; Obtaining a syntax element from the at least one block, the syntax element indicating a flip mode; Determining whether a reconstruction flip is to be applied to the at least one block; Predicting a flip pattern of the at least one block based on neighboring reconstruction samples of the at least one block and corresponding reference blocks of the at least one block in response to determining that the reconstruction flip is to be applied to the at least one block; Decoding the at least one block based on the flip mode and the predicted flip pattern; A non-transitory computer-readable medium including the above.
18. The non-transitory computer-readable medium according to claim 17, wherein the predicted flip pattern is one of a horizontal flip pattern, a vertical flip pattern, and a no-flip pattern.
19. The step of predicting the flip pattern includes calculating a template distortion of each flip pattern, The non-transitory computer-readable medium according to claim 18, wherein the predicted flip pattern is the flip pattern having the minimum distortion.
20. The at least one block includes a template including neighboring reconstruction samples on at least two sides of the at least one block, Determining the template distortion of the horizontal flip pattern includes determining the distortion between the template of the at least one block and the template applied to the corresponding reference block when the template is flipped horizontally. Determining the template distortion of the vertical flip pattern includes determining the distortion between the template of the at least one block and the template applied to a corresponding reference block with the template flipped in the vertical direction. Determining the template distortion of the no-flip pattern includes determining the distortion between the template of the at least one block and the template applied to a corresponding reference block with the template not flipped, the non-transitory computer-readable medium of claim 19. **Claim 21** A method executed by at least one processor in a video encoder, the method comprising: generating a coded video bitstream including a current picture including at least one block. The at least one block is predicted in RR-IBC (reconstruction-reordered intra block copy) mode. The at least one block includes a syntax element, and the syntax element indicates a flip mode. When a reconstruction flip is applied to the at least one block, a flip pattern of the at least one block is predicted based on neighboring reconstruction samples of the at least one block and a corresponding reference block of the at least one block. A method in which the at least one block is decoded based on the flip mode and the predicted flip pattern.