Method and apparatus for video decoding, and program

The video decoding apparatus addresses the challenges of video coding by decoding prediction information and determining block vectors to reconstruct samples efficiently, especially with variable CTU sizes, thereby improving encoding and decoding efficiency.

JP2025087842AActive Publication Date: 2025-06-10TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025035610
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-08-06
Filing Date
2025-03-06
Publication Date
2025-06-10
Estimated Expiration
2040-01-06

AI Technical Summary

Technical Problem

Current video coding technologies face challenges in efficiently encoding and decoding video data, particularly in reducing redundancy and managing large bitrates, especially when dealing with intra prediction and variable CTU sizes.

Method used

The proposed solution involves an apparatus for video decoding that includes a receiving circuit and a processing circuit. The processing circuit decodes prediction information for a current block in a current coding tree unit (CTU) from a coded video bitstream, determines a block vector pointing to a reference block, and reconstructs samples based on the reconstructed samples of the reference block read from the reference sample memory.

Benefits of technology

This approach enhances video encoding and decoding efficiency by effectively managing the search range and reference sample memory, particularly when the CTU size is variable, thereby reducing storage and bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087842000001_ABST
    Figure 2025087842000001_ABST
Patent Text Reader

Abstract

To provide a video encoding / decoding method configured to reduce redundancy in an input video signal, through compression.SOLUTION: A method of a video decoder includes: decoding prediction information of a current block in a current coding tree unit (CTU) from a coded video bitstream, the prediction information being indicative of an intra block copy mode, a size of the current CTU being smaller than a maximum size of a reference sample memory for storing reconstructed samples; determining a block vector that points to a reference block in the same picture as the current block, the reference block having reconstructed samples buffered in the reference sample memory; and reconstructing at least one sample of the current block based on the reconstructed samples of the reference block that are retrieved from the reference sample memory.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Incorporation by Reference This application claims the benefit of priority of U.S. Patent Application No. 16 / 533,719, filed on Aug. 6, 2019, titled "METHOD AND APPARATUS FOR VIDEO CODING", and claims the benefit of priority of U.S. Provisional Application No. 62 / 792,888, filed on Jan. 15, 2019, titled "SEARCH RANGE ADJUSTMENT WITH VARIABLE CTU SIZE FOR INTRA PICTURE BLOCK COMPENSATION". The entire disclosure of the prior applications is hereby incorporated by reference in its entirety.

[0002] Technical Field The present disclosure generally describes embodiments related to video coding.

Background Art

[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. The current research of the inventors, to the extent that the research is described in this background section and in the aspects of the description that may not otherwise be eligible as prior art at the time of filing, is not admitted as prior art to the present disclosure, either expressly or implicitly.

[0004] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of images, each of which has spatial dimensions, for example, of 1920×1080 luminance samples and associated color samples. The series of images can have a fixed or variable image rate (e.g., 60 images per second or 60 Hz). Uncompressed video requires a large bit rate. For example, 1080p60 4:2:0 video at 8 bits per sample (1920x1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbps. One hour of such video requires a storage area of more than 600 GB.

[0005] One purpose of video coding and decoding is to reduce the redundancy of the input video signal by compression. Compression can, in some cases, help reduce the bandwidth or storage space requirements by more than two orders of magnitude. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal can be used for its intended purpose. In the case of video, irreversible compression is widely used. The amount of allowable distortion depends on the application. For example, a user of a particular consumer streaming application may tolerate higher distortion than a user of a television distribution application. The achievable compression ratio can reflect the fact that higher allowable / tolerable distortion can result in a higher compression ratio.

[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy coding.

[0007] Video codec technology can include techniques known as intra coding. In intra coding, sample values are represented without referring to samples or other data from previously reconstructed reference images. In some video codecs, an image is spatially divided into blocks of samples. If all blocks of samples are coded in an intra mode, that image can be an intra image. Intra images and their derivatives such as independent decoder refresh images can be used to reset the decoder state and thus can be used as the first image in a coded video bitstream and video session or as a still image. Samples of an intra block can be subjected to a transform, and the transform coefficients can be quantized before entropy coding. Intra prediction can be a technique that minimizes sample values in the pre - transform domain. In some cases, the smaller the DC value after transform and the smaller the AC coefficients, the fewer bits are required for a given quantization step size to represent the block after entropy coding.

[0008] Conventional intra coding, such as that known from MPEG - 2 generation coding technology, does not use intra prediction. However, some newer video compression technologies include techniques that attempt to block data, for example, from surrounding, previously decoded in - decode order sample data and / or metadata obtained during the encoding / decoding of spatially proximate data. Such techniques are hereinafter referred to as "intra prediction" techniques. Note that in at least some cases, intra prediction uses only reference data from the current image being reconstructed and does not use reference data from reference images.

[0009] There can be various forms of intra prediction. If one or more of such techniques can be used in a given video coding technique, the technique in use can be coded in an intra prediction mode. In certain cases, the mode can have sub - modes and / or parameters, which can be coded individually or can be included in a mode codeword. The code name used for a given mode / sub - mode / parameter combination can affect the coding efficiency gain through intra prediction. Similarly, the entropy coding technique used to convert the code name into a bitstream can also potentially have an impact.

[0010] Certain intra prediction modes were introduced in H.264, improved in H.265, and further improved in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). The predictor block can be formed using neighboring sample values belonging to already available samples. The sample values of the neighboring samples are copied into the predictor block according to a direction. The reference to the direction in use can be coded in the bitstream or can itself be predicted.

[0011] Referring to FIG. 1, shown at the lower right is a subset of 9 predictor directions known from 33 possible predictor directions of H.265 (corresponding to 33 angular modes of 35 intra modes). The point (101) where the arrows converge represents the predicted sample. The arrows indicate the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples at an angle of 45 degrees from the horizontal and towards the upper right. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples (101) at an angle of 22.5 degrees from the horizontal direction and towards the lower left of sample (101).

[0012] Referring further to FIG. 1, shown at the upper left is a square block (104) of 4×4 samples (shown by the thick dashed line). The square block (104) contains 16 samples, each sample being labeled with an "S" and including its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within block (104). Since the block size is 4×4 samples, S44 is at the lower right. Further, reference samples are shown according to a similar numbering scheme. The reference samples are labeled with an "R", its Y position (e.g., row index) and X position (column index) relative to block (104). In both H.264 and H.265, since the predicted samples are adjacent to the block being reconstructed, there is no need to use negative values.

[0013] Intra-image prediction functions by copying the reference sample value from adjacent samples according to the signal prediction direction. For example, assume that the coded video bitstream includes a signal indicating a prediction direction that coincides with arrow (102) for this block. That is, the sample is predicted from one or more prediction samples at a 45-degree angle from the horizontal direction to the upper right. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Thereafter, sample S44 is predicted from reference sample R08.

[0014] In certain cases, especially when the direction is not evenly divisible by 45 degrees, the values of multiple reference samples can be combined, for example, by interpolation, to calculate the reference sample.

[0015] With the development of video coding technology, the number of possible directions is increasing. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and at the time of disclosure, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and certain techniques of entropy coding are used to represent those possible directions in a small number of bits, accepting a specific penalty for less likely directions. Furthermore, the direction itself can sometimes be predicted from the adjacent direction used in an adjacent, already decoded block.

[0016] FIG. 2 shows a schematic diagram (201) showing 65 intra prediction directions by JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of the intra prediction direction bits in the coded video bitstream representing the direction can vary from video coding technology to video coding technology and can range from a simple direct mapping of the prediction direction to complex adaptive schemes including the intra prediction mode, code name, most likely mode, and similar techniques. However, in any case, in video content, there may be a particular direction that is statistically less likely to occur than other particular directions. Since the goal of video compression is to reduce redundancy, in a well - operating video coding technique, the less likely direction will be represented by more bits than the more likely direction. SUMMARY OF THE INVENTION

[0018] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some embodiments, an apparatus for video decoding includes a receiving circuit and a processing circuit. For example, the processing circuit decodes prediction information for a current block in a current coding tree unit (CTU) from a coded video bitstream. The prediction information represents an intra - block copy mode. The size of the current CTU is smaller than the maximum size of a reference sample memory for storing reconstructed samples. The processing circuit determines a block vector pointing to a reference block in the same picture as the current block. The reference block has reconstructed samples buffered in the reference sample memory. Then, the processing circuit reconstructs at least one sample of the current block based on the reconstructed samples of the reference block read from the reference sample memory.

[0019] In some embodiments, the processing circuit is within the same CTU column as the current CTU and determines a block vector pointing to a reference block located in a region from the (N - 1)th CTU to the adjacent CTU to the left of the current CTU in the same CTU column as the current CTU. The maximum size of the reference sample memory is N times the size of the current CTU, and N is a positive number greater than 1.

[0020] In some embodiments, the processing circuit checks whether the top boundary of the reference block is within the same CTU column. Further, the processing circuit checks whether the bottom boundary of the reference block is within the same CTU column. Then, the processing circuit checks whether the left boundary of the reference block is to the right of the right side of the Nth CTU on the left, and checks whether the right boundary is to the left of the current CTU.

[0021] In some embodiments, the processing circuit checks whether the reference block is at least partially within the Nth CTU on the left within the same CTU column as the current CTU, and the maximum size of the reference sample memory is N times the size of the current CTU, where N is a positive number greater than 1.

[0022] Further, the processing circuit checks whether the left boundary of the reference block is within the Nth CTU on the left. When the reference block is at least partially within the Nth CTU on the left, the processing circuit determines whether the collocated block of the reference block is at least partially reconstructed within the current CTU. In one example, the processing circuit determines whether the upper left corner of the collocated block has been reconstructed. The processing circuit invalidates the block vector pointing to the reference block when the collocated block within the current CTU is at least partially reconstructed.

[0023] In some embodiments, the processing circuit determines a reference block region within the Nth CTU on the left that contains the reference block, and determines whether the collocated block region of the reference block region is at least partially reconstructed within the current CTU. Then, the processing circuit invalidates the block vector pointing to the reference block when the collocated block region within the current CTU is at least partially reconstructed.

[0024] Also, an aspect of the present disclosure provides a non-transitory computer-readable media storage storing instructions that, when executed by a computer for video decoding, cause the computer to execute a method for video coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 11

Figure 12

Figure 13

Figure 14

[0026] FIG. 3 shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can code video data (e.g., a stream of video images captured by the terminal device (310)) and transmit it to another terminal device (320) via the network (350). The encoded image data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to restore a video image, and display the video image according to the restored video data. Unidirectional data transmission can be common in media providing applications and the like.

[0027] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) and performs bidirectional transmission of coded video data that can occur, for example, during a video conference. For bidirectional data transmission, for example, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video images captured by the terminal device) to transmit to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can receive the coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to restore a video image, and display the video image on an accessible display device according to the restored video data.

[0028] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) can be shown as servers, personal computers, and smartphones, but the principles of the present invention are not limited thereto. Embodiments of the present invention find application in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that transmit encoded video data among the terminal devices (310), (320), (330), and (340), including, for example, wireline (wired) and / or wireless communication networks. The communication network (350) can exchange data within circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, unless otherwise described below, the architecture and topology of the network (250) may not be important for the operation of the present invention.

[0029] FIG. 4 shows the arrangement of a video encoder and a video decoder in a streaming environment as an application example of the disclosed subject matter. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital TV, and the storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0030] A streaming system can include, for example, a video source (401) that generates a stream of uncompressed video images (402), and a capture subsystem (413) that can include, for example, a digital camera. In one embodiment, the stream of video images (402) includes samples captured by a digital camera. The stream of video images (402), depicted as a thick line to emphasize the high data volume when compared to the encoded video data (404) (or coded video bitstream), can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof, and enables or implements aspects of the disclosed subject matter as will be described in more detail below. The encoded video data (404) (or encoded video bitstream (404)) is shown as a thin line to emphasize the lower data volume when compared to the stream of video images (402) and can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) and read copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and generates an output stream (411) of video images that can be rendered on a display (412) (such as a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be coded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265.For example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0031] Note that electronic devices (420) and (430) can include other components (not shown). For example, electronic device (420) can include a video decoder (not shown), and electronic device (430) can also include a video encoder (not shown).

[0032] FIG. 5 shows a block diagram of a video decoder (510) according to an embodiment of the present disclosure. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) of the example of FIG. 4.

[0033] The receiver (531) can receive one or more coded video sequences to be decoded by the video decoder (510), and in the same or another embodiment, can receive one coded video sequence at a time, where the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequence can be received from the channel (501), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (531) can receive the encoded video data together with other data, such as coded audio data and / or an accompanying data stream, and these data can be transferred using their respective entities (not shown). The receiver (531) can separate the coded video sequence from other data. To counter network jitter, a buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter "parser (520)"). For certain applications, the buffer memory (515) is part of the video decoder (510). In other cases, it can be outside the video decoder (510) (not shown). In yet another case, for example, to counter network jitter, a buffer memory (not shown) can be present outside the video decoder (510), and further, for example, to process the playout timing, another buffer memory (515) can be present inside the video decoder (510). If the receiver (531) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may be unnecessary or can be small.For use in a best-effort packet network such as the Internet, buffer memory (515) is required, relatively large, and advantageously may be of an adaptive size and may be implemented at least partially in an operating system or similar element (not shown) outside video decoder (510).

[0034] Video decoder (510) may include a parser (520) for reconstructing symbols (521) from a coded video sequence. The categories of these symbols, as shown in FIG. 5, include information used to manage the operation of video decoder (510) and potential information for controlling a rendering device (512) (e.g., a display screen) that is not an essential part of electronic device (530) but may be coupled to electronic device (530). The control information for the (multiple) rendering devices may be in the form of supplementary enhancement information (SEI messages) or video usability information (VUI) parameter set fragments (not shown). Parser (520) can parse / entropy decode the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser (520) can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to a group. The subgroups can include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. Parser (520) can also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.

[0035] The parser (520) may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (515) to generate symbols (521).

[0036] The reconstruction of the symbols (521) may include a plurality of different units depending on the type of the coded video image or a portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors. How the units are included can be controlled by subgroup control information parsed from the video sequence coded by the parser (520). The flow of such subgroup control information between the parser (520) and the following plurality of units is not illustrated for clarity.

[0037] In addition to the functional blocks already described, the video decoder (510) can conceptually be divided into several functional units as described below. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.

[0038] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives the quantized transform coefficients and control information including the transform, block size, quantization coefficient, quantization scaling matrix, etc. to be used as the symbol(s) (521) from the parser (520). The scaler / inverse transform unit (551) can output a block including sample values that can be input to the aggregator (555).

[0039] In some cases, the output samples of the scaler / inverse transform (551) can be related to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed image but can use prediction information from a previously reconstructed part of the current image. Such prediction information can be provided by the intra picture prediction unit (552). In some cases, the intra picture prediction unit (552) can use the surrounding already reconstructed information fetched from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, a partially reconstructed current image and / or a fully reconstructed current image. The aggregator (555) can, in some cases, add, for each sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0040] In other cases, the output samples of the scaler / inverse transform unit (551) can be related to inter-coding and potentially to motion compensated blocks. In such cases, the motion compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples to be used for prediction. According to the symbol (521) related to the block, after motion compensating the fetched samples, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, referred to as residual samples or residual signal) to generate the output sample information. The address in the reference picture memory (557) from which the motion compensation prediction unit (553) fetches the prediction samples can be controlled by the motion vector and is available to the motion compensation prediction unit (553) in the form of, for example, a symbol (521) having X, Y, and reference picture components. Motion compensation can also include interpolating the sample values so that they are fetched from the reference picture memory (557) when an exact motion vector of sub-samples is used, a motion vector prediction mechanism, etc.

[0041] The output samples of the aggregator (555) can undergo various loop filtering techniques within the loop filter unit (556). The video compression technology can include in-loop filter technologies, which are controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream), made available to the loop filter unit (556) as symbols (521) from the parser (520), and can respond to meta information obtained between the decoding of the coded image or coded video sequence and the preceding part in the decoding order, and can also respond to previously reconstructed and loop-filtered sample values, and can include in-loop filter technologies.

[0042] The output of the loop filter unit (556) can be output to the rendering device (512) and can be a sample stream that can be stored in the reference image memory (557) for use in future intra-image prediction.

[0043] Once the coded image is completely reconstructed, it can be used as a reference image for future prediction. For example, once the coded image corresponding to the current image is completely reconstructed and the coded image is identified as a reference image (e.g., by the parser (520)), the current image buffer (4558) can become part of the reference image memory (557), and the new current image buffer can be reallocated before starting the reconstruction of subsequent coded images.

[0044] The video decoder (510) can perform a decoding operation according to a predetermined video compression technique of a standard such as ITU-T Rec.H.265. The coded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence conforms to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile can select specific tools as the only tools available under that profile from all the tools available in the video compression technique or standard. Also, for compliance, it can be necessary that the complexity of the coded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level restricts the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (e.g., measured in megasamples per second), the maximum reference picture size, etc. The restrictions set by the level can, in some cases, be further restricted by the specifications of a Hypothetical Reference Decoder (HRD) and the metadata of HRD buffer management signaled in the coded video sequence.

[0045] In one embodiment, the receiver (531) can receive additional (redundant) data along with the encoded video. The additional data can be included as part of the coded (multiple) video sequences. The additional data can be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0046] FIG. 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). Instead of the video encoder (403) in the example of FIG. 4, the video encoder (603) can be used.

[0047] The video encoder (603) may receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that can capture the (multiple) video images to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0048] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit,...), any color space (e.g., BT.601 Y CrCB, RGB,...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media delivery system, the video source (601) can be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual images that provide motion when viewed in sequence. The images themselves can be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0049] According to one embodiment, a video encoder (603) can code and compress images of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required by an application. Achieving an appropriate coding speed is one function of a controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units as described below. The coupling is not shown for clarity. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, ...), picture size, picture group layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.

[0050] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop can include a source coder (630) (which is responsible for generating symbols such as a symbol stream, based on, for example, an input image and a reference image to be coded) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in the same manner as (and such that any compression between the symbols and the coded video bitstream is reversible in the video compression techniques contemplated by the disclosed subject matter) a (remote) decoder would also create. The reconstructed sample stream (sample data) is input into the reference image memory (634). Since the decoding of the symbol stream yields bit-exact results that are independent of the decoder location (local or remote), the content in the reference image memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as reference image samples that the decoder "sees" when using prediction during decoding. This basic principle of reference image synchronization (and the drift that results, for example, if synchronization cannot be maintained due to channel errors) is similarly used in some related arts.

[0051] The operation of the "local" decoder (533) can be the same as that of a "remote" decoder such as the video decoder (410), as already described in detail above in connection with FIG. 5. However, referring briefly to FIG. 5 as well, since symbols are available and the coding / decoding of the symbols into the coded video sequence by the entropy coder (645) and the parser (520) can be reversible, the entropy decoding portion of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).

[0052] The observations that can be made in this regard are that any decoder technology other than the parse / entropy decoder existing within the decoder must also exist in substantially the same functional form within the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. Since the description of encoder technology is the reverse of the comprehensively described decoder technology, it can be omitted. More detailed explanations are necessary only in certain fields and are provided below.

[0053] During operation, in some examples, the source coder (630) may perform motion-compensated predictive coding that predictively codes an input image with reference to one or more previously coded images from a video sequence designated as a "reference image". In this way, the coding engine (632) codes the difference between a pixel block of the input image and a pixel block of the (one or more) reference images that may be selected as the (one or more) prediction references for the input image.

[0054] The local video decoder (633) may decode the coded video data of an image that may be designated as a reference image based on the symbols generated by the source coder (630). The operation of the coding engine (632) may advantageously be an irreversible process. If the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) repeats the decoding process that is performed by the video decoder on the reference image and results in a reconstructed reference image to be stored in the reference image cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference image having common content as the reconstructed reference image that would be obtained by a distal video decoder.

[0055] Predictor (635) can perform a predictive search on the coding engine (632). That is, for a new image to be coded, the predictor (635) can search the reference image memory (634) for specific metadata such as reference image motion vectors, block shapes, etc. that can serve as appropriate predictive references for the new image, or sample data (as candidates for reference pixel blocks). The predictor (635) can operate on a per-sample block basis to find an appropriate predictive reference. In some cases, the input image can have a predictive reference drawn from a plurality of reference images stored in the reference image memory (634) as determined by the search results obtained by the predictor (635).

[0056] The controller (650) can manage the coding operations of the source coder (630), including, for example, setting parameters and subgroup parameters used to encode video data.

[0057] All outputs of the above-described functional units can undergo entropy coding in the entropy coder (645). The entropy coder (645) converts the symbols generated by the various functional units into a coded video sequence by reversibly compressing the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.

[0058] The transmitter (640) can buffer the coded video sequence created by the entropy coder (645) and prepare it for transmission via a communication channel (660) that can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) can merge the coded video data from the video coder (603) with other data to be transmitted, such as, for example, coded audio data and / or an auxiliary data stream (not shown).

[0059] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a specific coded image type to each coded image, which may affect the coding technique applicable to each image. For example, an image is often assigned as one of the following image types:

[0060] An intra picture (I picture) can be coded and decoded without using other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, for example, an independent decoder refresh (IDR) picture. Those skilled in the art recognize these variations of I pictures, as well as their respective uses and characteristics.

[0061] A predicted picture (P picture) can be coded and decoded using inter prediction or intra prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0062] A bi-directionally predicted picture (B picture) can be coded and decoded using inter prediction or intra prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use two or more reference pictures and associated metadata for the reconstruction of one block.

[0063] The source image is typically spatially divided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and coded block by block. The blocks can be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to each respective image of the block. For example, blocks of an I image are either non-predictively coded or they can be predictively coded with reference to already coded blocks of the same image (spatial prediction or inter prediction). Pixel blocks of a P image can be predictively coded via spatial prediction or temporal prediction with reference to one previously coded reference image. Blocks of a B image can be predictively coded via spatial prediction or temporal prediction with reference to one or two previously coded reference images.

[0064] The video encoder (603) can perform coding operations according to a predetermined video coding technique or a standard such as ITU-T Rec. H.265. In its operation, the video encoder (603) can perform various compression operations including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technique or standard being used.

[0065] In one embodiment, the transmitter (640) can transmit additional data along with the encoded video. The source coder (630) can include such data as part of the coded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0066] Video can be captured as a plurality of source images (video pictures) in a temporal sequence. Intra picture prediction (often abbreviated as intra prediction) uses the spatial correlation in a given picture, while inter picture prediction uses the correlation (temporal or otherwise) between pictures. In one example, a particular picture being coded / decoded, referred to as the current picture, is divided into blocks. If a block within the current picture is similar to a reference block within a reference picture that has been previously coded and is still buffered within the video, the block within the current picture can be coded by a vector referred to as a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension identifying the reference picture if multiple reference pictures are being used.

[0067] In some embodiments, a bi-prediction technique can be used in inter picture prediction. According to the bi-prediction technique, two reference pictures such as a first reference picture and a second reference picture that both precede the current picture in the video in decode order (but can be past and future respectively in display order) are used. A block within the current picture can be coded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0068] Furthermore, to improve coding efficiency, a merge mode technique can be used for inter picture prediction.

[0069] According to some embodiments of the present disclosure, predictions such as inter-image prediction and intra-image prediction are performed in units of blocks. For example, according to the HEVC standard, images in a video image sequence are partitioned into coding tree units (CTUs) for compression, and the CTUs in an image have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs) which are one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree partitioned into one or more coding units (CUs). For example, a 64×64 pixel CTU can be partitioned into 1 CU of 64×64 pixels, 4 CUs of 32×32 pixels, or 16 CUs of 16×16 pixels. In the example, each CU is analyzed to determine the prediction type of the CU, such as an inter-prediction type or an intra-prediction type. The CU is partitioned into one or more prediction units (PUs) depending on the temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0070] FIG. 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values in a current video image within a video image sequence and encode the processing block into a coded image that is part of a coded video sequence. In one embodiment, the video encoder (703) is used in place of the video encoder (403) of the example of FIG. 4.

[0071] In an example of HEVC, a video encoder (703) receives a matrix of sample values for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) determines whether the processing block is best coded using an intra mode, an inter mode, or a bi - directional prediction mode, for example, using rate - distortion optimization. If the processing block is coded in the intra mode, the video encoder (703) can use intra prediction techniques to encode the processing block into the coded picture. If the processing block is coded in the inter mode or the bi - directional prediction mode, the video encoder (703) can use inter prediction techniques or bi - directional prediction techniques, respectively, to code the processing block into the coded picture. In certain video coding techniques, the merge mode can be an inter - picture prediction sub - mode in which the motion vector is derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictor. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one embodiment, the video encoder (703) includes other components, such as a mode - decision module (not shown) for determining the mode of the processing block.

[0072] In the example of FIG. 7, the video encoder (703) includes an entropy encoder (725), an inter - encoder (730), an intra - encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), and a general - purpose controller (721) that are coupled together as shown in FIG. 7.

[0073] The inter-encoder (730) receives samples of the current block (e.g., a processing block), compares the block with one or more reference blocks in a reference image (e.g., blocks in a preceding image and a subsequent image), generates inter-prediction information (e.g., a description of redundant information by an inter-encoding technique, a motion vector, merge mode information), and is configured to calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference image is a decoded reference image decoded based on the encoded video information.

[0074] The intra-encoder (722) receives samples of the current block (e.g., a processing block), optionally compares the block with a block already coded in the same image, generates quantized coefficients after transformation, and is optionally also configured to generate intra-prediction information (e.g., intra-prediction direction information according to one or more intra-coding techniques). In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks in the same image.

[0075] The general-purpose control device (721) is configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one embodiment, the general-purpose controller (721) determines the mode of a block and supplies a control signal to a switch (726) based on the mode. For example, when the mode is the intra mode, the general-purpose controller (721) controls the switch (726) to select the result of the intra mode used by the residual calculator (723), controls the entropy encoder (725) to select the intra-prediction information, includes the intra-prediction information in the bitstream, and when the mode is the inter mode, the general-purpose controller (721) controls the switch (726) to select the inter-prediction result used by the residual calculator (723), controls the entropy encoder (725) to select the inter-prediction information, and includes the inter-prediction information in the bitstream.

[0076] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) operates based on the residual data and is configured to encode the residual data to generate a conversion coefficient. In one embodiment, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate a conversion coefficient. Thereafter, the conversion coefficient is subjected to quantization processing to obtain a quantized conversion coefficient. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse conversion and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded image in some embodiments, and the decoded image can be buffered in a memory circuit (not shown) and used as a reference image.

[0077] The entropy encoder (725) is configured to format the bitstream to include the encoded blocks. The entropy encoder (725) is configured to include various information according to an appropriate standard such as the HEVC standard. In one embodiment, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. Note that when coding a block in either the merge submode of the inter mode or the bi - directional prediction mode according to the disclosed subject matter, there is no residual information.

[0078] FIG. 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive a coded image that is part of a coded video sequence and decode the coded image to generate a reconstructed image. In one embodiment, the video decoder (810) is used in place of the video decoder (410) of the embodiment of FIG. 3.

[0079] In the example of FIG. 8, the video decoder (810) includes an intra decoder (872), an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and others, which are combined together as shown in FIG. 8.

[0080] The entropy decoder (871) can be configured to reconstruct specific symbols representing syntax elements from which the coded picture is made up from the coded picture. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, inter mode, bi - directional prediction mode, merge sub - mode or inter mode, bi - directional prediction mode, etc. in another sub - mode), prediction information (e.g., intra prediction information or inter prediction information), and they can identify specific samples or metadata used by the intra decoder (822) or the inter decoder (880), respectively, such as residual information in the form of quantized transform coefficients, etc. As an example, when the prediction mode is inter mode or bi - directional prediction mode, inter prediction information is provided to the inter decoder (880), and when the prediction type is intra prediction type, intra prediction information is provided to the intra decoder (872). The residual information can undergo inverse quantization and is provided to the residual decoder (873).

[0081] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0082] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0083] The residual decoder (873) is configured to perform inverse quantization to extract de - quantized transform coefficients, and process the de - quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require specific control information (including quantization parameter (QP)), and such information can be provided by the entropy decoder (871) (the data path is not shown as it is only low - volume control information).

[0084] The reconstruction module (874) is configured to combine, in a spatial domain, a residual as an output by the residual decoder (873) and a prediction result (optionally as an output by an inter or intra prediction module) to form a reconstruction block, which can be part of a reconstructed image and thus can be part of a reconstructed video. Note that other appropriate operations such as a deblocking operation can be performed to improve visual quality.

[0085] Note that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) can be realized using one or more processors that execute software instructions.

[0086] Aspects of the present disclosure provide encoding / decoding techniques for intra image block compensation, particularly techniques for adjusting a search range with a variable CTU size.

[0087] Block-based compensation can be used for both inter prediction and intra prediction. For inter prediction, block-based compensation can also be performed from a previously reconstructed area within the same picture. Block-based compensation from a reconstructed area within the same picture is called intra picture block compensation, current picture referencing (CPR), or intra block copy (IBC). A displacement vector indicating the offset between the current block and the reference block within the same picture is called a block vector (abbreviated as BV). Different from the motion vector in motion compensation, the block vector can be any value (positive or negative, either in the x or y direction), but the block vector has some constraints to ensure that the reference block is available and has already been reconstructed. Also, in some examples, to consider parallel processing, some reference areas that are tile boundaries or wavefront ladder shape boundaries are excluded.

[0088] The coding of the block vector can be either explicit or implicit. In the explicit mode (also called the advanced motion vector prediction (AMVP) mode in inter coding), the difference between the block vector and its predictor is signaled, and in the implicit mode, the block vector is restored from a predictor (referred to as a block vector predictor) in the same way as the motion vector in the merge mode. The resolution of the block vector is limited to integer positions in some implementations, but in other systems, the block vector is allowed to point to fractional positions.

[0089] In some embodiments, the use of intra-block copy at the block level can be signaled using a block-level flag called the IBC flag. In an embodiment, the IBC flag is signaled when the current block is not coded in merge mode. In other examples, the use of intra-block copy at the block level is signaled by a reference index approach. The current picture during decoding is then treated as a reference picture. In one embodiment, such a reference picture is placed at the last position in a list of reference pictures. This special reference picture is managed together with other temporal reference pictures in a buffer such as a decoded picture buffer (DPB).

[0090] Also, there are several variations of intra-block copy, such as flipped intra block copy (the reference block is flipped horizontally or vertically before being used to predict the current block), and line based intra block copy (each compensation unit within an MxN coding block is an Mx1 or 1xN line).

[0091] FIG. 9 shows an example of intra-block copy according to an embodiment of the present disclosure. The current picture (900) is being decoded. The current picture (900) includes a reconstructed area (910) (the dotted area) and an area to be decoded (920) (the white area). The current block (930) is being reconstructed by the decoder. The current block (930) can be reconstructed from a reference block (940) within the reconstructed area (910). The position offset between the reference block (940) and the current block (930) is called a block vector (950) (or BV (950)).

[0092] In some examples (e.g., VVC), the search range of the intra block copy mode is restricted to be within the current CTU. At this time, the memory requirement for storing reference samples for the intra block copy mode is 1 (largest) CTU size of samples. In one example, the (largest) CTU has a sample size of 128×128. The CTU is divided into 4 block regions, each having a size of 64×64 samples, in some examples. Thus, in some embodiments, the total memory (e.g., cache memory having an access speed faster than the main storage) can store samples of size 128×128, and the total memory includes an existing reference sample memory portion for storing samples reconfigured into the current block, such as a 64×64 region, and an additional memory portion for storing samples of the other 3 regions of size 64×64. In this way, in some examples, the effective search range of the intra block copy mode is extended to a part of the left CTU (e.g., 1 CTU size, 4 times the total 64×64 reference sample memory) while the total memory requirement for storing reference pixels remains unchanged.

[0093] In some embodiments, an update process is executed to update the stored reference samples from the left CTU to the samples reconstructed from the current CTU. Specifically, in some examples, the update process is executed based on 64×64 luma samples. In one embodiment, for each of the 4 64×64 block regions in the CTU size memory, the reference samples in the region from the left CTU can be used to predict the coding block in the current CTU having the CPR mode until any block in the same region of the current CTU is being coded or has been coded.

[0094] Figures 10A to 10D show examples of the effective search range for the intra-block copy mode according to an embodiment of the present disclosure. In some examples, the encoder / decoder includes a cache memory that can store samples of one CTU, such as 128×128 samples. Further, in the examples of Figures 10A to 10D, the current block region for prediction has a size of 64×64 samples. Note that the examples can be appropriately modified for other suitable sizes of the current block region.

[0095] Each of FIGS. 10A to 10D shows a current CTU (1020) and a left CTU (1010). The left CTU (1010) includes four block regions (1011) to (1014), and each block region has a sample size of 64×64 samples. The current CTU (1020) includes four block regions (1021) to (1024), and each block region has a sample size of 64×64 samples. The current CTU (1020) is the CTU that includes the current block region being reconstructed (indicated by the label "current" and the vertical stripe pattern). The left CTU (1010) is directly adjacent to the left of the current CTU (1020). As shown in FIGS. 10A to 10D, the gray blocks are the already reconstructed block regions, and the white blocks are the block regions to be reconstructed.

[0096] In FIG. 10A, the current block region being reconstructed is block region (1021). The cache memory stores the reconstructed samples within block regions (1012), (1013), and (1014), and the cache memory will be used to store the reconstructed samples of the current block region (1021). In the example of FIG. 10A, the valid search range for the current block region (1021) includes block regions (1012), (1013), and (1014) within the left CTU (1010) that have the reconstructed samples stored in the cache memory. In one embodiment, note that the reconstructed samples of block region (1011) are stored in the main memory which has a slower access speed than the cache memory (e.g., they are copied from the cache memory to the main memory before the reconstruction of block region (1021)).

[0097] In FIG. 10B, the current block region being reconstructed is block region (1022). The cache memory stores the reconstructed samples within block regions (1013), (1014), and (1021), and the cache memory will be used to store the reconstructed samples of the current block region (1022). In the example of FIG. 10B, the valid search range for the current block region (1022) includes block regions (1013) and (1014) within the left CTU (1010) that have the reconstructed samples stored in the cache memory and block region (1021) within the current CTU (1010). In one embodiment, note that the reconstructed samples of block region (1012) are stored in the main memory which has a slower access speed than the cache memory (e.g., they are copied from the cache memory to the main memory before the reconstruction of block region (1022)).

[0098] In FIG. 10C, the current block region being reconstructed is block region (1023). The cache memory stores the reconstructed samples within block regions (1014), (1021), and (1022), and the cache memory will be used to store the reconstructed samples of the current block region (1023). In the example of FIG. 10C, the valid search range for the current block region (1023) includes block region (1014) within the left CTU (1010) having the reconstructed samples stored in the cache memory, and block regions (1021) and (1022) within the current CTU (1010). Note that in one embodiment, the reconstructed samples of block region (1013) are stored in a main memory that is slower to access than the cache memory (e.g., copied from the cache memory to the main memory before the reconstruction of block region (1023)).

[0099] In FIG. 10D, the current block region being reconstructed is block region (1024). The cache memory stores the reconstructed samples within block regions (1021), (1022), and (1023), and the cache memory will be used to store the reconstructed samples of the current block region (1024). In the example of FIG. 10D, the valid search range for the current block region (1024) includes block regions (1021), (1022), and (1023) within the current CTU (1020) having the reconstructed samples stored in the cache memory. Note that in one embodiment, the reconstructed samples of block region (1014) are stored in a main memory that is slower to access than the cache memory (e.g., copied from the cache memory to the main memory before the reconstruction of block region (1024)).

[0100] In the above example, the cache memory has a total memory space of 1 (maximum) CTU size. The example can be appropriately adjusted to other suitable CTU sizes. Note that the cache memory is referred to as a reference sample memory in some embodiments.

[0101] The proposed methods can be used separately or in combination in any order. Further, each of the methods (or embodiments), encoders, and decoders can be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. Hereinafter, the term block may be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU.

[0102] Aspects of the present disclosure provide techniques for adjusting a search range when the CTU size changes, e.g., to be smaller than the maximum CTU size. In an implementation, a designated memory that stores reference samples of CUs coded in advance for future intra-block copy references is referred to as a reference sample memory (and in some embodiments as a cache memory). In the present disclosure, methods for improving intra-block copy performance under specific reference area restrictions are proposed. More specifically, the size of the reference sample memory for the search is restricted. In the following description, the size of the reference sample memory is fixed to 128×128 luma samples (along with corresponding chroma samples). In some embodiments (e.g., the VVC standard), one maximum CTU size of the reference samples is considered as the designated memory size. The proposed method can be further extended to various combinations of memory sizes / CTU sizes, e.g., 64×64 luma samples (plus corresponding chroma samples) for the CTU size and 128×128 luma samples (and corresponding chroma samples) for the memory size.

[0103] According to one aspect of the disclosure, the collocated blocks of the present disclosure refer to a pair of blocks having the same size and the same shape. One of the collocated blocks is within a previously coded CTU, and the other of the collocated blocks is within the current CTU. One block of the pair is referred to as the collocated block of the other block of the pair. Further, when the memory buffer size is designed to store a CTU of the maximum size (e.g., 128×128), in one embodiment, the previous CTU refers to a CTU having a CTU width luma sample offset of one CTU to the left of the current CTU. Further, these two collocated blocks have the same position offset value with respect to the upper left corner of their respective CTUs. In other words, the collocated blocks have the same y coordinate with respect to the upper left corner of the image, but in some embodiments, they are two blocks with a difference in CTU width in the x coordinate from each other.

[0104] FIG. 11 shows an example of collocated blocks according to some embodiments of the present disclosure. In the example of FIG. 11, the current CTU and the left CTU during decoding are shown. The reconstructed regions are shown in gray and the regions to be reconstructed are shown in white. FIG. 11 shows three examples of reference blocks in the left CTU for the current block in the intra-block copy mode during decoding. The three examples are shown as reference block 1, reference block 2, and reference block 3. FIG. 11 also shows collocated block 1 for reference block 1, collocated block 2 for reference block 2, and collocated block 3 for reference block 3. In the example of FIG. 11, the reference sample memory size is the CTU size. The reconstructed samples of the current CTU and the left CTU are stored in the reference sample memory in a complementary way. When the reconstructed samples of the current CTU are written to the reference sample memory, the reconstructed samples are written to the positions of the collocated samples of the left CTU. As an example, for reference block 3, since the collocated block 3 is within the current CTU and has not yet been reconstructed, reference block 3 can be found from the reference sample memory. The reference sample memory still stores the samples of reference block 3 from the left CTU and can be accessed quickly to retrieve the samples of reference block 3, and reference block 3 can be used to reconstruct the current block in the intra-block copy mode in one embodiment.

[0105] In other embodiments, for reference block 1, the reconstruction of the collocated block 1 within the current CTU is completed, and thus the reference sample memory stores the samples of the collocated block 1. Further, the samples of reference block 1 are stored in off-chip storage that has a relatively high latency, for example, compared to the reference sample memory. Thus, in one embodiment, reference block 1 is not found in the reference sample memory, and in one embodiment, reference block 1 cannot be used to reconstruct the current block in the intra-block copy mode.

[0106] Similarly, in other embodiments, for reference block 2, a part of the collocated block 2 has been reconstructed, and thus the reference sample memory is updated to store the samples of the collocated block 2. Thus, in one embodiment, reference block 2 cannot be a valid reference block for reconstructing the current block in the intra-block copy mode.

[0107] Generally, in the intra-block copy mode, for reference blocks in a previously decoded CTU, if the collocated blocks within the current CTU have not yet been reconstructed, the samples of the reference blocks are available in the reference sample memory, and the samples of the reference blocks can be read from the reference sample memory for use as a reference for reconstruction in the intra-block copy mode.

[0108] It should be noted that in the above embodiments, the samples at the upper left corner of the collocated block within the current CTU, also referred to as the samples at the upper left corner of the reference block, are checked. If the collocated samples within the current CTU have not yet been reconstructed, all the remaining samples of the reference block are available for use as a reference for the intra-block copy.

[0109] Also, in the above embodiments, it should be noted that the memory size of the reference sample memory is the size of one CTU, and the previously decoded CTU means the CTU adjacent to the left of the current CTU.

[0110] According to one aspect of the disclosure, the memory size of the reference sample memory may be larger than the size of one CTU.

[0111] FIG. 12 shows an example of collocated blocks according to some embodiments of the present disclosure. In the example of FIG. 12, the reference sample memory is configured to have a size that is N times (N is an integer greater than or equal to 2) the size of a CTU, and thus the reference sample memory can store reconstructed samples from N + 1 CTUs. For example, as shown in FIG. 12, the current CTU (1210), the leftmost CTU (1230), and the (plural) intermediate left CTUs (1220) etc. between the current CTU (1210) and the leftmost CTU (1230). The number of left CTUs is equal to N - 1. The reconstructed samples of the current CTU (1210) and the leftmost CTU (1230) are stored in the reference sample memory in a complementary manner. To store the reconstructed sample of the current CTU (1210) in the reference sample memory, the reconstructed sample of the current CTU (1210) is written at the position of the collocated sample of the leftmost CTU (1230). The reconstructed area is shown in gray and the area to be reconstructed is shown in white. In some embodiments, the left CTUs are numbered from the left adjacent CTU to the leftmost CTU. For example, the left adjacent CTU to the current CTU (1210) is the first CTU on the left, the leftmost CTU (1230) is the Nth CTU on the left, and the CTU to the right of the leftmost CTU is the (N - 1)th CTU on the left.

[0112] In the embodiment of FIG. 12, the reference sample memory size is N times the size of the CTU. Therefore, it should be noted that all the samples of the (plural) leftmost intermediate CTUs (1220) are available during the reconstruction of the samples in the current CTU (1210). However, the samples of the leftmost CTU (1230) are only partially available in the reference sample memory and are under similar limitations as shown in the embodiment of FIG. 11. For example, among the current blocks in the intra-block copy mode during decoding, three examples of the reference blocks in the leftmost CTU (1230) are shown as reference block 1, reference block 2, and reference block 3. FIG. 12 also shows collocated block 1 for reference block 1, collocated block 2 for reference block 2, and collocated block 3 for reference block 3. In this case, the x-coordinate offset between a pair of collocated blocks or samples is N times the CTU width. Other descriptions of FIG. 12 are the same as those of FIG. 11 and have been described above, so they are omitted here for clarity.

[0113] It should be noted that in some embodiments, the reconstructed sample portion of the leftmost CTU (1230) and the reconstructed sample portion of the current CTU (1210) are complementary in the reference sample memory. When the samples of the current CTU (1210) are reconstructed, the reconstructed samples of the current CTU (1210) are written into the reference sample memory at the positions of the reconstructed samples of the leftmost CTU (1230). The reconstructed samples of the leftmost CTU (1230) are overwritten in the reference sample memory by the reconstructed samples of the current CTU (1210).

[0114] Note that the example of FIG. 12 can be used in scenarios where the reference sample memory size is greater than or equal to the CTU size. For example, if the reference sample memory size is 4 times the CTU size, the (multiple) leftmost intermediate CTUs (1220) include 3 CTUs, and if the reference sample memory size is 16 times the CTU size, the (multiple) leftmost intermediate CTUs (1220) include 15 CTUs. For example, if the reference sample memory size is equal to the CTU size, there is no leftmost intermediate CTU (1220).

[0115] Aspects of the present disclosure provide techniques for adjusting the search range in the intra-block copy mode when the current CTU size is less than the maximum CTU size and the reference sample memory size is equal to 1 maximum CTU size. The reference sample memory can then buffer a plurality of previously decoded CTUs. If the reference block is from the left adjacent CTU, no additional condition checks are required for availability. In this case, all left CTUs are available for intra-block copy reference. In some embodiments, a unified condition check is used for the case where the current CTU size is equal to the maximum CTU size and for the case where the current CTU size is less than the maximum CTU size.

[0116] Various parameters are used in the following description of the embodiments.

[0117] MaxCtbLog2SizeY indicates in the log2 domain the maximum CTU size allowed for one side (height or width) when the CTU is square. For example, if the maximum allowable CTU is 128×128 luma samples in height, MaxCtbLog2SizeY is equal to 7.

[0118] CtbLog2SizeY indicates in the log2 domain the CTU size for one side (height or width) when the CTU is square.

[0119] cbHeight represents the height of the coding block (also referred to as the current block), and cbWidth represents the width of the coding block. The position of the upper left corner of the coding block is indicated by (xCb, yCb), the position of the upper right corner of the coding block is indicated by (xCb + cbWidth - 1, yCb), and the position of the lower left corner of the coding block is indicated by (xCb, yCb + cbHeight - 1).

[0120] mvL0 represents a block vector. mvL0[0] represents the x component of the block vector mvL0 at 1 / 16 pel resolution, and mvL0[1] represents the y component of the block vector mvL0 at 1 / 16 pel resolution. Therefore, the integer value of the x component is obtained by right-shifting mvL0[0] by 4 bits, and the integer value of the y component is obtained by right-shifting mvL0[1] by 4 bits.

[0121] According to some aspects of the present disclosure, to determine whether mvL0 is a valid block vector pointing to a reference block in the intra-block copy mode and whether the reference block is completely stored in the reference sample memory, the block vector mvL0 is restricted using a two-stage restriction process. Therefore, the reference sample memory can be accessed without reading the reference samples from the main memory (e.g., off-chip memory).

[0122] In the first stage, the block vector mvL0 is checked to determine whether the block vector is within a potential valid block vector that refers to a CTU-based search range including any CTU having all or part of the reconstructed samples in the reference sample memory. For example, the CTU-based search range includes the current CTU, the leftmost CTU, and any left CTU between the current CTU and the leftmost CTU. In one embodiment, the CTU size is equal to the maximum allowable CTU size, the CTU-based search range includes the current CTU and the leftmost CTU (referred to as the left CTU in FIG. 11), and there is no intermediate left CTU between the current CTU and the leftmost CTU.

[0123] If the block vector mvL0 is a potential valid block vector, the block vector mvL0 is further checked in the second stage to determine whether the block vector is a valid block vector. For example, in the second stage, if the block vector refers to a reference block within an intermediate left CTU, the block vector is a valid block vector. In the second stage, if the block vector refers to a reference block within the leftmost CTU, the block vector mvL0 is further checked to determine whether the collocated block of the reference block has been reconstructed.

[0124] In the first embodiment, the CTU size is equal to the maximum allowable CTU size, and thus the reconstructed samples stored in the reference sample memory are from either the current CTU or the leftmost CTU (the left CTU in FIG. 11).

[0125] In the first stage, mvL0 is checked to determine whether the block vector refers to a CTU-based search range including the current CTU and the leftmost CTU. In some examples, CTUs with the same y value form a CTU row, and CTUs with the same x value form a CTU column. In some embodiments, the restrictions on the block vector mvL0 are represented by equations (1) to (4): (yCb + (mvL0[1] >> 4)) >> CtbLog2SizeY = yCb >> CtbLog2SizeY (Equation 1) (yCb + (mvL0[1] >> 4) + cbHeight - 1) >> CtbLog2SizeY = yCb >> CtbLog2SizeY (Equation 2) (xCb + (mvL0[0] >> 4)) >> CtbLog2SizeY ≥ (xCb >> CtbLog2SizeY) - 1 (Equation 3) (xCb + (mvL0[0] >> 4) + cbWidth - 1) >> CtbLog2SizeY ≤ (xCb >> CtbLog2SizeY) (Equation 4)

[0126] When Equation (1) is satisfied, the top of the reference block is in the same CTU row as the current block. When Equation (2) is satisfied, the bottom of the reference block is in the same CTU row as the current block. When Equation (3) is satisfied, the left side of the reference block is in the same CTU column as the current block, or in an example, in the CTU column immediately to the left of the current block. When Equation (4) is satisfied, the right side of the reference block is in the same CTU column as the current block, or in an example, in the CTU column to the left of the current block. Thus, in the embodiment, when Equations 1 to 4 are satisfied, the reference block is within the CTU-based search range.

[0127] In the second stage, when the reference block is within the leftmost CTU, for example, when Equation (5) is satisfied, the block vector mvL0 is checked and it is determined whether the collocated block of the reference block has been reconstructed. (xCb + (mvL[0] >> 4)) >> CtbLog2SizeY = (xCb >> CtbLog2SizeY) - 1 (Equation 5)

[0128] In the embodiment, The upper left corner of the current block is at (xCb, yCb), where the upper left corner of the reference block is at (xCb + (mvL[0] >> 4), yCb + (mvL0[1] >> 4)), and the upper left corner of the collocated block is at (xCb + (mvL[0] >> 4) + (1 << CtbLog2SizeY), yCb + (mvL0[1] >> 4)). In the example, the upper left corner of the collocated block is used to check a map that tracks the image reconstruction process. Note that if the position on the map is "false", in one example it indicates that the collocated block has not been reconstructed, and thus the reference block is available in the reference sample memory and can be used to reconstruct the current block in the intra-block copy mode. In that case, the block vector is a valid block vector. However, if the position on the map is "true", in one example it indicates that at least a part of the collocated block has been reconstructed, and thus the samples of the collocated block are stored in the reference sample memory at the positions of the samples of the reference block, and the samples of the reference block are not available in the reference sample memory. Therefore, the block vector mvL0 is not a valid block vector.

[0129] In the first embodiment, the second stage is based on the collocated block, and the process of the first embodiment is referred to as a coding unit (CU)-based update process.

[0130] In the second embodiment, the second stage is based on the block region. For example, when the CTU has 128×128 samples, the CTU can be divided into four block regions each having 64×64 samples. Similarly, in the second embodiment, the CTU size is equal to the maximum allowable CTU size, and thus the reconstructed samples stored in the reference sample memory are by either the current CTU or the leftmost CTU farthest away (the leftmost CTU in FIG. 11).

[0131] In the first stage, the block vector mvL0 is checked to determine whether the block vector mvL0 indicates a CTU-based search range that includes the current CTU and the leftmost CTU farthest away, by using equations (1) to (4) in a manner similar to the first embodiment.

[0132] In the second stage, when Equation 5 is satisfied, for example, when the reference block is within the leftmost CTU farthest away, if Equation (5) is satisfied, the block vector mvL0 is checked, and it is determined whether the collocated block region (e.g., a 64×64 block region) for the reference block has been reconstructed.

[0133] In an example, the upper left corner of the current block is at (xCb, yCb), where the upper left corner of the reference block is at (xCb+(mvL[0]>>4), yCb+(mvL0[1]>>4)), and the upper left corner of the collocated block region for the reference block is at ((((xCb+(mvL[0]>>4)+(1<<CtbLog2SizeY))>>(CtbLog2SizeY-1))<<(CtbLog2SizeY-1)), (((yCb+(mvL0[1]>>4))>>(CtbLog2SizeY-1))<<(CtbLog2SizeY-1))). In the example, the upper left corner of the collocated block region is used to check a map that tracks the image reconstruction process. If the position on the map is "false", in one example, it indicates that the collocated block region has not been reconstructed, so the reference block is available in the reference sample memory and can be used to reconstruct the current block in the intra-block copy mode. In that case, the block vector is a valid block vector. However, if the position on the map is "true", it indicates that the collocated block region has been reconstructed or partially reconstructed, and in that case, the block vector mvL0 is not a valid block vector.

[0134] In the second embodiment, the second stage is based on the collocated blocks, and the process of the second embodiment is referred to as a block region-based update process. In the second embodiment, when any sample in the 64×64 block region (collocated block region) within the current CTU is reconstructed, the corresponding region in the reference sample memory to which the collocated sample (of the current sample) belongs is not available for intra-block copy reference.

[0135] According to one aspect of the present disclosure, a change in the CTU size usually occurs as the width and / or height doubles or is halved. For example, when the width and height are halved, samples for 4 CTUs can be stored in the reference sample memory of 1 maximum CTU size.

[0136] In some embodiments, when the current CTU size CtbLog2SizeY in the Log2 domain is used, the number of CTUs in which reference sample data can be stored in the reference sample memory buffer is variable depending on the relationship between MaxCtbLog2SizeY and CtbLog2SizeY.

[0137] In the third embodiment, the CTU size is CtbLog2SizeY in the Log2 domain, and thus the reconstructed samples stored in the reference sample memory are by any one of the current CTU, the leftmost CTU farthest away, and the (multiple) intermediate left CTUs located between the current CTU and the leftmost CTU farthest away.

[0138] In the first stage, block mvL0 is checked to determine whether the block vector mvL0 points to a CTU-based search range including the current CTU, the leftmost CTU farthest away, and the (multiple) intermediate left CTUs. In some examples, the restrictions on the block vector mvL0 are represented by Equations (6) to (9): (yCb+(mvL0[1]>>4))>>CtbLog2SizeY=yCb>>CtbLog2SizeY Equation (6) (yCb + (mvL0[1] >> 4) + cbHeight - 1) >> CtbLog2SizeY = yCb >> CtbLog2SizeY, Equation (7) (xCb + (mvL0[0] >> 4)) >> CtbLog2SizeY >= (xCb >> CtbLog2SizeY) - 1 << (2 * (MaxCtbLog2SizeY - CtbLog2SizeY)), Equation (8) (xCb + (mvL0[0] >> 4) + cbWidth - 1) >> CtbLog2SizeY <= (xCb >> CtbLog2SizeY), Equation (9)

[0139] When Equation (6) is satisfied, the top of the reference block is in the same CTU row as the current block. When Equation (7) is satisfied, the bottom of the reference block is in the same CTU row as the current block. When Equation (8) is satisfied, the left side of the reference block is in the same CTU column as one of the current CTU, (a plurality of) intermediate left CTUs, and the farthest left CTU. When Equation (9) is satisfied, the right side of the reference block is in the same CTU column as one of the current CTU or the CTU to the left of the current CTU (e.g., (a plurality of) intermediate left CTUs, the farthest left CTU, etc.). Thus, in the embodiment, when Equations (6) to (9) are satisfied, the reference block is within the CTU-based search range.

[0140] When the block vector mvL0 is a potential valid block vector, the block vector mvL0 is further checked in the second stage to determine whether the block vector is a valid block vector. For example, in the second stage, when the block vector points to a reference block within the intermediate left CTU, for example, if Equation (10) is not satisfied, the block vector is a valid block vector. In the second stage, when the reference block is within the farthest left CTU, for example, if Equation (10) is satisfied, the block vector mvL0 is further checked to determine whether the collocated block of the reference block has been reconstructed. (xCb+(mvL[0]>>4))>>CtbLog2SizeY=(xCb>>CtbLog2SizeY)-(1<<(2*(MaxCtbLog2SizeY-CtbLog2SizeY))), Equation (10)

[0141] In an embodiment In the current block, the upper left corner is at (xCb, yCb). At this time, the upper left corner of the reference block is at (xCb+(mvL[0]>>4), yCb+(mvL0[1]>>4)), and the upper left corner of the block collocated with the reference block is at (xCb+(mvL[0]>>4)+(1<<4^(MaxCtbLog2SizeY-CtbLog2SizeY)), yCb+(mvL0[1]>>4)). In one example, the upper left corner of the collocated block is used to check a map that tracks the image reconstruction process. If the position on the map is "false", in one example, it indicates that the collocated block has not been reconstructed. Therefore, the reference block is available in the reference sample memory and can be used to reconstruct the current block in the intra-block copy mode. In that case, the block vector is a valid block vector. However, if the position on the map is "true", in one example, it indicates that at least a part of the collocated block has been reconstructed. Therefore, the samples of the collocated block are stored in the reference sample memory at the positions of the samples of the reference block, and the samples of the reference block are not available in the reference sample memory. Therefore, the block vector mvL0 is not a valid block vector.

[0142] In the third embodiment, the second stage is based on the collocated block, and the process of the third embodiment is referred to as a CU-based update process.

[0143] In the fourth embodiment, the second stage is based on block regions. For example, the CTU size is CtbLog2SizeY in the Log2 domain, and when the block region size is CtbLog2SizeY-1 in the Log2 domain, the CTU can be divided into four block regions of the same size. Similarly, in the fourth embodiment, the reconstructed samples stored in the reference sample memory are from the current CTU, the leftmost CTU, and the intermediate left CTU between the current CTU and the leftmost CTU.

[0144] In the first stage, check the block vector mvL0 to determine whether the block vector mvL0 is a potentially valid block vector that refers to a CTU-based search range including the current CTU, the leftmost CTU, and the (multiple) intermediate left CTUs between the current CTU and the leftmost CTU, using equations (6) to (9) in a manner similar to the third embodiment.

[0145] If the block vector mvL0 is a potentially valid block vector, the block vector mvL0 is further checked in the second stage to determine whether the block vector is a valid block vector. For example, in the second stage, if the block vector refers to a reference block within the intermediate left CTU, the block vector is a valid block vector if equation (10) is not satisfied. In the second stage, when the reference block is within the leftmost CTU, for example, if equation (10) is satisfied, the block vector mvL0 is checked to determine whether the block collocated with the reference block has been reconstructed.

[0146] In an embodiment, the upper left corner of the current block is at (xCb, yCb). At this time, the upper left corner of the reference block is at (xCb + (mvL[0] >> 4), yCb + (mvL0[1] >> 4)), and the upper left corner of the collocated block area for the reference block is at ((((xCb + (mvL[0] >> 4) + (1 << 4 ^ (MaxCtbLog2SizeY - CtbLog2SizeY))) >> (CtbLog2SizeY - 1)) << (CtbLog2SizeY - 1), ((yCb + (mvL0[1] >> 4)) >> (CtbLog2SizeY - 1)) << (CtbLog2SizeY - 1)). In the example, the upper left corner of the collocated block area is used to check a map that tracks the image reconstruction process. Note that when the position on the map is "false", in one example, it indicates that the collocated block area has not been reconstructed. Therefore, the reference block is available in the reference sample memory and can be used to reconstruct the current block in the intra-block copy mode. In that case, the block vector is a valid block vector. However, when the position on the map is "true", it indicates that the collocated block area has been reconstructed or partially reconstructed. In that case, the block vector mvL0h is not a valid block vector.

[0147] In the fourth embodiment, the second stage is based on the collocated block, and the process of the fourth embodiment is referred to as a block area-based update process. In the fourth embodiment, when any sample of the collocated block area within the current CTU (with respect to the reference block) has been reconstructed, the corresponding area in the reference sample memory that stores the samples of the collocated block area is not available for intra-block copy reference.

[0148] According to some aspects of the present disclosure, the leftmost CTU is intentionally excluded from the search region, and then the block vector mvL0 is restricted using a one-step restriction process to determine whether mvL0 is a valid block vector pointing to a reference block in the intra-block copy mode and whether the reference block is fully stored in the reference sample memory.

[0149] In the fifth embodiment, similar to the third embodiment, the CTU size is CtbLog2SizeY in the Log2 domain, and thus the reconstructed samples stored in the reference sample memory are by any one of the current CTU, the leftmost CTU farthest away, and the (plural) intermediate left CTUs located between the current CTU and the leftmost CTU farthest away. In the fifth embodiment, the block mvL0 is checked to determine whether the block vector mvL0 points to a search range including the current CTU, the leftmost CTU farthest away, and the (plural) intermediate left CTUs. It should be particularly noted that in the fifth embodiment, the left CTU is excluded from the search range. In some examples, the restrictions on the block vector mvL0 are represented by Equations (11) to (14): (yCb+(mvL0[1]>>4))>>CtbLog2SizeY=yCb>>CtbLog2SizeY Equation (11) (yCb+(mvL0[1]>>4)+cbHeight-1)>>CtbLog2SizeY=yCb>>CtbLog2SizeY Equation (12) (xCb+(mvL0[0]>>4))>>CtbLog2SizeY>(xCb>>CtbLog2SizeY)-1<<(2*(MaxCtbLog2SizeY-CtbLog2SizeY)) Equation (13) (xCb+(mvL0[0]>>4)+cbWidth-1)>>CtbLog2SizeY>(xCb>>CtbLog2SizeY)-1<<(2*(MaxCtbLog2SizeY-CtbLog2SizeY)) Equation (14)

[0150] When Equation (11) is satisfied, the top of the reference block is in the same CTU row as the current block. When Equation (12) is satisfied, the bottom of the reference block is in the same CTU row as the current block. When Equation (13) is satisfied, the left side of the reference block is in the same CTU column as the current CTU or one of the (multiple) intermediate left CTUs. When Equation (14) is satisfied, the right side of the reference block is in the same CTU column as the current CTU or one of the (multiple) intermediate left CTUs. In that case, in the embodiment, when Equations 11 to 14 are satisfied, the reference block is within the search range, and the block vector mvL0 is a valid block vector.

[0151] FIG. 13 shows a flowchart showing an overview of a process (1300) according to an embodiment of the present disclosure. The process (1300) can be used for reconstructing a block coded in an intra mode, and thus generates a prediction block for the block being reconstructed. In various embodiments, the process (1300) is executed by a processing circuit in the terminal devices (310), (320), (330) and (340), a processing circuit that executes the functions of the video encoder (403), a processing circuit that executes the functions of the video decoder (410), a processing circuit that executes the functions of the video decoder (510), a processing circuit that executes the functions of the video encoder (603), and the like. In some embodiments, the process (1300) is implemented by software instructions, and thus when the processing circuit executes the software instructions, the processing circuit executes the process (1300). The process starts from (S1301) and proceeds to (S1310).

[0152] In (S1310), the prediction information of the current block in the current CTU is decoded from the coded video bitstream. The prediction information represents the intra block copy mode. The size of the current CTU is smaller than the maximum size corresponding to the storage capacity of the reference sample memory. In some embodiments, the reference sample memory has an access speed faster than that of the main memory for storing samples reconstructed from the coded video bitstream. For example, the reference sample memory is an on-chip memory on the same chip as the decoder circuit, and the main memory is an off-chip memory outside the chip having the decoder circuit. Note that the reference sample memory can be implemented using off-chip memory in some examples.

[0153] In (S1320), a block vector is determined. The block vector points to a reference block within the same image as the current block, and the reference block is reconstructed from samples buffered in the reference sample memory. In some embodiments, the search region is defined to include CTUs having reconstructed samples buffered in the reference sample memory, such as the current CTU, the middle left CTU, and the farthest left CTU. The farthest left CTU has at least one reconstructed sample overwritten with the reconstructed samples of the current CTU in the reference sample memory.

[0154] In (S1330), the current block is reconstructed based on the reconstructed samples of the reference block read from the reference sample memory. For example, the reference sample memory is accessed to read the reconstructed samples of the reference block, and then the samples of the current block are reconstructed based on the reconstructed samples read from the reference sample memory. Thereafter, the process proceeds to (S1399) and ends.

[0155] The above technology can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 14 shows a computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter.

[0156] The computer software can be coded using any suitable machine code or computer language that can be the subject of an assembly, compilation, linking, or similar mechanism, and can be executed directly or create code that includes instructions that can be executed via implementation, microcode execution, etc. by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0157] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, the Internet of Things, etc.

[0158] The components shown in FIG. 18 for the computer system (1400) are of an exemplary nature and are not intended to imply any limitations regarding the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Also, the configuration of the components should not be construed as having dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system (1400)

[0159] The computer system (1400) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, switching, movement of a data glove), voice input (e.g., voice, clapping), visual input (e.g., gesture), olfactory input (not shown). Also, the human interface device may be used to capture specific media that is not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., 2D video, 3D video including stereoscopic images).

[0160] The input human interface device may include one or more of a keyboard (1401), a mouse (1402), a trackpad (1403), a touch screen (1410), a data glove (not shown), a joystick (1405), a microphone (1806), a scanner (1407), a camera (1408).

[0161] The computer system (1400) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may be, for example, tactile output devices (e.g., tactile feedback by a touch screen (1410), a data glove (not shown), or a joystick (1405), although it can also be a tactile feedback device that does not function as an input device), audio output devices (e.g., speakers (1409), headphones (not shown)), visual output devices (e.g., a screen (1410) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each with or without tactile feedback capabilities, some of which can enable 2D visual output or output of three or more dimensions through means such as stereoscopic output like virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown), may be included.

[0162] The computer system (1400) can also include human-accessible memory devices and their accessible media. The media can include, for example, an optical media drive (1420) including CD / DVD ROM / RW by media such as CD / DVD (1421), a USB memory (1422), a removable hard drive or a solid state drive (1423), conventional magnetic media such as tapes, floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles, etc.

[0163] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0164] The computer system (1400) can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include Ethernet, wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., cellular networks including cable TV, satellite TV, and terrestrial broadcast TV, industrial and vehicle including CANBus. A particular network generally requires an external network interface adapter (e.g., a USB port of the computer system (1400)) connected to a particular general-purpose data port or peripheral bus (1449), and other networks are generally integrated into the core of the computer system (1400) by being connected to the system bus described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be unidirectional communication, receive-only (e.g., broadcast TV) communication, unidirectional transmit-only (e.g., CAN bus to a particular CAN bus device) communication, or two-way communication to other computer systems using, for example, local or wide area digital networks. Specific protocols and protocol stacks can be used with each of those networks and network interfaces as described above.

[0165] The aforementioned human interface device, human-accessible storage device, and network interface can be connected to the core (1440) of the computer system (1400).

[0166] The core (1440) can include one or more central processing devices (CPUs) (1441), graphics processing devices (GPUs) (1442), special programmable processing devices in the form of field programmable gate arrays (FPGAs) (1443), hardware accelerators (1444) for specific tasks, etc. These devices can be connected via a system bus (1448) together with read-only memory (ROM) (1445), random access memory (1446), internal mass storage devices, such as internal non-user-accessible hard drives, SSDs, etc. (1447). In some computer systems, the system bus (1448) can be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be directly connected to the system bus (1448) of the core or connected via a peripheral bus (1449). The architecture of the peripheral bus includes PCI, USB, etc.

[0167] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can be combined to execute specific instructions that can constitute the above-mentioned computer code. The computer code can be stored in the ROM (1445) or RAM (1446). Migration data can also be stored in the RAM (1446), but permanent data can be stored, for example, in the internal mass storage device (1447). By using cache memory that can be closely associated with one or more CPUs (1441), GPUs (1442), mass storage devices (1447), ROM (1445), RAM (1446), etc., fast storage and retrieval to any of the memory devices can be enabled.

[0168] A computer-readable medium can have computer code thereon for performing various computer-executable operations. The media and the computer code can be those specially designed and created for the present disclosure, or they can be of the kind that are well-known and available to those having the skills in computer software arts.

[0169] As an example, and not by way of limitation, a computer system having an architecture (1400), specifically a core (1440), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with a mass storage device accessible by a user as described above, as well as a specific storage device of the core (1440) of a non-transitory nature such as a core-internal mass storage device (1447) or a ROM (1445). The software implementing various embodiments of the present disclosure can be stored on such a device and executed by the core (1440). A computer-readable medium can include one or more memory devices or chips, depending on specific needs. Software can cause a core (1440) and specifically processors (including CPUs, GPUs, FPGAs, etc.) therein to define data structures stored in a RAM (1446) and modify such data structures according to processes defined by the software, thereby causing the computer system to execute specific processes or specific portions described herein. Additionally or alternatively, a computer system can provide functionality as a result of logic wired in a circuit (e.g., an accelerator (1444)) or otherwise embodied, which can operate instead of or in conjunction with software to execute specific processes or specific portions of specific processes described herein. References to software include logic and, if necessary, vice versa. References to a computer-readable medium can include a circuit (such as an integrated circuit (IC)) that stores software for execution, a circuit that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0170] Appendix A: Abbreviations JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplemental Enhancement Information VUI: Video User Information GOP: Group of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal-to-Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long-Term Evolution CAN bus: Controller Area Network bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field-Programmable Gate Array SSD: Solid State Drive IC: Integrated Circuit CU: Coding Unit

[0171] Although this disclosure describes some exemplary embodiments, there are changes, substitutions, and various equivalents that fall within the scope of the present invention. Therefore, those skilled in the art will understand that although not explicitly shown or described herein, they can implement the principles of the present invention and thus create numerous systems and methods that are within the concept and scope thereof.

Claims

[Claim 1] 1. A processor-implemented method for obtaining reference samples used to compute a residual for a current block in a current coding tree unit (CTU) in an encoder, comprising: determining prediction information for a current block in a current coding tree unit (CTU), the prediction information representing an intra block copy mode, a size of the current CTU being smaller than a size of a reference sample memory for storing reconstructed samples, the size of the reference sample memory being equal to a maximum allowed CTU size; determining a block vector pointing to a reference block of the same image as the current block, the reference block having reconstructed samples buffered in the reference sample memory; reconstructing at least one sample of the current block based on the reconstructed samples of the reference block read from the reference sample memory; The method includes:

Citation Information

Patent Citations

  • Intra-block copy-merge mode and padding for ibc reference regions not available

    JP2018530249A

  • Method and apparatus for video coding

    WO2020113156A1