Improvements to IBC Signaling
Patent Information
- Application Number
- CN202580010996.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-03-28
- Publication Date
- 2026-08-18
AI Technical Summary
[0010]因此,公开了用于对视频进行编码和解码的方法以及设备和系统。这样的方法、设备和系统可以补充或替代用于视频编码/解码的常规方法、设备和系统。
Smart Images

Figure CN122603508A_ABST
Abstract
Description
Related applications
[0001] This application is a continuation of U.S. Patent Application No. 19 / 093,050, filed March 27, 2025, which claims priority to U.S. Provisional Patent Application No. 63 / 572,273, filed March 31, 2024, entitled “Improvements on IBC signaling,” each of which is incorporated herein by reference in its entirety. Technical Field
[0002] The disclosed implementations generally relate to video encoding and decoding, including but not limited to systems and methods for IntraBlock Copy (IBC or Intra-BC). Background Technology
[0003] Digital video is supported by a variety of electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital camera devices, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc. Electronic devices send and receive or otherwise transmit digital video data across communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video encoding can be used to compress video data according to one or more video encoding standards before transmission or storage. Video encoding can be performed by hardware and / or software on servers or electronic / client devices providing cloud services.
[0004] Video coding typically uses prediction methods that leverage the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video coding aims to compress video data into a form using a lower bitrate while avoiding or minimizing video quality degradation. Several video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H (Moving Picture Experts Group-H) project. The ITU-T (International Telecommunication Union-Telecommunication Standardization Sector) and ISO / IEC (International Organization for Standardization / International Electrotechnical Commission) released the HEVC / H.265 standard in 2013 (Revision 1), 2014 (Revision 2), 2015 (Revision 3), and 2016 (Revision 4). Versatile Video Coding (VVC / H.266) is a video compression standard designed as a successor to HEVC. The ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (Revision 1) and 2022 (Revision 2), respectively. AOMedia Video 1 (Alliance for Open MediaVideo 1, AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, a validation version 1.0.0 of the specification with errata table 1 was released. Summary of the Invention
[0005] Among other things, this disclosure describes a set of methods for video (image) compression, more specifically relating to conditionally signaling indicators associated with intra-block copying information. For example, the indicators associated with intra-block copying information can be conditionally signaled based on one or more characteristics of the current block, such as the current block size, superblock size, and corresponding search region. By conditionally signaling the indicators associated with intra-block copying information, compression efficiency can be improved when a smaller local search region is determined to be suitable for finding a prediction block. When a smaller local search region is determined to be unsuitable for finding a prediction block, signaling overhead can be reduced by not signaling block copying information, while using other prediction methods to maintain encoding / decoding accuracy.
[0006] According to some embodiments, a method for video decoding is provided. The method includes (i): receiving a video bitstream having a plurality of blocks, the plurality of blocks including a current block; (ii) when one or more features of the current block satisfy one or more criteria: (a) parsing one or more indicators from the video bitstream, the one or more indicators indicating IBC information; (b) identifying a predicted block of the current block based on the IBC information; and (c) reconstructing the current block using the predicted block; and (iii) when one or more features of the current block do not satisfy one or more criteria, decoding the current block without parsing one or more indicators indicating IBC information.
[0007] According to some implementations, a video encoding method includes: (i) receiving video data comprising a plurality of blocks, the plurality of blocks including a current block; (ii) when one or more features of the current block satisfy one or more criteria: (a) encoding the current block using an IBC mode; and (b) signaling one or more indicators in the video bitstream, the one or more indicators indicating IBC information of the current block; and (iii) when one or more features of the current block do not satisfy one or more criteria, not signaling IBC information of the current block.
[0008] According to some embodiments, a computing system, such as a streaming system, server system, personal computer system, or other electronic device, is provided. The computing system includes a control circuitry system and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).
[0009] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions executable by a computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.
[0010] Therefore, methods, apparatus, and systems for encoding and decoding video are disclosed. Such methods, apparatus, and systems can supplement or replace conventional methods, apparatus, and systems for video encoding / decoding.
[0011] The features and advantages described in this specification are not necessarily all included, and in particular, some additional features and advantages will be apparent to those skilled in the art in light of the accompanying drawings, specification, and claims provided in this disclosure. Furthermore, it should be noted that the language used in this specification has been chosen primarily for readability and instructional purposes and is not necessarily intended to depict or limit the subject matter described herein. Attached Figure Description
[0012] To provide a more detailed understanding of this disclosure, a more specific description can be made by referring to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the drawings only show relevant features of this disclosure and are therefore not necessarily intended to be limiting, as those skilled in the art will understand upon reading this disclosure that other valid features may be permissible.
[0013] Figure 1 This is a block diagram illustrating an example communication system according to some implementations.
[0014] Figure 2A This is a block diagram illustrating example elements of an encoder component according to some embodiments.
[0015] Figure 2B This is a block diagram illustrating example elements of a decoder component according to some embodiments.
[0016] Figure 3 This is a block diagram illustrating an example server system according to some implementation methods.
[0017] Figure 4A An example intra-frame block copying technique according to some implementations is shown.
[0018] Figure 4B Examples of different intra-block copying modes according to some implementation methods are shown.
[0019] Figure 5A An example video decoding process according to some implementation methods is shown.
[0020] Figure 5B An example video encoding process according to some implementation methods is shown.
[0021] By convention, the various features shown in the accompanying drawings are not necessarily drawn to scale, and similar reference numerals may be used to indicate similar features throughout the specification and the drawings. Detailed Implementation
[0022] This disclosure describes video / image compression techniques, including conditionally signaling indicators associated with intra-block copying information (e.g., IBC enable / disable flags, block vectors, block vector prediction values, etc.) based on one or more features of the current block. Searching for the prediction block within a smaller local search region can improve the coding efficiency of the intra-block copying method. However, when one or more features of the current block indicate a reduced probability of finding a suitable prediction block within a smaller local search region, the encoder may forgo signaling flags and syntax associated with the intra-block copying method. If a larger global search region is available to find the prediction block, the intra-block copying method can still be used to encode the current block even if the encoder has not signaled the intra-block copying flags. Otherwise, the intra-block copying method is not used to encode the current block. The advantage of conditionally signaling indicators associated with intra-block copying information based on one or more features of the current block is that it reduces signaling overhead (e.g., using fewer bits when it is determined that the local search region is not suitable for finding the prediction block), thereby improving coding efficiency. Using a smaller local search region to find the predicted block based on one or more features of the current block improves the coding efficiency of the intra-block copying method.
[0023] Example systems and devices
[0024] Figure 1 This is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0025] Source device 102 includes a video source 104 (e.g., a camera device component or media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera device (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from video source 104 can have a higher data volume compared to the encoded video bitstream 108 generated by encoder component 106. Because the encoded video bitstream 108 has a lower data volume (less data) compared to the video stream from video source 104, it requires less bandwidth to transmit and less storage space to store compared to the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., configured to transmit uncompressed video to network 110).
[0026] One or more networks 110 represent any number of networks that transmit information between source device 102, server system 112, and / or electronic device 120, including, for example, wired (connected) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet.
[0027] One or more networks 110 include server systems 112 (e.g., distributed / cloud computing systems). In some embodiments, server system 112 is a streaming server or includes streaming servers (e.g., configured to store and / or distribute video content, such as encoded video streams from source device 102). Server system 112 includes codec components 114 (e.g., configured to encode and / or decode video data). In some embodiments, codec components 114 include encoder components and / or decoder components. In various embodiments, codec components 114 are instantiated as hardware, software, or a combination of hardware and software. In some embodiments, codec components 114 are configured to decode encoded video bitstream 108 and re-encode video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, server system 112 is configured to generate multiple video formats and / or encodings based on encoded video bitstream 108. In some embodiments, server system 112 is used as a Media-Aware Network Element (MANE). For example, server system 112 can be configured to trim the encoded video bitstream 108 to customize potentially different bitstreams for one or more electronic devices 120. In some implementations, the MANE is provided separately from server system 112.
[0028] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode encoded video data 116 to generate an outgoing video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., communicatively coupled to an external display device and / or includes a media storage device). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.
[0029] The source device and / or multiple electronic devices 120 are sometimes referred to as “terminal devices” or “user devices”. In some embodiments, one or more of the electronic devices 120 and / or the source device 102 are instances of server systems, personal computers, portable devices (e.g., smartphones, tablets, or laptops), wearable devices, video conferencing equipment, and / or other types of electronic devices.
[0030] In an example operation of communication system 100, source device 102 transmits an encoded video bitstream 108 to server system 112. For example, source device 102 may encode a stream of images captured by the source device. Server system 112 receives the encoded video bitstream 108 and may decode and / or encode the encoded video bitstream 108 using codec component 114. For example, server system 112 may apply encoding to the video data that is better suited for network transmission and / or storage. Server system 112 may transmit encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display video images.
[0031] Figure 2A This is a block diagram illustrating example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 YCrCb or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0 or YCrCb 4:4:4). In some embodiments, the video source 104 is a storage device storing previously captured / prepared video. In some embodiments, the video source 104 is a camera device that captures local image information as a video sequence. Video data can be provided as multiple individual images that are given motion when viewed sequentially. Each image can be organized as a spatial array of pixels, where, depending on the sampling structure, color space, etc., each pixel may include one or more samples. The relationship between pixels and samples will be readily understood by those skilled in the art.
[0032] Encoder component 106 is configured to encode and / or compress images of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, encoder component 106 is configured to perform a conversion between the source video sequence and a visual media data bitstream (e.g., a video bitstream). Implementing an appropriate encoding rate is a function of controller 204. In some embodiments, controller 204 controls and is functionally coupled to other functional units as described below. Parameters set by controller 204 may include rate control-related parameters (e.g., image skipping, quantizer and / or rate-distortion optimization technique λ value), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. Other functions of controller 204 can be readily identified by those skilled in the art, as these functions may belong to encoder component 106 optimized for a particular system design.
[0033] In some implementations, encoder component 106 is configured to operate within an encoding / decoding loop. In a simplified example, the encoding / decoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder 210. Decoder 210 reconstructs the symbols in a manner similar to that of the (remote) decoder to create sample data (assuming lossless compression between the symbols and the encoded video bitstream). The reconstructed sample stream (sample data) is input to a reference image memory 208. Since decoding the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference image memory 208 are also bit-accurate between the local and remote encoders. In this way, the encoder's prediction portion interprets the same sample values as the sample values that the decoder interprets during prediction as reference image samples.
[0034] The operation of decoder 210 can be combined with a remote decoder, for example, as shown below. Figure 2B The operation of the decoder component 122 is the same as described in the detailed description. However, a brief reference is provided. Figure 2B Since the symbols are available and the encoding of the symbols into an encoded video sequence by the entropy encoder 214 and the decoding of the symbols by the parser 254 can be lossless, the entropy decoding portion of the decoder component 122, including the buffer memory 252 and the parser 254, does not need to be fully implemented in the local decoder 210.
[0035] Besides parsing / entropy decoding, the decoder techniques described in this paper can exist in the corresponding encoders in essentially the same form. For this reason, the subject matter focuses on decoder operations. Furthermore, the description of encoder techniques can be simplified, as encoder techniques are inverses of decoder techniques.
[0036] As part of the operation of the source encoder 202, the source encoder 202 can perform motion-compensated predictive coding, which predictively encodes the input frame with reference to one or more previously encoded frames from the video sequence designated as reference frames. In this way, the encoding engine 212 encodes the difference between pixel blocks of the input frame and pixel blocks of the reference frame, which can be selected as the prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding the video data.
[0037] Decoder 210 decodes encoded video data based on symbols created by source encoder 202 that can be designated as reference frames. The operation of encoding engine 212 can advantageously handle lossy processing. When encoded video data is processed by video decoder (… Figure 2A When decoded at a location (not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. Decoder 210 replicates the decoding process performed on the reference frame by a remote video decoder, and the reconstructed reference frame can be stored in reference image memory 208. In this way, encoder component 106 locally stores a copy of the reconstructed reference frame that shares common content (no transmission errors) with the reconstructed reference frame to be obtained by the remote video decoder.
[0038] Predictor 206 can perform a prediction search against encoding engine 212. That is, for a new frame to be encoded, predictor 206 can search the reference image memory 208 for sample data (as candidate reference pixel blocks) or specific metadata such as reference image motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new image. Predictor 206 can operate pixel-by-pixel based on the sample blocks to find appropriate prediction references. As determined by the search results obtained by predictor 206, the input image can have prediction references obtained from multiple reference images stored in reference image memory 208.
[0039] The outputs of all the functional units mentioned above can undergo entropy encoding in the entropy encoder 214. The entropy encoder 214 converts the symbols, such as those generated by the various functional units, into an encoded video sequence by lossless compression according to techniques known to those skilled in the art (e.g., Huffman coding, variable-length coding, and / or arithmetic coding).
[0040] In some implementations, the output of entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer encoded video sequences, such as those created by entropy encoder 214, in preparation for transmission via communication channel 218, which can be a hardware / software link to a storage device storing the encoded video data. The transmitter can be configured to combine encoded video data from source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown). In some implementations, the transmitter can transmit additional data along with the encoded video. Source encoder 202 can include such data as part of the encoded video sequence. Additional data can include temporal / spatial / SNR (Signal-to-Noise Ratio, SNR) enhancement layers, other forms of redundant data such as redundant images and slices, Supplementary Enhancement Information (SEI) messages, fragments of Visual Usability Information (VUI) parameter sets, etc.
[0041] Controller 204 can manage the operation of encoder component 106. During encoding, controller 204 can assign a specific type of encoded picture to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, pictures can be assigned as intra-pictures (I-pictures), predictive pictures (P-pictures), or bidirectional predictive pictures (B-pictures). Intra-pictures can be encoded and decoded without using any other frames in the sequence as prediction sources. Some video codecs allow different types of intra-pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are familiar with those variations of I-pictures and their corresponding applications and characteristics, and therefore they will not be repeated here. Predictive pictures can be encoded and decoded using inter-frame prediction or intra-frame prediction that uses at most one motion vector and reference index to predict sample values for each block. Bidirectional predictive pictures can be encoded and decoded using inter-frame prediction or intra-frame prediction that uses at most two motion vectors and reference indexes to predict sample values for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0042] Source images can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded on a block-by-block basis. These blocks can be predictively encoded with reference to other (already encoded) blocks, determined by the coding assignments of the corresponding images applied to the blocks. For example, blocks of image I can be unpredictably encoded, or blocks of image I can be predictively encoded (spatial prediction or intra-frame prediction) with reference to already encoded blocks of the same image. Pixel blocks of image P can be unpredictably encoded with reference to a previously encoded reference image via spatial prediction or temporal prediction. Blocks of image B can be unpredictably encoded with reference to one or two previously encoded reference images via spatial prediction or temporal prediction.
[0043] Video can be captured as multiple source images (video images) in a time-series manner. Intra-frame image prediction (often simply called intra-prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In the example, a specific image during encoding / decoding—referred to as the current image—is segmented into blocks. Where a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference images.
[0044] Encoder component 106 can perform encoding operations according to any predetermined video coding technique or standard described herein. In operation, encoder component 106 can perform various compression operations, including predictive coding operations utilizing temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0045] Figure 2B This is a block diagram illustrating example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 is coupled to channel 218 and display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to loop filter 256 and configured to transmit data to display 124 (e.g., via a wired or wireless connection).
[0046] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence can be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive encoded video data along with other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective user entities (not depicted). The receiver may separate the encoded video sequences from other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. This additional data may be included as part of the encoded video sequence. Decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0047] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra-frame image prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference image memory 266, and a current image memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuit systems. The decoder component 122 can be implemented at least partially in software.
[0048] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to combat network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, a separate buffer memory is provided outside decoder component 122 (e.g., to combat network jitter) in addition to buffer memory 252 inside decoder component 122 (e.g., configured to handle broadcast timing). Buffer memory 252 may not be necessary or may be small when receiving data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network. Buffer memory 252 may be required to make the best use of packet networks such as the Internet. Buffer memory 252 may be relatively large and / or have an adaptive size and may be implemented at least partially in an operating system or similar component outside decoder component 122.
[0049] Parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. Symbols may include, for example, information for managing the operation of decoder component 122, and / or information for controlling rendering devices such as display 124. Control information for the rendering device may be in the form of, for example, Supplemental Enhancement Information (SEI) messages or Video Availability Information (VUI) parameter set fragments (not depicted). Parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow principles known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 254 may extract a subgroup parameter set from the encoded video sequence for pixels in a subgroup of at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to a group. Subgroups may include picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser 254 can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0050] Depending on the type of encoded video picture or a portion thereof (e.g., inter-frame and intra-frame pictures, inter-frame and intra-frame blocks) and other factors, the reconstruction of symbol 270 may involve multiple different units. Which units are involved and how they are involved can be controlled by parser 254 through subgroup control information parsed from the encoded video sequence. For clarity, the flow of such subgroup control information between parser 254 and the following multiple units is not depicted.
[0051] The decoder component 122 can be conceptually subdivided into multiple functional units, and in some implementations, these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, this document retains the conceptual subdivision of functional units.
[0052] The scaler / inverse transform unit 258 receives quantized transform coefficients as symbols 270 and control information (e.g., which transform to use, block size, quantization factor, and / or quantization scaling matrix) from the parser 254. The scaler / inverse transform unit 258 can output blocks including sample values, which can be input to the aggregator 268. In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-frame encoded blocks; that is, blocks that do not use predictive information from previously reconstructed images, but can use predictive information from previously reconstructed portions of the current image. Such predictive information can be provided by the intra-frame image prediction unit 262. The intra-frame image prediction unit 262 can use surrounding reconstructed information obtained from the current (partially reconstructed) image from the current image memory 264 to generate blocks of the same size and shape as the blocks in the reconstruction. The aggregator 268 can add the prediction information already generated by the intra-frame image prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 based on each sample.
[0053] In other cases, the output samples of the scaler / inverse transform unit 258 belong to inter-frame encoded blocks and potentially to motion-compensated blocks. In such cases, the motion compensation prediction unit 260 can access the reference image memory 266 to obtain samples for prediction. After motion compensation of the obtained samples according to the symbols 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to in this case as residual samples or residual signals) to generate output sample information. The address in the reference image memory 266 from which the motion compensation prediction unit 260 obtains the predicted samples can be controlled by motion vectors. Motion vectors can be used by the motion compensation prediction unit 260 in the form of symbols 270, which can have, for example, X components, Y components, and reference image components. Motion compensation can also include, for example, interpolation of sample values obtained from the reference image memory 266 when using subsampled precise motion vectors, and motion vector prediction mechanisms.
[0054] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video bitstream and available to loop filter unit 256 as symbols 270 from parser 254. However, video compression techniques may also respond to metadata obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to sample values obtained from previous reconstruction and loop filtering. The output of loop filter unit 256 may be a sample stream, which can be output to a rendering device such as display 124 and stored in reference picture memory 266 for use in future inter-frame picture prediction.
[0055] Once reconstructed, certain encoded images can be used as reference images for future predictions. Once an encoded image has been reconstructed and that encoded image (e.g., by parser 254) has been identified as a reference image, the current reference image can become part of the reference image memory 266, and a new current image memory can be reallocated before reconstructing subsequent encoded images begins.
[0056] Decoder component 122 can perform decoding operations according to a predetermined video compression technique that may be recorded in any standard such as those described herein. As specified in a video compression technique document or standard, and particularly in a configuration file therein, an encoded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard. Furthermore, to conform to some video compression techniques or standards, the complexity of the encoded video sequence may be within a range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limitations set by the hierarchy can be further restricted by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer used to represent signals in the encoded video sequence.
[0057] Figure 3This is a block diagram illustrating a server system 112 according to some embodiments. The server system 112 includes a control circuitry system 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry system 302 includes one or more processors (e.g., a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and / or a DPU (Data Processing Unit)). In some embodiments, the control circuitry system includes a field-programmable gate array, a hardware accelerator, and / or an integrated circuit (e.g., an application-specific integrated circuit).
[0058] Network interface 304 can be configured to interface with one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). Communication networks can be local, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of communication networks include: local area networks such as Ethernet, wireless LANs; cellular networks including GSM (Global System for Mobile Communications), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; wired or wireless wide area digital networks for TV including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus (Controller Area Network-BUS), etc. Such communication can be one-way receiving (e.g., broadcast TV), one-way transmitting (e.g., to a CANBus device), or bidirectional (e.g., to other computer systems using local area digital networks or wide area digital networks). Such communication can include communication across one or more cloud computing networks.
[0059] User interface 306 includes one or more output devices 308 and / or one or more input devices 310. Input devices 310 may include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. Output devices 308 may include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or monitors), etc.
[0060] Memory 314 may include high-speed random access memory (e.g., DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), DDR RAM (Double Data Rate Random Access Memory), and / or other random access solid-state memory devices) and / or non-volatile memory (e.g., one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state memory devices). Memory 314 may optionally include one or more storage devices located remotely from the control circuitry system 302. Memory 314, or alternatively, the non-volatile solid-state memory devices within memory 314 include non-transitory computer-readable storage media. In some embodiments, memory 314 or the non-transitory computer-readable storage media of memory 314 stores programs, modules, instructions, and data structures, or subsets or supersets thereof:
[0061] • Operating system 316, which includes procedures for handling various basic system services and for performing hardware-related tasks;
[0062] • Network communication module 318, which is used to connect server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections);
[0063] • Encoding / decoding module 320, which performs various functions related to encoding and / or decoding data, such as video data. In some embodiments, encoding / decoding module 320 is an example of codec component 114. Encoding / decoding module 320 includes, but is not limited to, one or more of the following:
[0064] ◦ Decoding module 322, which performs various functions related to decoding encoded data, such as those previously described with respect to decoder component 122; and
[0065] ◦ Encoding module 340, which performs various functions related to encoding data, such as those previously described with respect to encoder component 106; and
[0066] • For example, an image memory 352 for storing images and image data, used in conjunction with the encoding / decoding module 320. In some embodiments, the image memory 352 includes one or more of the following: a reference image memory 208, a buffer memory 252, a current image memory 264, and a reference image memory 266.
[0067] In some implementations, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to the parser 254), a transformation module 326 (e.g., configured to perform various functions previously described with respect to the scaler / inverse transformation unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra-frame image prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to the loop filter 256).
[0068] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions previously described with respect to source encoder 202 and / or encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described with respect to predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include... Figure 3 A subset of the modules shown. For example, a prediction module shared by both decoding module 322 and encoding module 340.
[0069] Each of the modules identified above and stored in memory 314 corresponds to a set of instructions for performing the functions described herein. The modules identified above (e.g., instruction sets) do not need to be implemented as separate software programs, processes, or modules, and therefore various subsets of these modules can be combined or otherwise rearranged in various implementations. For example, the codec module 320 may optionally not include separate decoding and encoding modules, but instead use the same set of modules to perform both sets of functions. In some implementations, memory 314 stores a subset of the modules and data structures identified above. In some implementations, memory 314 stores additional modules and data structures not described above.
[0070] Although Figure 3 A server system 112 according to some embodiments is shown, but Figure 3 This is intended more as a functional description of various features that can exist in one or more server systems than as a structural diagram of the implementation described herein. In practice, items shown individually may be combined, and some items may be separated. For example, Figure 3 Some items shown individually can be implemented on a single server, and a single item can be implemented by one or more servers. The actual number of servers used to implement server system 112 and how features are distributed among said servers will vary depending on the implementation method, and optionally, it depends in part on the amount of data traffic processed by the server system during peak usage periods and during average usage periods.
[0071] Example encoding and decoding techniques
[0072] The encoding and decoding processes and techniques described below can be performed at the devices and systems described above (e.g., source device 102, server system 112, and / or electronic device 120). According to some implementations, an example method for conditionally signaling an indicator associated with intra-block copying information is described below.
[0073] Intra-block copying methods use block vectors to identify predicted blocks for the current block. Block vectors are used to identify another block in the same image that may be adjacent to or not adjacent to the current block. Block vectors can be explicitly signaled or implicitly derived. When block vectors are explicitly signaled, it is usually called an intra-block copying method. When block vectors are implicitly derived, for example by comparing the template region (a set of neighboring reconstructed samples located adjacent to the current block) between the current block and the candidate predicted block, it is usually called an intra-template matching method.
[0074] In some implementations, a "local search region" refers to a neighboring reconstructed sample region of the current block that can be used to identify the predicted block. For example, a region covered by a block region of a given size (e.g., 64×64) associated with the current block can be a local search region. As used herein, a global search region can refer to a reconstructed sample region farther away from the current block that can be used to identify the predicted block. For example, a region beyond a neighboring block region of a given size (e.g., all reconstructed sample regions beyond two previously encoded superblocks (or so-called coding tree blocks)).
[0075] "Superblock size" can refer to the maximum block size applied to encode an image / video picture or video sequence. Block size or region size can refer to the width and / or height of the block / region, the area of the block / region, the number of samples in the block / region, the maximum (or minimum) value between the width and height of the block / region, and / or the aspect ratio of the block / region.
[0076] Figure 4A An example of intra-block copying technology is illustrated. In this example, intra-block copying technology includes identifying a predicted block 404 within the same frame 400 as the current block 402 via a predicted block vector 408. Figure 4 shows an example where the predicted block 404 is not adjacent to the current block 402. In some implementations, the predicted block 404 is adjacent to the current block 402. In some implementations, the predicted block vector 408 is selected from a list of candidate block vectors. The list of candidate block vectors may be populated with block vectors used in neighboring blocks and / or block vectors from a block vector library. The block vector prediction value index can be signaled to indicate which candidate block vector in the candidate list was used to predict the block vector of the current block 402.
[0077] In some implementations, the Block Vector Difference (BVD) 410 represents the difference between the predicted block vector and the actual block vector 406, which is the vector between corresponding portions of the current block (e.g., current block 402) and the predicted block (e.g., predicted block 404). Although Figure 4A The BVD 410 depicted has components along both the horizontal and vertical dimensions, but the BVD can extend along a single dimension or span two or more dimensions. In some implementations, the BVD 410 may include information representing the magnitude and / or direction of the block vector difference, one or both of which may be signaled in the bitstream as intra-frame block copy information or syntax.
[0078] Figure 4B Examples of different intra-block copying modes according to some implementation methods are shown. Figure 4B An example image 420 with multiple samples is shown. As an example, the dark gray sample in the global search region 422 and the cross-shaded sample in the local search region 424 both correspond to the reconstructed samples in image 420. In some implementations, the current block 426 in the superblock 428 is searched from the reconstructed samples (e.g., similar to...). Figure 4A The predicted block (e.g., similar to the current block 402) in the current block (402) Figure 4A Predicted block 404). Superblock 428 in Figure 4B The diagram is shown schematically, and the number of dashed squares does not correspond to the actual number of samples. Superblock 428 can be, for example, 128×128, 64×64, or other sizes. Superblock 428 can be divided into smaller coded blocks, such as current block 426 and coded block 430.
[0079] In some implementations, searching for a prediction block from a coded block that covers all available reconstructed samples requires more processing time than searching only a local search region. In some implementations, the reconstructed sample closest to the current block 426 includes information content more similar to the current block 426 (e.g., in a scene where image 420 includes more localized information, such as localized texture), and searching for a prediction block within the local search region 424 can be more time-efficient and provide a more suitable prediction block. For example, the features of the current block 426 can reflect the information content of the current block. If the current block contains more textural information, the current block may be smaller (e.g., 8×8, 16×16, 32×32, etc.). If the information content within the current block is smoother (e.g., without much texture), the current block may be larger (e.g., 64×64, etc.).
[0080] In some implementations, the first intra-block copy mode includes searching for the prediction block in a more restricted, smaller local region 424, rather than searching in a larger global region (e.g., the sum of regions 422 and 424, or only region 422). In some implementations, the second intra-block copy mode includes searching for the prediction block in a larger global region that includes at least region 422. In some implementations, when the first intra-block copy mode is unavailable (e.g., no prediction block is searched in the smaller local region 424, but the current block is still encoded using the intra-block copy method), the second intra-block copy mode (e.g., searching using a larger global region) is the default intra-block copy mode.
[0081] In some implementations, the encoder may signal one or more flags and / or syntax indicating the use of an intra-block copying method to the bitstream. For example, an intra-block copying enable / disable flag may be signaled to indicate whether the intra-block copying method is used on a coded block. Intra-block copying related parameters may also be signaled, such as block vectors (e.g., 406) and block vector prediction indexes (e.g., indicating which block vector in the candidate block vector list to use). In some implementations, whether to signal an intra-block copying enable flag (e.g., to indicate the use of the intra-block copying method) may depend on at least one of the following: coded block size, superblock size, selection of a search region, and / or the relative position of the current block with respect to an associated block of a given size.
[0082] In some implementations, the block size threshold s2 is derived based on a given superblock size (e.g., denoted as s0) and the size of the local search region (e.g., denoted as s1). In some implementations, the size s1 of the local search region can vary depending on the location of the current block. For example, for a current block at the edge of image 420, the local search region s1 can be smaller compared to a current block inside image 420. In some implementations, whether to use the first intra-block copy mode depends on the size of the current block. In some implementations, when the current block is less than or equal to the block size threshold s2, a flag and syntax indicating the use of the intra-block copy method are signaled to indicate that the current block should be encoded using the intra-block copy method. In some implementations, the syntax further indicates whether to search for the predicted block within the local search region (e.g., region 424) or within both the local search region (region 424) and a larger global search region (e.g., region 422).
[0083] In some implementations, no signaling flags and syntax are used when the current block is larger than a block size threshold s2. For example, in this case, there may not be a sufficient number of reconstructed samples and / or potential predicted blocks, making it less likely to find a good predicted block. Therefore, if intra-block copying is used to encode the current block in this situation, the local search region 424 is disabled, and the search for predicted blocks is performed only in the global search region. In some implementations, intra-block copying is not used to encode the current block; instead, other prediction modes (e.g., angular prediction modes and / or non-angular prediction modes) are used instead.
[0084] In some implementations, when the intra-block copy flag and / or intra-block copy syntax are not signaled, the second intra-block copy mode is used by default (e.g., searching for prediction blocks within the global search region 422). In some implementations, when the intra-block copy flag and / or intra-block copy syntax are not signaled, the intra-block copy method is not used for the current block (e.g., the current block uses either angular intra-prediction mode or non-angular intra-prediction mode).
[0085] In some implementations, the block size threshold s2 is equal to the local search region s1, or the block size threshold s2 is s1 / N, where N is a power of 2, such as 2, 4, 8, 16, etc. As an example, s1 can be 64×64, and the block size threshold s2 can be derived as 64×64 / 4 (e.g., N is 4), which corresponds to 32×32. When the current block size is less than or equal to 32×32, a signal is used to indicate the flags and syntax for using the intra-block copying method (e.g., indicating the use of a first intra-block copying mode or a second intra-block copying mode). In some implementations, when the intra-block copying flag and / or intra-block copying syntax are not signaled, the second intra-block copying mode is used by default (e.g., searching for prediction blocks within the global search region 422). In some implementations, when the intra-block copying flag and / or intra-block copying syntax are not signaled, the intra-block copying method is not used for the current block (e.g., the current block uses an angular intra-prediction mode or a non-angular intra-prediction mode).
[0086] In some implementations, the block size threshold s2 is equal to the smaller of the superblock size s0 and the local search region s1 (e.g., also denoted as...). In some implementations, the block size threshold s2 is equal to... Where N is a power of 2, such as 2, 4, 8, 16, etc. As an example, the superblock size s0 can be 32×32, the local search region s1 can be 64×64, and the block size threshold s2 can be derived as follows: In this example, it is 16×16. In this scenario, when the coded block size (e.g., the size of region 426) is less than or equal to 16×16, the encoder signals the intra-block copy enable flag and intra-block copy syntax (e.g., block vector, block vector prediction index, and / or block vector difference). In some implementations, when the intra-block copy flag and / or intra-block copy syntax are not signaled, the second intra-block copy mode is used by default (e.g., searching for prediction blocks within global search region 422). In some implementations, when the flag and / or syntax are not signaled, the intra-block copy method is not used (e.g., the current block uses angular intra-prediction mode or non-angular intra-prediction mode).
[0087] In some implementations, the block size threshold s2 is predefined and set to a fixed value for all video sequences. In some implementations, the block size threshold s2 is signaled to the bitstream using a high-level syntax, such as a sequence header, frame header, slice header, etc.
[0088] In some implementations, when the current block is located at a selected relative position within a block of a predefined block size, the encoder does not signal intra-block copying related syntax (e.g., the encoder forgoes signaling intra-block copying related syntax). For example, if the current block is located at the upper left of a 64×64 block in a 64×64-block grid (e.g., a segmented grid) of the current image, the encoder does not signal intra-block copying related syntax (e.g., the encoder forgoes signaling intra-block copying related syntax). In some implementations, such as Figure 4B As shown, the current block 426 is in the upper left portion of the superblock 428, and the search range for the predicted block for the current block 426 will appear in the lower left portion 432, thus reducing the likelihood of a good predicted block being selected. For example, the lower left portion 432 may be adjacent to the unprocessed region 434 of the current image 420. Therefore, the encoder abandons signaling the syntax related to intra-block copying. In contrast, in some implementations, for the coded block 430 in the lower right corner of the superblock 428, the search range for the predicted block for coded block 430 is not limited to the lower left portion 432, and the encoder can signal the syntax related to intra-block copying for coded block 430 to the bitstream.
[0089] In some implementations, when the current block size is larger than the reconstructed region size of the associated 64×64 block in the 64×64 segmentation grid of the current image (or larger than a scaling value of that reconstructed region size), the encoder does not signal the intra-block copying related syntax (e.g., the encoder waives signaling the intra-block copying related syntax). For example, instead of Figure 4BIn the global search region 422 shown in the image, four complete rows of gray reconstructed samples are used. In a scenario where only one or two rows of reconstructed samples are available in image 420 and the current block size is larger than at least a portion of the area of the reconstructed samples, the encoder does not signal the syntax related to intra-block copying and intra-block copying is not used to encode the current block.
[0090] In some implementations, when both the global search region (e.g., region 422) and the local search region (e.g., region 424) and the global search region (e.g., region 422) are applicable to the search of the prediction block, the conditions for signaling the syntax related to the intra-block copying method may be different from the conditions for signaling the syntax related to the intra-block copying method when only the local search region is applicable.
[0091] In some implementations, the parameters related to intra-block copying (e.g., block vector or block vector difference) may be signaled to depend on one or more features of the current block as described above.
[0092] In some implementations, the maximum value of one or more intra-block copying related parameters (e.g., block vectors or block vector differences) can be derived, and the intra-block copying related parameters can be signaled to depend on this maximum value. For example, if the search for the predicted block is performed in a smaller local search region 424, there may be an upper limit to the block vector (e.g., the block vector is not too large). In some implementations, if searching for the predicted block in the smaller local search region 424 is disabled (e.g., by the encoder not signaling the IBC flag and / or syntax), searching for the predicted block in the global search region 422 may still be allowed if the current block is to be encoded using intra-block copying. Alternatively, the current block can be encoded using a non-IBC method (e.g., angular prediction mode or non-angular prediction mode).
[0093] In some implementations, block-level syntax elements such as block splitting flags, transform block splitting flags, reference frame indexes, motion vectors / block vectors, prediction modes, EOB (End of Block), transform skip flags, transform type flags, etc., are not signaled in the bitstream, but are instead implicitly inferred by the decoding unit when such frames are identified by separate frame-level flags. In some implementations, previous IBC frames are used as references unless otherwise explicitly signaled, in which case all blocks are skipped and zero motion vectors are used.
[0094] In some implementations, the local search region extends to the local memory of 128×128 blocks within the luminance samples. For example, for sub-blocks of 64×64 or smaller, reconstructed samples in the left superblock (if available) and already encoded samples in the current superblock can be used as IBC references. The availability of reconstructed samples on the left superblock is based on the position of the current 64×64 block. As an example, for the current block within the top-left 64×64 sub-block of the current superblock, only the already reconstructed samples within the current sub-block, the reconstructed samples of the top-right 64×64 sub-block, the bottom-left 64×64 sub-block of the left superblock, and / or the bottom-right 64×64 sub-block can be used. In another example, for the current block within the top-right 64×64 sub-block of the current superblock, only the reconstructed samples within the current sub-block, the reconstructed samples of the top-left 64×64 sub-block of the current superblock, the bottom-left 64×64 sub-block of the left superblock, and / or the bottom-right 64×64 sub-block can be used. In another example, for the current block within the bottom-left 64×64 sub-block of the current superblock, only the reconstructed samples from the top-left and top-right sub-blocks of the current superblock, and / or the reconstructed samples from the bottom-right 64×64 sub-block of the left superblock, can be used. In yet another example, for the current block within the bottom-right 64×64 sub-block of the current superblock, only the reconstructed samples from the sub-blocks of the current superblock can be used.
[0095] In some implementations, the IBC derivation is performed as shown in Table 1 below.
[0096]
[0097] Table 1 - Derivation of Intra-Block Copying
[0098] Figure 5A This is a flowchart illustrating a method 500 for decoding video according to some embodiments. Method 500 may be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 500 is executed by executing instructions stored in the computing system's memory (e.g., memory 314).
[0099] The system receives (502) a video bitstream (e.g., an encoded video sequence) comprising multiple blocks (e.g., corresponding to a set of pictures), said multiple blocks including the current block. When one or more features of the current block satisfy (504) one or more criteria: the system parses (506) one or more indicators from the video bitstream, which indicate IBC information. The system identifies (508) a predicted block of the current block based on the IBC information, and the system reconstructs (510) the current block using the predicted block. When one or more features of the current block do not satisfy (514) one or more criteria: the system decodes (514) the current block without parsing one or more indicators indicating IBC information.
[0100] In some implementations, the use of intra-block copying methods is conditionally indicated by signals, such as intra-block copying enable / disable flags, intra-block copying related parameters such as block vectors, block vector prediction indexes, and the conditions may include parameters related to: coded block size, superblock size, selection of search region, and / or the relative position of the current block with respect to associated blocks of a given size.
[0101] In some implementations, given a superblock size (denoted as s0) and a local search region size (denoted as s1), the block size s2 is derived. The flags and syntax indicating the use of the intra-block copying method are signaled only if the coded block is less than or equal to s2. Otherwise, the flags and syntax are not signaled.
[0102] In some implementations, s2 is equal to s1, or s2 is s1 / N, where N is a power of 2, such as 2, 4, 8, 16, etc. For example, s1 can be 64×64, and s2 can be derived as 64×64 / 4, which is 32×32. The flags and syntax indicating the use of the intra-block copying method are only signaled when the coded block size is less than or equal to 32×32.
[0103] In some implementations, s2 equals , or s2 is Where N is a power of 2, such as 2, 4, 8, 16, etc. For example, s0 can be 32×32, s1 can be 64×64, and s2 is derived as... That is, 16×16. The flags and syntax indicating the use of the intra-block copying method are only signaled when the coded block size is less than or equal to 16×16.
[0104] In some implementations, when a coded block is located at a selected relative position within a given block size having a predefined block size, no signal is needed to notify intra-block copying related syntax. For example, when the current coded block is located to the upper left of an associated 64×64 block within a 64×64 segmentation grid of the current image, no signal is needed to notify intra-block copying related syntax.
[0105] In another example, when the current coded block size is larger than the reconstructed region size of the associated 64×64 block in the 64×64 segmentation grid of the current image (or larger than the scaling value of that reconstructed region size), no signal is used to notify the syntax related to intra-block copying.
[0106] In some implementations, the conditions for signaling the syntax related to the intra-block copying method may differ from the conditions for signaling the syntax related to the intra-block copying method when a global search region or both a local search region and a global search region are applicable.
[0107] In some implementations, when a global search region or both a local search region and a global search region are applicable, the conditional signaling of the intra-block copy flag and syntax described above is not applied.
[0108] In some implementations, the size of s2 is set to a predefined and fixed value for all video sequences. In other implementations, the size of s2 is signaled to the bitstream using a high-level syntax, such as a sequence header, frame header, slice header, etc.
[0109] In some implementations, intra-block copying is not applied to the current block when no signal is used to indicate the flags and syntax for using the intra-block copying method.
[0110] In some implementations, the parameters related to intra-block copying (e.g., block vectors or block vector differences) may be signaled to depend on the conditions described above. As an example, a maximum value for the parameters related to intra-block copying (e.g., block vectors or block vector differences) may be derived, and the parameters related to intra-block copying may be signaled to depend on this maximum value.
[0111] Figure 5B This is a flowchart illustrating a method 550 for encoding video according to some embodiments. Method 550 can be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 550 is executed by executing instructions stored in the computing system's memory (e.g., memory 314). In some embodiments, method 550 is executed by the same system as method 500 described above.
[0112] The system receives (552) video data (e.g., a source video sequence) comprising multiple blocks (e.g., corresponding to a set of images), said multiple blocks including the current block. When one or more features of the current block satisfy (554) one or more criteria: the system encodes the current block using IBC mode (556), and the system signals (558) one or more indicators in the video bitstream, which indicate the IBC information of the current block. When one or more features of the current block do not satisfy (560) one or more criteria: the system does not signal (562) the IBC information of the current block. As previously described, the encoding process can mirror the decoding process described herein (e.g., the conditional signaling of the intra-mode copy implementation described above). For brevity, these details will not be repeated here.
[0113] Although Figure 5A and Figure 5B Multiple logical stages are shown in a specific order, but stages that are not in order can be reordered, and other stages can be combined or split. Some reorderings or other groupings not specifically mentioned will be obvious to those skilled in the art, and therefore the orderings and groupings presented herein are not exhaustive. Furthermore, it should be understood that the stages can be implemented in hardware, firmware, software, or any combination thereof.
[0114] Now let's turn to some example implementations.
[0115] (A1) In one aspect, some implementations include a method for video decoding (e.g., method 500). In some implementations, the method is performed at a computing system (e.g., server system 112) having memory and control circuitry. In some implementations, the method is performed at an encoding / decoding module (e.g., encoding / decoding module 320). In some implementations, the method is performed at a source encoding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes (i) receiving a video bitstream (e.g., an encoded video sequence) comprising a plurality of blocks, said plurality of blocks including a current block. When one or more features of the current block satisfy one or more criteria: the method includes: (ii) parsing one or more indicators from the video bitstream, said one or more indicators indicating IBC information; (iii) identifying a predicted block of the current block based on the IBC information; and (iv) reconstructing the current block using the predicted block. When one or more features of the current block do not meet one or more criteria, the method includes: (v) decoding the current block without parsing one or more indicators indicating IBC information. For example, flags and / or syntax elements indicating the use of an intra-block copying method (e.g., intra-block copying enable / disable flags, and / or parameters related to intra-block copying such as block vectors, block vector prediction indexes, etc.) are conditionally signaled based on at least one of the following: coded block size, superblock size, selection of a search region (e.g., selecting one of global IBC and local IBC), and / or the position of the current block relative to an associated block of a given size. In some embodiments, one or more indicators indicating IBC information are parsed from the video bitstream based on the determination that one or more features of the current block meet one or more criteria. In some embodiments, the current block is decoded without parsing one or more indicators indicating IBC information (e.g., the system abandons parsing one or more indicators) based on the determination that one or more features of the current block do not meet one or more criteria.
[0116] (A2) In some implementations of A1, one or more criteria include at least one of the following: the size of the coded block of the current block; the size of the superblock corresponding to the current block; the selection of the search region for the current block; and the position of the current block relative to an associated block of a given size.
[0117] (A3) In some embodiments of A1 or A2, the IBC information includes at least one of the following: an IBC mode flag; a block vector; a block vector prediction index; and a block vector difference. For example, signaling IBC-related parameters (e.g., block vector, block vector difference, and / or block vector accuracy) is based on whether one or more features of the current block meet one or more criteria. In some embodiments, the IBC mode flag indicates whether IBC mode is enabled for the current block. In some embodiments, the IBC mode flag indicates whether global IBC mode or local IBC mode is enabled for the current block.
[0118] (A4) In some embodiments of A1 through A3, decoding the current block without parsing one or more indicators indicating IBC information includes decoding the current block according to a non-IBC mode. For example, a non-IBC mode could be an intra-prediction mode (e.g., angular intra-prediction mode or non-directional intra-prediction mode). As an example, IBC is not applied to the current block when the flags and syntax indicating the use of IBC are not signaled / parsed. In some embodiments, decoding the current block without parsing one or more indicators indicating IBC information includes deriving the IBC information based on encoded / decoded information (e.g., available at the decoder).
[0119] (A5) In some implementations of any of A1 to A4, one or more criteria include a size threshold, which is based on the superblock size and the local search region size; and the current block satisfies one or more criteria when the current block size is smaller than the size threshold. For example, the size threshold s2 is derived based on the superblock size s0 and the local search region size s1. For coded blocks smaller than or equal to the size threshold s2, a flag and / or syntax indicating the use of IBC is signaled; otherwise, the flag and / or syntax is not signaled. If the block size is large relative to the superblock size and / or search region, there may not be enough pixels to search (e.g., in the local search region) to obtain good candidates. In this case, local IBC should not be used because the chance of finding good candidates is low. Therefore, to save signaling overhead, the system may forgo signaling IBC information. In some implementations, local IBC is disabled in this case (e.g., the system uses global IBC by default).
[0120] (A6) In some implementations of A5, the size threshold is equal to the local search region size divided by N, where N is a power of 2. For example, N can be 1, 2, 4, 8, or 16. For example, s2 is equal to s1 or s1 / N, where N is a power of 2. As a specific example, when s1 is 64×64, s2 can be derived as 64×64 / 4, which is 32×32. In this case, the flags and syntax indicating the use of IBC are signaled only when the coded block size is less than or equal to 32×32. Therefore, when the block size is greater than s1 / N, local IBC is not allowed (disabled).
[0121] (A7) In some implementations of A5, the size threshold is equal to S / N, where S is equal to the smaller of the superblock size and the local search region size, and N is a power of 2. For example, s² equals , or s2 is Where N is a power of 2 (e.g., 2, 4, 8, 16). As a specific example, s0 is 32×32, s1 is 64×64, and s2 is derived as... That is, 16×16. In this case, the flags and syntax for using IBC are signaled only when the code block size is less than or equal to 16×16.
[0122] (A8) In some implementations of A5, the size threshold is a predefined value. For example, the size threshold s2 is set to a predefined and fixed value for all video sequences.
[0123] (A9) In some embodiments of A5, the method further includes: parsing a second indicator indicating a size threshold from the video bitstream. For example, the size threshold s2 is signaled to the video bitstream in a high-level syntax, such as a sequence header, frame header, or slice header.
[0124] (A10) In some implementations of any of A1 to A9, one or more criteria include a block size criterion and a block relative position criterion. For example, when the encoded block is located at a selected relative position within a given block having a predefined block size, no signal is used to notify the IBC-related syntax.
[0125] (A11) In some implementations of A10, the block relative position criterion includes a criterion regarding whether the current block is located in the upper-left position within the segmentation grid. For example, when the current coded block is located in the upper-left position of an associated 64×64 block within the 64×64 segmentation grid of the current image, no signal is needed to notify the intra-block copying related syntax. In this case, the search region may have a lower chance of generating good candidates. For a coded block in the upper-left position, the search region may be in the lower-left position, which may also result in a lower chance of generating good candidates.
[0126] (A12) In some implementations of A11, the block size criterion includes a criterion regarding whether the size of the current block is greater than the search region size. For example, if the current encoded block size is greater than the reconstructed region size (or a scaling value of the region size) of the associated 64×64 block in the 64×64 segmentation grid of the current image, then no signal is given to the syntax related to intra-block copying (e.g., if the search region is not large enough for the block, then local IBC is not allowed).
[0127] (A13) In some embodiments of any of A1 to A12, the method further includes: determining a first set of criteria as one or more criteria when the current block allows global IBC mode; and determining a second set of criteria as one or more criteria when global IBC mode is not allowed for the current block. For example, the conditions for signaling IBC-related syntax may differ from the conditions for signaling IBC-related syntax when only a local search region applies, or when both a global search region and a local search region apply.
[0128] (A14) In some embodiments of A1 to A13, the method further includes: decoding the current block without considering whether one or more features of the current block satisfy one or more criteria when the current block allows global IBC mode. For example, conditional parsing of IBC information is not applied when a global search region or both a local search region and a global search region are applicable. In some embodiments, when the current block allows global IBC mode, one or more indicators are parsed from the video bitstream without considering whether one or more features of the current block satisfy one or more criteria.
[0129] (A15) In some embodiments of any of A1 to A14, the method further includes: deriving a maximum value of the IBC parameter, wherein one or more criteria include criteria regarding the maximum value of the IBC parameter. For example, a maximum value of an IBC-related parameter (e.g., a block vector or block vector difference) can be derived, and the IBC-related parameter can be signaled to depend on that maximum value. For example, if the maximum value of the block vector or block vector difference is too large, the local IBC mode can be disabled.
[0130] (B1) In another aspect, some implementations include a video encoding method (e.g., method 550). In some implementations, the method is performed at a computing system (e.g., server system 112) having memory and control circuitry. In some implementations, the method is performed at an encoding / decoding module (e.g., encoding / decoding module 320). The method includes: receiving video data (e.g., a source video sequence) comprising a plurality of blocks, said plurality of blocks including a current block. When one or more features of the current block satisfy one or more criteria: the method includes: encoding the current block using an IBC mode; and signaling one or more indicators in the video bitstream, said one or more indicators indicating the IBC information of the current block. When one or more features of the current block do not satisfy one or more criteria, the method includes: not signaling the IBC information of the current block. In some implementations, when one or more features of the current block do not satisfy one or more criteria, the current block is encoded in a non-IBC mode. In some implementations, when one or more features of the current block do not satisfy one or more criteria, the current block is encoded in a global IBC mode.
[0131] (B2) In some implementations of B1, the method includes: signaling the encoded current block in the video bitstream.
[0132] (B3) In some implementations of B1 or B2, when the current block allows global IBC mode, the method includes determining a first set of criteria as one or more criteria. When global IBC mode is not allowed for the current block, the method includes determining a second set of criteria as one or more criteria.
[0133] (B4) In some implementations of any of B1 to B3, when the current block allows the global IBC mode, the method includes: encoding the current block without considering whether one or more features of the current block meet one or more criteria.
[0134] (C1) In another aspect, some embodiments include a method for processing visual media data. In some embodiments, the method is performed at a computing system (e.g., server system 112) having memory and control circuitry. In some embodiments, the method is performed at an encoding / decoding module (e.g., encoding / decoding module 320). The method includes: obtaining a source video sequence comprising multiple frames; and performing a conversion between the source video sequence and a video bitstream of visual media data according to format rules. The video bitstream includes a set of encoded blocks containing a current block, and when one or more features of the current block satisfy one or more criteria, one or more indicators are signaled indicating IBC information of the current block. The format rules specify that when one or more features of the current block satisfy one or more criteria, IBC information is parsed from the video bitstream, a predicted block of the current block is identified based on the IBC information, and the current block is reconstructed using the predicted block. In some embodiments, the format rules also specify that when one or more features of the current block do not satisfy one or more criteria, the current block is decoded without parsing one or more indicators indicating IBC information. In some implementations, the video bitstream does not include one or more indicators when one or more features of the current block do not meet one or more criteria.
[0135] On the other hand, some implementations include a computing system (e.g., server system 112) that includes a control circuitry system (e.g., control circuitry system 302) and a memory (e.g., memory 314) coupled to the control circuitry system, the memory storing one or more sets of instructions configured to be executed by the control circuitry system, the set of one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A15, B1 to B4 and C1 above).
[0136] In another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by a control circuitry of a computing system, said set of instructions including instructions for performing any of the methods described herein (e.g., A1 to A15, B1 to B4, and C1 above). In some embodiments, the non-transitory computer-readable storage medium stores a video bitstream generated by any of the video coding methods described herein.
[0137] Unless otherwise stated, any syntax element described herein (e.g., indicator) can be a High-Level Syntax (HLS). As used herein, the HLS is signaled at a level higher than the block level. For example, the HLS can correspond to a sequence level, frame level, slice level, or tile level. As another example, HLS elements can be signaled in a Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptation Parameter Set (APS), slice header, picture header, tile header, and / or CTU (Coding Tree Unit) header.
[0138] It should be understood that while the terms “first,” “second,” etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of embodiments and appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or,” as used herein, refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that, when used in this specification, the terms “comprising” and / or “including” specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0139] As used herein, the term "if" may be interpreted, depending on the context, as meaning "when the prerequisite is true," "after the prerequisite is true," "in response to determining that the prerequisite is true," "based on the determination that the prerequisite is true," or "in response to detecting that the prerequisite is true." Similarly, the phrases "if it is determined [the prerequisite is true]," "if [the prerequisite is true]," or "when [the prerequisite is true]" may be interpreted, depending on the context, as meaning "after determining that the prerequisite is true," "in response to determining that the prerequisite is true," "based on the determination that the prerequisite is true," "after detecting that the prerequisite is true," or "in response to detecting that the prerequisite is true."
[0140] For illustrative purposes, the foregoing description has been described with reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the claims to the precise forms disclosed. In view of the above teachings, many modifications and variations are possible. The embodiments were chosen and described in order to best illustrate the principles of operation and practical application, thereby enabling others skilled in the art to implement them.
Claims
1. A method for video decoding performed at a computing system having a memory and one or more processors, the method comprising: Receive a video bitstream comprising multiple blocks, including the current block; When one or more features of the current block satisfy one or more criteria: Parse one or more indicators from the video bitstream, the one or more indicators indicating intra-block copy (IBC) information; The predicted block for the current block is identified based on the IBC information; as well as Reconstruct the current block using the predicted block; as well as If one or more features of the current block do not meet the one or more criteria, the current block is decoded without parsing the one or more indicators that indicate the IBC information.
2. The method according to claim 1, wherein, The one or more criteria include at least one of the following: The size of the encoded block of the current block; The size of the superblock corresponding to the current block; Selection of the search region for the current block; and The position of the current block relative to an associated block of a given size.
3. The method according to claim 1, wherein, The IBC information includes at least one of the following: IBC pattern logo; Block vector; Block vector predicted value index; and Block vector difference.
4. The method according to claim 1, wherein, Decoding the current block without parsing the one or more indicators that indicate the IBC information includes: decoding the current block according to a non-IBC mode.
5. The method according to claim 1, wherein: The one or more criteria include a size threshold, which depends on the superblock size and the local search region size; and When the size of the current block is less than the size threshold, the current block satisfies one or more of the criteria.
6. The method according to claim 5, wherein, The size threshold is equal to the size of the local search region divided by N, where N is a power of 2.
7. The method according to claim 5, wherein, The size threshold is equal to S / N, where S is equal to the smaller of the superblock size and the local search region size, and where N is a power of 2.
8. The method according to claim 5, wherein, The size threshold is a predefined value.
9. The method according to claim 5, further comprising: A second indicator indicating the size threshold is parsed from the video bitstream.
10. The method according to claim 1, wherein, The one or more criteria include block size criteria and block relative position criteria.
11. The method according to claim 10, wherein, The relative position criteria of the blocks include criteria regarding whether the current block is located at the top left position in the segmented grid.
12. The method according to claim 10, wherein, The block size criterion includes a standard regarding whether the size of the current block is greater than the size of the search region.
13. The method according to claim 1, further comprising: When the current block allows global IBC mode, the first set of criteria is determined as one or more of the criteria; as well as When the current block does not allow global IBC mode, the second set of criteria is determined as one or more of the criteria.
14. The method according to claim 1, further comprising: When the current block allows global IBC mode, the current block is decoded without considering whether one or more features of the current block meet one or more criteria.
15. The method according to claim 1, further comprising: Derive the maximum value of the IBC parameter, wherein the one or more criteria include criteria regarding the maximum value of the IBC parameter.
16. A method of video encoding performed at a computing system having a memory and one or more processors, the method comprising: Receive video data comprising multiple blocks, including the current block; When one or more features of the current block satisfy one or more criteria: The current block is encoded using the Intra-Block Copy (IBC) mode; as well as In the video bitstream, one or more indicators are signaled, which indicate the IBC information of the current block; as well as When one or more features of the current block do not meet one or more criteria, the IBC information of the current block is not notified without signaling.
17. The method of claim 16, further comprising: The current block that has been encoded is notified by a signal in the video bitstream.
18. The method of claim 16, further comprising: When the current block allows global IBC mode, the first set of criteria is determined as one or more of the criteria; as well as When global IBC mode is not allowed for the current block, the second set of criteria is determined as one or more of the criteria.
19. The method of claim 16, further comprising: When the current block allows global IBC mode, the current block is encoded without considering whether one or more features of the current block meet one or more criteria.
20. A non-transitory computer-readable storage medium storing a video bitstream generated by a video coding method, the method comprising: Receive video data comprising multiple blocks, including the current block; When one or more features of the current block satisfy one or more criteria: The current block is encoded using the Intra-Block Copy (IBC) mode; as well as In the video bitstream, one or more indicators are signaled, and the one or more indicators indicate the IBC information of the current block; as well as When one or more features of the current block do not meet one or more criteria, the IBC information of the current block is not notified by signal.