Video encoding and decoding method, computer device, equipment and computer readable medium

By disabling multiple encoding tools in video encoding and adopting intra-block copy and string copy modes, optimizing encoding for screen content is solved, and the problem of low video encoding efficiency in screen content in the prior art is solved, and more efficient video compression is achieved.

CN115486072BActive Publication Date: 2025-05-06TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180030883.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-16
Filing Date
2021-09-24
Publication Date
2025-05-06
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

Existing video encoding technology has low encoding efficiency when processing screen content, especially in videos containing a large amount of text and graphics. Traditional encoding tools cannot effectively utilize the characteristics of these contents, resulting in low encoding efficiency.

Method used

By introducing advanced syntax control in video encoding, disabling multiple encoding tools, and using specific encoding tools such as intra-block copy and string copy modes to optimize encoding for screen content.

Benefits of technology

Improve the encoding efficiency of screen content videos, reduce the size of the bitstream, and improve the video compression effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115486072B_ABST
    Figure CN115486072B_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods for video decoding and devices including processing circuits. The processing circuit can decode encoding information of multiple blocks from an encoded video bitstream. The encoding information can indicate advanced control flags for the multiple blocks. The advanced control flags can indicate whether multiple encoding tools are disabled for at least one of the multiple blocks, wherein at least one of the multiple blocks includes a current block. The processing circuit can determine whether multiple encoding tools are disabled for at least one of the multiple blocks based on the advanced control flags. The processing circuit can reconstruct the current block without the multiple encoding tools based on the multiple encoding tools being determined to be disabled.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation by reference

[0002] This application claims the benefit of priority to U.S. Patent Application No. 17 / 477,151, filed on September 16, 2021, "METHOD AND APPARATUS FOR VIDEO CODING," which claims the benefit of priority to U.S. Provisional Application No. 63 / 152,178, filed on February 22, 2021, "High level syntax control for Screen Content Coding." The entire disclosure of the prior application is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure describes embodiments generally related to video decoding, and more particularly to a method, computer device, apparatus, and computer-readable medium for video encoding and decoding. Background Art

[0004] The purpose of the background description provided herein is to generally present the context of the present disclosure. To the extent that the work of the presently named inventors is described in this background section, the work of the presently named inventors and aspects of the description that may not otherwise be considered prior art at the time of filing are neither explicitly nor implicitly admitted as prior art to the present disclosure.

[0005] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. An uncompressed digital video may include a series of pictures, each of which has a spatial size of, for example, 1920×1080 luma samples and associated chroma samples. The series of pictures may have a fixed or variable picture rate (also informally referred to as a frame rate), such as 60 pictures per second or 60 Hz. Uncompressed video has specific bit rate requirements. For example, 1080p60 4:2:0 video at 8 bits per sample (1920×1080 luma sample resolution at 60 Hz frame rate) requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires more than 600 GBytes of storage space.

[0006] One purpose of video encoding and decoding can be to reduce the redundancy of the input video signal by compression. Compression can help reduce the above-mentioned bandwidth requirements and / or storage space requirements, in some cases by two orders of magnitude or more. Both lossless compression and lossy compression and their combinations can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal based on the compressed original signal. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough to enable the reconstructed signal to be used for the intended application. In the case of video, lossy compression is widely used. The amount of distortion tolerated depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that a higher allowable / tolerable distortion can produce a higher compression ratio.

[0007] Video encoders and decoders may utilize techniques from several broad categories including, for example, motion compensation, transforms, quantization, and entropy coding.

[0008] Video coding techniques may include techniques known as intra-frame coding. In intra-frame coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into sample blocks. When all sample blocks are encoded in intra-frame mode, the picture may be an intra-frame picture. Intra-frame pictures and their derivatives (e.g., independent decoder refresh pictures) may be used to reset decoder states, and may therefore be used as the first picture in an encoded video bitstream and video session or as a still image. Samples of intra-frame blocks may be subjected to transformation, and transform coefficients may be quantized before entropy coding. Intra-frame prediction may be a technique for minimizing sample values ​​in a pre-transform domain. In some cases, the smaller the DC value after transformation, and the smaller the AC coefficient, the fewer bits are required to represent the block after entropy coding at a given quantization step size.

[0009] Conventional intra-frame coding, such as that known from, for example, the MPEG-2 generation of coding techniques, does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to do so based on metadata and / or surrounding sample data obtained during encoding and / or decoding of, for example, spatially adjacent and preceding data blocks in decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. Note that, in at least some cases, intra-frame prediction uses only reference data from the current picture being reconstructed, and not reference data from reference pictures.

[0010] There can be many different forms of intra-frame prediction. When more than one such technique can be used in a given video coding technique, the technique used can be encoded in the intra-frame prediction mode. In some cases, a mode can have sub-modes and / or parameters, and these sub-modes and / or parameters can be encoded separately or included in the mode codeword. What codeword is used for a given mode, sub-mode and / or parameter combination can have an impact on the coding efficiency gain through intra-frame prediction, and therefore the entropy coding technique used to convert the codeword into a bitstream can also have an impact on it.

[0011] Some modes of intra prediction were introduced with H.264, refined in H.265, and further refined in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). Neighboring sample values ​​belonging to samples that are already available can be used to form a prediction block. Sample values ​​of neighboring samples are copied into the prediction block according to the direction. A reference to the direction used can be encoded in the bitstream, or it can itself be predicted. Summary of the invention

[0012] According to an embodiment, a method for video decoding is provided. The method includes: decoding encoding information of a plurality of blocks from an encoded video bitstream, the encoding information including advanced control flags for the plurality of blocks, the advanced control flag indicating whether a plurality of encoding tools are disabled for at least one of the plurality of blocks, the at least one of the plurality of blocks including a current block; determining whether the plurality of encoding tools are disabled for at least one of the plurality of blocks based on the advanced control flag; and reconstructing the current block without the plurality of encoding tools based on the plurality of encoding tools being determined to be disabled.

[0013] According to an embodiment, a method for video encoding is provided, the method comprising: determining encoding information of a plurality of blocks, the encoding information comprising advanced control flags for the plurality of blocks, the advanced control flags indicating whether a plurality of encoding tools are disabled for at least one of the plurality of blocks, the at least one of the plurality of blocks comprising a current block; determining whether a plurality of encoding tools are disabled for at least one of the plurality of blocks; and reconstructing the current block without the plurality of encoding tools based on the plurality of encoding tools being determined to be disabled.

[0014] According to an embodiment, a computer device is provided, which includes: one or more computer-readable non-transitory storage media configured to store computer program codes; and one or more computer processors configured to access the computer program codes and execute the above method for video decoding according to the instructions of the computer program codes.

[0015] According to an embodiment, a device for video decoding is provided, comprising a processing circuit. The processing circuit is configured to execute the above method for video decoding.

[0016] According to an embodiment, a non-transitory computer-readable storage medium is provided, which stores instructions that, when executed by a computer, cause the computer to perform the above method for video decoding.

[0017] According to the video decoding method, computer device, apparatus and non-transitory computer readable medium of the present invention, the syntax structure in HLS is designed to simultaneously disable multiple encoding tools, for example, by disabling syntax elements of multiple encoding tools. Disabling multiple encoding tools simultaneously rather than individually can save encoder running time and thus improve encoding efficiency. 。 BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0019] FIG. 1A is a diagram illustrating an exemplary subset of intra prediction modes.

[0020] FIG. 1B is a diagram of exemplary intra prediction directions.

[0021] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merging candidates in an example.

[0022] Figure 3 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.

[0023] Figure 4 is a schematic diagram of a simplified block diagram of a communication system according to another embodiment.

[0024] Figure 5 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.

[0025] Figure 6 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.

[0026] Figure 7 A block diagram of an encoder according to another embodiment is shown.

[0027] Figure 8 A block diagram of a decoder according to another embodiment is shown.

[0028] Fig. 9 An example of intra block copying according to an embodiment of the present disclosure is shown.

[0029] Fig.10 An example of intra block copy according to another embodiment of the present disclosure is shown.

[0030] Fig.11 An example of intra block copying according to still another embodiment of the present disclosure is shown.

[0031] FIG. 12A to FIG. 12D An example of intra block copying according to yet another embodiment of the present disclosure is shown.

[0032] Fig.13 An example of a string copy mode according to an embodiment of the present disclosure is shown.

[0033] Fig.14 An exemplary bitstream structure according to an embodiment of the present disclosure is shown.

[0034] FIG. 15A to FIG. 15B An exemplary syntax table (1500) according to an implementation of the present disclosure is shown.

[0035] Fig.16 A syntax table in a sequence parameter set (SPS) according to an embodiment of the present disclosure is shown.

[0036] FIG. 17A to FIG. 17B An exemplary syntax for block-level flags according to an embodiment of the present disclosure is shown.

[0037] Fig.18 A flow chart outlining a process (1800) according to an implementation of the present disclosure is shown.

[0038] Fig.19 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION

[0039] 1A, a subset of nine prediction directions known from the 33 possible prediction directions of H.265 (corresponding to 33 angular modes in 35 intra modes) is depicted at the bottom right. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted based on one or more samples at an angle of 45 degrees to the horizontal direction at the top right. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more samples at an angle of 22.5 degrees to the horizontal direction at the bottom left of sample (101).

[0040] Still referring to FIG. 1A , a square block (104) of 4×4 samples is depicted at the upper left (indicated by the bold dashed line). The square block (104) includes 16 samples, each of which is labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in both the Y and X dimensions in block (104). Since the size of the block is 4×4 samples, S44 is at the lower right. Also shown are reference samples that follow a similar numbering scheme. The reference samples are labeled with R, their Y position (e.g., row index), and X position (column index) relative to the block (104). In both H.264 and H.265, the prediction samples are adjacent to the block being reconstructed; therefore, there is no need to use negative values.

[0041] Intra-picture prediction can work by appropriately copying reference sample values ​​from neighboring samples according to the signaled prediction direction. For example, assume that the encoded video bitstream includes signaling that, for this block, the signaling indicates a prediction direction consistent with arrow (102) - that is, the sample is predicted based on one or more prediction samples to the upper right at a 45 degree angle to the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Sample S44 is then predicted based on reference sample R08.

[0042] In some cases, the values ​​of multiple reference samples may be combined, such as by interpolation, in order to compute a reference sample; in particular, when the direction is not divisible by 45 degrees.

[0043] As video coding techniques have evolved, the number of possible directions has increased. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions when disclosed. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent these possible directions with a small number of bits, accepting certain penalties for less likely directions. In addition, the direction itself can sometimes be predicted based on neighboring directions used in adjacent decoded blocks.

[0044] FIG. 1B shows a schematic diagram ( 180 ) depicting 65 intra prediction directions according to JEM to illustrate that the number of prediction directions increases over time.

[0045] The mapping of intra-prediction direction bits representing directions in the coded video bitstream may vary from one video coding technique to another; and the mapping may range, for example, from a simple direct mapping of prediction directions to intra-prediction modes, to codewords, to complex adaptive schemes involving most probable modes, and the like. However, in all cases, there may be some directions that are statistically less likely to appear in the video content than some other directions. Since the goal of video compression is to reduce redundancy, in a well-functioning video coding technique, those less probable directions will be represented by a larger number of bits than more probable directions.

[0046] Motion compensation may be a lossy compression technique and may involve the following technique: a block of sample data from a previously reconstructed picture or portion thereof (reference picture) is used to predict a newly reconstructed picture or portion of a picture after being spatially shifted in a direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, the third dimension being an indication of the reference picture in use (the third dimension may indirectly be a temporal dimension).

[0047] In some video compression techniques, an MV applicable to a particular region of sample data may be predicted based on other MVs, for example, based on an MV of the sample data that is spatially related to another region adjacent to the region being reconstructed and that precedes the MV in decoding order. Doing so can significantly reduce the amount of data required to encode the MV, thereby eliminating redundancy and increasing compression. MV prediction can work effectively, for example, because when encoding an input video signal derived from a camera (called natural video), there is a statistical possibility that regions larger than the region to which a single MV applies move in similar directions, and thus similar motion vectors derived from MVs of adjacent regions can be used for prediction in some cases. This makes the MV obtained for a given region similar or identical to the MV predicted from the surrounding MVs, and it can in turn be represented with fewer bits after entropy coding than would be used if the MV was encoded directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself can be lossy, for example due to rounding errors when calculating a predictor based on several surrounding MVs.

[0048] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). Described here is a technique referred to as "spatial merging" among the various MV prediction mechanisms provided by H.265.

[0049] 2, the current block (201) includes samples obtained by the encoder during the motion search process that can be predicted from a previous block of the same size that has been spatially shifted. Instead of encoding the MV directly, the MV can be derived from metadata associated with one or more reference pictures, for example, the most recent (in decoding order) reference picture, using the MV associated with any of the five surrounding samples represented by A0, A1 and B0, B1, B2 (202 to 206, respectively). In H.265, MV prediction can use a predictor from the same reference picture that a neighboring block is also using.

[0050] Figure 3 A simplified block diagram of a communication system (300) according to an embodiment of the present disclosure is shown. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). Figure 3In the example of , a first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can encode video data (e.g., a video picture stream captured by the terminal device (310)) for transmission to another terminal device (320) via a network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video picture, and display the video picture based on the recovered video data. Unidirectional data transmission can be common in media service applications and the like.

[0051] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340), which perform bidirectional transmission of encoded video data that may occur, for example, during a video conference. For the bidirectional transmission of data, in the example, each of the terminal devices (330) and (340) can encode video data (e.g., a video picture stream captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), and can decode the encoded video data to restore the video picture, and can display the video picture at an accessible display device based on the restored video data.

[0052] exist Figure 3 In the example of , terminal devices (310), (320), (330) and (340) can be shown as servers, personal computers and smart phones, but the principles of the present disclosure may not be so limited. Implementations of the present disclosure are applicable to laptop computers, tablet computers, media players and / or dedicated video conferencing equipment. Network (350) represents any number of networks that transmit encoded video data between terminal devices (310), (320), (330) and (340), including, for example, wired (wired) and / or wireless communication networks. Communication network (350) can exchange data in circuit switching channels and / or packet switching channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purposes of this discussion, unless otherwise specified below, the architecture and topology of network (350) may be irrelevant to the operation of the present disclosure.

[0053] As examples of applications of the disclosed subject matter, Figure 4The arrangement of a video encoder and a video decoder in a streaming environment is shown. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example: video conferencing; digital TV; storing compressed video on digital media including CDs, DVDs, memory sticks, etc., etc.

[0054] The streaming system may include a capture subsystem (413), which may include a video source (401), such as a digital camera, that creates, for example, an uncompressed video picture stream (402). In an example, the video picture stream (402) includes samples captured by the digital camera. The video picture stream (402) is depicted as a thick line to emphasize the high amount of data when compared to the encoded video data (404) (or encoded video bitstream), and the video picture stream (402) may be processed by an electronic device (420) coupled to the video source (401) including a video encoder (403). The video encoder (403) may include hardware, software, or a combination thereof to implement or implement aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)) is depicted as a thin line to emphasize the lower amount of data when compared to the video picture stream (402), and the encoded video data (404) can be stored on the streaming server (405) for future use. One or more streaming client subsystems (e.g. Figure 4 The client subsystems (406) and (408) in the video transmission system can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and creates an outgoing video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or other presentation device (not depicted). In some streaming systems, the encoded video data (404), (407) and (409) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In the example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0055] Note that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0056] Figure 5 A block diagram of a video decoder (510) according to an embodiment of the present disclosure is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used instead of Figure 4 A video decoder (410) in an example of FIG.

[0057] A receiver (531) may receive one or more encoded video sequences to be decoded by a video decoder (510); in the same or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data and other data such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and an entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other applications, the buffer memory (515) may be external to the video decoder (510) (not depicted). In still other applications, there may be a buffer memory (not depicted) external to the video decoder (510) to, for example, prevent network jitter, and further there may be another buffer memory (515) internal to the video decoder (510) to, for example, handle playout timing. When the receiver (531) receives data from a store / forward device with sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory (515) may not be required, or the buffer memory (515) may be smaller. For use over a best effort packet network such as the Internet, the buffer memory (515) may be required, the buffer memory (515) may be relatively large and may advantageously have an adaptive size, and may be implemented at least in part in an operating system or similar element (not depicted) external to the video decoder (510).

[0058] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The types of symbols include information for managing the operation of the video decoder (510) and may include information for controlling a presentation device such as a presentation device (512) (e.g., a display screen) that is not part of the electronic device (530) but can be coupled to the electronic device (530), such as Figure 5 As shown. The control information for (one or more) rendering devices may be in the form of a Supplemental Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (520) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be based on a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a subgroup parameter set for at least one subgroup of the pixel subgroup from the coded video sequence in the video decoder based on at least one parameter corresponding to the group. The subgroup may include: Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The parser (520) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0059] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515), thereby creating symbols (521).

[0060] The reconstruction of the symbol (521) may involve a number of different units depending on the type of the coded video picture or portion thereof (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by the parser (520). For clarity, the flow of such subgroup control information between the parser (520) and the various units described below is not depicted.

[0061] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into a number of functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0062] The first unit is a sealer / inverse transform unit (551). The sealer / inverse transform unit (551) receives quantized transform coefficients as (one or more) symbols (521) from the parser (520) and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (551) can output a block including sample values, which can be input into an aggregator (555).

[0063] In some cases, the output samples of the sealer / inverse transform (551) may relate to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding reconstructed information obtained from a current picture buffer (558). For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (555) adds the prediction information already generated by the intra-prediction unit (552) to the output sample information as provided by the sealer / inverse transform unit (551) on a per-sample basis.

[0064] In other cases, the output samples of the scaler / inverse transform unit (551) may be associated with a block that is inter-coded and possibly motion compensated. In such a case, the motion compensated prediction unit (553) may access a reference picture memory (557) to obtain samples for prediction. After the obtained samples are motion compensated according to the symbols (521) associated with the block, these samples may be added to the output of the scaler / inverse transform unit (551) (in this case referred to as residual samples or residual signals) by an aggregator (555) to generate output sample information. The address within the reference picture memory (557) from which the motion compensated prediction unit (553) obtains the predicted samples may be controlled by a motion vector in the form of the symbols (521) that can be obtained by the motion compensated prediction unit (553), which motion vector may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values ​​obtained from the reference picture memory (557) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0065] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as a coded video bitstream) and obtained by the loop filter unit (556) as symbols (521) from the parser (520), but the in-loop filter techniques may also be responsive to meta-information obtained during decoding of a previous portion (in decoding order) of the coded picture or coded video sequence, and to previously reconstructed and loop filtered sample values.

[0066] The output of the loop filter unit (556) may be a sample stream that may be output to a rendering device (512) and stored in a reference picture memory (557) for future inter-picture prediction.

[0067] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557) and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0068] The video decoder (510) may perform decoding operations according to a predetermined video compression technique in a standard such as ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence follows both the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as tools available only under the profile. For compliance, it is also required that the complexity of the encoded video sequence is within the limits defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, millions of samples per second), the maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0069] In an embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence (one or more). The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0070] Figure 6 A block diagram of a video encoder (603) according to an embodiment of the present disclosure is shown. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used instead of Figure 4 A video encoder (403) in an example.

[0071] The video encoder (603) can be used to obtain the video source (601) (which is not Figure 6 In an example of an electronic device (620) receiving video samples, a video source (601) can capture (one or more) video images to be encoded by a video encoder (603). In another example, the video source (601) is a part of the electronic device (620).

[0072] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by a video encoder (603), the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The picture itself may be organized as a spatial pixel array, wherein each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by those skilled in the art. The following description focuses on samples.

[0073] According to an embodiment, the video encoder (603) can encode and compress the pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units described below and is functionally coupled to the other functional units. The coupling is not depicted for clarity. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology...), picture size, group of pictures (group of pictures, GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions related to the video encoder (603) optimized for a specific system design.

[0074] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an extremely simplified description, in an example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and (one or more) reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder would create sample data (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-accurate between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" as reference picture samples exactly the same sample values ​​as the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, eg due to channel errors) is also used in some related techniques.

[0075] The operation of the "local" decoder (633) can be combined with the "remote" decoder such as has been described above. Figure 5 The operation of the video decoder (510) described in detail is the same. However, in addition, briefly refer to Figure 5 , since the symbols are available and the encoding of the symbols into an encoded video sequence by the entropy encoder (645) and the decoding of the symbols by the parser (520) can be lossless, the entropy decoding portion of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).

[0076] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is contrary to the decoder technology described comprehensively. A more detailed description is only needed in certain areas, and the description is provided below.

[0077] In some examples, during operation, the source encoder (630) may perform motion compensated predictive coding that predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as “reference pictures.” In this manner, the encoding engine (632) encodes the differences between pixel blocks of an input picture and pixel blocks of reference picture(s) that may be selected as prediction reference(s) for the input picture.

[0078] The local video decoder (633) can decode the encoded video data of the picture that can be designated as the reference picture based on the symbol created by the source encoder (630). The operation of the encoding engine (632) can advantageously be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 6 When the video encoder (603) is decoded at a remote end (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) may store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture that will be obtained by the remote video decoder.

[0079] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may be used as suitable prediction references for the new picture. The predictor (635) may operate on a sample-block-by-pixel-block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).

[0080] The controller (650) can manage encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0081] The outputs of all the above-mentioned functional units may be subjected to entropy coding in the entropy encoder (645). The entropy encoder (645) converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0082] The transmitter (640) can buffer the encoded video sequence(s) created by the entropy encoder (645) in preparation for transmission via a communication channel (660), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0083] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a specific coded picture type to each coded picture, which may affect the coding techniques that may be applied to the corresponding picture. For example, a picture may generally be assigned one of the following picture types:

[0084] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of those variations of I pictures and their corresponding applications and features.

[0085] A predictive picture (P picture) may be a picture that may be encoded and decoded using inter prediction or intra prediction to predict sample values ​​of each block using at most one motion vector and a reference index.

[0086] Bi-directional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using inter-prediction or intra-prediction using up to two motion vectors and reference indices to predict sample values ​​for each block. Similarly, multi-predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.

[0087] The source picture may typically be spatially subdivided into a number of blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples, respectively) and coded block by block. The blocks may be predictively coded with reference to other (already coded) blocks, which are determined by the coding allocation applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively coded, or may be predictively coded (spatial prediction or intra-prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0088] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In its operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0089] In an embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0090] Video can be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlation in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an example, a particular picture being encoded / decoded, which is referred to as the current picture, is divided into blocks. In the case where a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case of using multiple reference pictures, the motion vector may have a third dimension that identifies the reference picture.

[0091] In some embodiments, bidirectional prediction techniques may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0092] In addition, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0093] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), namely a luminance CTB and two chrominance CTBs. Each CTU can be recursively partitioned into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be partitioned into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In the example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. The CU is partitioned into one or more prediction units (PUs) based on temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, the prediction operation in the codec (encoding / decoding) is performed in units of prediction blocks. Using the luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0094] Figure 7 A diagram of a video encoder (703) according to another embodiment of the present disclosure is shown. The video encoder (703) is configured to receive a processed block (e.g., a predicted block) of sample values ​​within a current video picture in a sequence of video pictures, and encode the processed block into an encoded picture that is part of an encoded video sequence. In an example, the video encoder (703) is used instead of Figure 4A video encoder (403) in an example.

[0095] In the HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) uses, for example, rate-distortion optimization to determine whether to best encode the processing block using intra mode, inter mode, or bidirectional prediction mode. In the case where the processing block is to be encoded in intra mode, the video encoder (703) may encode the processing block into an encoded picture using intra prediction techniques; and in the case where the processing block is to be encoded in inter mode or bidirectional prediction mode, the video encoder (703) may encode the processing block into an encoded picture using inter prediction or bidirectional prediction techniques, respectively. In some video coding techniques, the merge mode may be an inter picture prediction submode in which a motion vector is derived from one or more motion vector predictors without resorting to encoded motion vector components external to the predictors. In some other video coding techniques, there may be motion vector components applicable to the object block. In the example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0096] exist Figure 7 In the example of FIG. 7 , the video encoder ( 703 ) includes Figure 7 An inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), an overall controller (721), and an entropy encoder (725) are shown coupled together.

[0097] The inter-frame encoder (730) is configured to: receive samples of a current block (e.g., a processing block); compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture); generate inter-frame prediction information (e.g., description of redundant information, motion vectors, merge mode information according to an inter-frame coding technique); and calculate an inter-frame prediction result (e.g., a prediction block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the encoded video information.

[0098] The intra encoder (722) is configured to: receive samples of a current block (e.g., a processing block); in some cases compare the block with an already encoded block in the same picture; generate quantized coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). In an example, the intra encoder (722) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.

[0099] The overall controller (721) is configured to determine overall control data and control other components of the video encoder (703) based on the overall control data. In an example, the overall controller (721) determines a mode of a block and provides a control signal to a switch (726) based on the mode. For example, when the mode is an intra-frame mode, the overall controller (721) controls the switch (726) to select an intra-frame mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select intra-frame prediction information and include the intra-frame prediction information in the bitstream; and when the mode is an inter-frame mode, the overall controller (721) controls the switch (726) to select an inter-frame prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select inter-frame prediction information and include the inter-frame prediction information in the bitstream.

[0100] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra-frame encoder (722) or the inter-frame encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate a transform coefficient. In an example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate a transform coefficient. The transform coefficient is then subjected to quantization to obtain a quantized transform coefficient. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra-frame encoder (722) and the inter-frame encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. In some examples, the decoded blocks are appropriately processed to generate decoded pictures, and these decoded pictures can be buffered in a memory circuit (not shown) and used as reference pictures.

[0101] The entropy encoder (725) is configured to format the bitstream to include the encoded blocks. The entropy encoder (725) is configured to include various information according to a suitable standard, such as the HEVC standard. In an example, the entropy encoder (725) is configured to include overall control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. Note that according to the disclosed subject matter, when the block is encoded in the inter-frame mode or the merge sub-mode of the bidirectional prediction mode, there is no residual information.

[0102] Figure 8A diagram of a video decoder (810) according to another embodiment of the present disclosure is shown. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an example, the video decoder (810) is used instead of Figure 4 A video decoder (410) in an example of FIG.

[0103] exist Figure 8 In the example of FIG. 8 , the video decoder ( 810 ) includes Figure 8 An entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) are shown coupled together.

[0104] The entropy decoder (871) can be configured to reconstruct certain symbols from the encoded picture, which represent syntax elements constituting the encoded picture. Such symbols may include, for example, a mode for encoding a block (e.g., an intra-frame mode, an inter-frame mode, a bidirectional prediction mode, a merged sub-mode of the latter two, or another sub-mode), prediction information (e.g., intra-frame prediction information or inter-frame prediction information) that can identify a certain sample or metadata for use by an intra-frame decoder (872) or an inter-frame decoder (880) for prediction, residual information in the form of quantized transform coefficients, etc. In an example, when the prediction mode is an inter-frame mode or a bidirectional prediction mode, the inter-frame prediction information is provided to the inter-frame decoder (880); and when the prediction type is an intra-frame prediction type, the intra-frame prediction information is provided to the intra-frame decoder (872). The residual information can be subjected to inverse quantization and provided to the residual decoder (873).

[0105] The inter-frame decoder (880) is configured to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.

[0106] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0107] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (Quantizer Parameter, QP)), and this information may be provided by the entropy decoder (871) (the data path is not depicted because this may only be a small amount of control information).

[0108] The reconstruction module (874) is configured to combine the residual output by the residual decoder (873) with the prediction result (output by the inter-frame prediction module or the intra-frame prediction module according to the situation) in the spatial domain to form a reconstructed block, which can be part of a reconstructed picture, which can be part of a reconstructed video. Note that other suitable operations such as deblocking operations can be performed to improve visual quality.

[0109] Note that the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using any suitable technology. In an embodiment, the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using one or more processors that execute software instructions.

[0110] Aspects of the present disclosure include techniques for advanced syntax control of encoding tools, such as screen content encoding.

[0111] Block-based compensation can be used for inter-frame prediction and intra-frame prediction. For inter-frame prediction, block-based compensation based on different pictures is called motion compensation. Block-based compensation can also be done, for example, in intra-frame prediction based on previously reconstructed areas within the same picture. Block-based compensation based on reconstructed areas within the same picture is called intra-picture block compensation, current picture referencing (CPR) or intra-block copy (IBC). The displacement vector indicating the offset between the current block and the reference block (also called the prediction block) in the same picture is called a block vector (BV), where the current block can be encoded / decoded based on the reference block. Unlike the motion vector in motion compensation, which can be any value (positive or negative in the x or y direction), BV has some constraints to ensure that the reference block is available and has been reconstructed. In addition, in some examples, some reference areas that are tile boundaries, slice boundaries, or wavefront trapezoid boundaries are excluded for parallel processing considerations.

[0112] The encoding of the block vector can be explicit or implicit. In explicit mode, the BV difference between the block vector and its predictor is signaled. In implicit mode, the block vector is recovered from the predictor (called the block vector predictor) in a similar manner to the motion vector in merge mode, without using the BV difference. The explicit mode may be referred to as the non-merged BV prediction mode. The implicit mode may be referred to as the merged BV prediction mode.

[0113] In some implementations, the resolution of the block vectors is limited to integer positions. In other systems, the block vectors are allowed to point to fractional positions.

[0114] In some examples, a block level flag, such as an IBC flag, may be used to signal the use of intra block copy at the block level. In an embodiment, the block level flag is signaled when the current block is explicitly encoded. In some examples, a reference index method may be used to signal the use of intra block copy at the block level. The current picture in decoding is then considered as a reference picture or a special reference picture. In an example, such a reference picture is placed in the last position of the reference picture list. The special reference picture and other temporal reference pictures are also managed in a buffer, such as a decoded picture buffer (DPB).

[0115] There may be variations for the IBC mode. In the example, the IBC mode is treated as a third mode different from the intra prediction mode and the inter prediction mode. Therefore, the BV prediction in the implicit mode (or merge mode) and the explicit mode is separated from the conventional inter mode. A separate merge candidate list can be defined for the IBC mode, and in the IBC mode, the entries in the separate merge candidate list are BVs. Similarly, in the example, the BV prediction candidate list in the IBC explicit mode includes only BVs. The general rule applied to these two lists (i.e., a separate merge candidate list and a BV prediction candidate list) is that in terms of candidate derivation processing, these two lists can follow the same logic as the merge candidate list used in the conventional merge mode or the AMVP predictor list used in the conventional AMVP mode. For example, five spatially adjacent positions (e.g., A0, A1 and B0, B1, B2 in Figure 2) are accessed for the IBC mode, such as the HEVC or VVC inter-frame merge mode, to derive a separate merge candidate list for the IBC mode.

[0116] As described above, the BV of the current block being reconstructed in the picture may have certain constraints, and therefore, the reference block of the current block is within the search range. The search range refers to the portion of the picture from which the reference block can be selected. For example, the search range may be within certain portions of the reconstructed area in the picture. The size, position, shape, etc. of the search range may be limited. Alternatively, the BV may be constrained. In the example, the BV is a two-dimensional vector including an x ​​component and a y component, and at least one of the x component and the y component may be constrained. Constraints may be specified for the BV, the search range, or a combination of the BV and the search range. In various examples, when certain constraints are specified for the BV, the search range is constrained accordingly. Similarly, when certain constraints are specified for the search range, the BV is constrained accordingly.

[0117] Fig. 9 An example of intra-block copying according to an implementation of the present disclosure is shown. The current picture (900) will be reconstructed during decoding. The current picture (900) includes a reconstructed area (910) (gray area) and an area to be decoded (920) (white area). The current block (930) is being reconstructed by the decoder. The current block (930) can be reconstructed based on a reference block (940) that is in the reconstructed area (910). The position offset between the reference block (940) and the current block (930) is called a block vector (950) (or BV (950)). In Fig. 9 In the example of , the search range (960) is within the reconstruction region (910), the reference block (940) is within the search range (960), and the block vector (950) is constrained to point to the reference block (940) within the search range (960).

[0118] Various constraints may be applied to the BV and / or the search range. In an embodiment, the search range of the current block being reconstructed in the current CTB is restricted to within the current CTB.

[0119] In an embodiment, the effective memory requirement for storing reference samples to be used in intra-frame block copy is one CTB size. In the example, the CTB size is 128×128 samples. The current CTB includes the current region under reconstruction. The size of the current region is 64×64 samples. Since the reference memory can also store the reconstructed samples in the current region, the reference memory can store 3 more regions of 64×64 samples when the reference memory size is equal to the CTB size of 128×128 samples. Therefore, the search range can include some parts of the previously reconstructed CTB, while the total memory requirement for storing reference samples remains unchanged (for example, 1 CTB size with 128×128 samples or a total of 4 64×64 reference samples). In the example, for example Fig.10 As shown, the previously reconstructed CTB is the left neighbor of the current CTB.

[0120] Fig.10An example of intra-block copying according to an embodiment of the present disclosure is shown. The current picture (1001) includes a current CTB (1015) being reconstructed and a previously reconstructed CTB (1010) which is a left neighbor of the current CTB (1015). The CTBs in the current picture (1001) have a CTB size such as 128×128 and a CTB width such as 128 samples. The current CTB (1015) includes 4 regions (1016) to (1019), where the current region (1016) is under reconstruction. The current region (1016) includes multiple coding blocks (1021) to (1029). Similarly, the previously reconstructed CTB (1010) includes 4 regions (1011) to (1014). The coding blocks (1021) to (1025) have been reconstructed, the current block (1026) is under reconstruction, and the coding blocks (1026) to (1027) and the regions (1017) to (1019) are to be reconstructed.

[0121] The current region (1016) has a collocated region (i.e., region (1011) in the previously reconstructed CTB (1010)). The relative position of the collocated region (1011) with respect to the previously reconstructed CTB (1010) may be the same as the relative position of the current region (1016) with respect to the current CTB (1015). Fig.10 In the example shown, the current region (1016) is the upper left region in the current CTB (1015), and therefore, the collocated region (1011) is also the upper left region in the previously reconstructed CTB (1010). Since the position of the previously reconstructed CTB (1010) is offset by the CTB width from the position of the current CTB (1015), the position of the collocated region (1011) is offset by the CTB width from the position of the current region (1016).

[0122] In an embodiment, the collocated region of the current region (1016) is in a previously reconstructed CTB, wherein the position of the previously reconstructed CTB is offset from the position of the current CTB (1015) by one or more CTB widths, and thus the position of the collocated region is also offset from the position of the current region (1016) by a corresponding one or more CTB widths. The position of the collocated region may be shifted left, up, etc. from the current region (1016).

[0123] As described above, the size of the search range of the current block (1026) is limited by the CTB size. In some embodiments, the size of the search range can be limited to other sizes. Fig.10In the example of , the search range may include the area (1012) to (1014) in the previously reconstructed CTB (1010), and the reconstructed portion of the current area (1016), such as coding blocks (1021) to (1025). The search range further excludes the collocated area (1011), so that the size of the search range is within the CTB size. Fig.10 , the reference block (1091) is located in the region (1014) of the previously reconstructed CTB (1010). The block vector (1020) indicates the offset between the current block (1026) and the corresponding reference block (1091). The reference block (1091) is within the search range.

[0124] Fig.10 The example shown can be appropriately applied to other scenarios where the current area is located at another position in the current CTB (1015). In the example, when the current block is in area (1017), the collocated area of ​​the current block is area (1012). Therefore, the search range can include areas (1013) to (1014), area (1016), and the reconstructed portion of area (1017). The search range further excludes area (1011) and the collocated area (1012), so that the size of the search range is within the CTB size. In the example, when the current block is in area (1018), the collocated area of ​​the current block is area (1013). Therefore, the search range can include areas (1014), areas (1016) to (1017), and the reconstructed portion of area (1018). The search range further excludes areas (1011) to (1012) and the collocated area (1013), so that the size of the search range is within the CTB size. In the example, when the current block is in region (1019), the collocated region of the current block is region (1014). Therefore, the search range may include regions (1016) to (1018) and the already reconstructed portion of region (1019). The search range further excludes the previously reconstructed CTB (1010), so that the size of the search range is within the CTB size.

[0125] In the above description, the reference block may be in a previously reconstructed CTB (1010) or a current CTB (1015).

[0126] Fig.11An example of intra-block copying according to an embodiment of the present disclosure is shown. The current picture (1101) includes a current CTB (1115) being reconstructed and a previously reconstructed CTB (1110) which is a left neighbor of the current CTB (1115). The CTB in the current picture (1101) has a CTB size and a CTB width. The current CTB (1115) includes 4 regions (1116) to (1119), wherein the current region (1116) is being reconstructed. The current region (1116) includes a plurality of coding blocks (1121) to (1129). Similarly, the previously reconstructed CTB (1110) includes 4 regions (1111) to (1114). In the current region (1116), the current block (1121) being reconstructed is first to be reconstructed, and coding blocks (1122) to (1129) are to be reconstructed. In the example, the CTB size is 128×128 samples, and each of the regions (1111) to (1114) and (1116) to (1119) is 64×64 samples. The reference memory size is equal to the CTB size and is 128×128 samples, and therefore, when limited by the reference memory size, the search range includes 3 regions and a portion of the additional region.

[0127] With reference Fig.10 Similar to the description, the current region (1116) has a collocated region (i.e., region (1111) in the previously reconstructed CTB (1110)). The reference block of the current block (1121) may be in region (1111), and thus the search range may include regions (1111) to (1114). For example, when the reference block is in region (1111), the collocated region of the reference block is region (1116), wherein samples in region (1116) are not reconstructed before reconstructing the current block (1121). However, as described with reference to Fig.10 As described, for example, after reconstructing the coding block (1121), the region (1111) is no longer suitable for inclusion in the search range for reconstructing the coding block (1122). Therefore, strict synchronization and timing control of the reference memory buffer will be used, and this strict synchronization and timing control may be challenging.

[0128] According to some embodiments, when a current block is to be reconstructed first in a current region of a current CTB, a search range may exclude a collocated region of the current region in a previously reconstructed CTB, wherein the current CTB and the previously reconstructed CTB are in the same current picture. A block vector may be determined such that a reference block is within a search range excluding the collocated region in the previously reconstructed CTB. In an embodiment, the search range includes coded blocks reconstructed after the collocated region and before the current block in decoding order.

[0129] In the following description, the CTB size may vary and the maximum CTB size is set to be the same as the reference memory size. In an example, the reference memory size or the maximum CTB size is 128×128 samples. These descriptions may be appropriately adapted to other reference memory sizes or maximum CTB sizes.

[0130] In an embodiment, the CTB size is equal to the reference memory size. The previously reconstructed CTB is a left neighbor of the current CTB, the position of the collocated region is offset from the position of the current region by the CTB width, and the coding blocks within the search range are in at least one of: the current CTB and the previously reconstructed CTB.

[0131] FIG. 12A to FIG. 12D An example of intra-block copying according to an embodiment of the present disclosure is shown. FIG. 12A to FIG. 12D , the current picture (1201) includes a current CTB (1215) being reconstructed and a previously reconstructed CTB (1210) which is a left neighbor of the current CTB (1215). The CTB in the current picture (1201) has a CTB size and a CTB width. The current CTB (1215) includes 4 regions (1216) to (1219). Similarly, the previously reconstructed CTB (1210) includes 4 regions (1211) to (1214). In an embodiment, the CTB size is the maximum CTB size and is equal to the reference memory size. In the example, the CTB size and the reference memory size are 128×128 samples, and therefore, regions (1211) to (1214) and (1216) to (1219) each have a size of 64×64 samples.

[0132] exist FIG. 12A to FIG. 12D In the example shown, the current CTB (1215) includes an upper left region, an upper right region, a lower left region, and a lower right region corresponding to regions (1216) to (1219), respectively. The previously reconstructed CTB (1210) includes an upper left region, an upper right region, a lower left region, and a lower right region corresponding to regions (1211) to (1214), respectively.

[0133] Reference Fig. 12A , the current region (1216) is being reconstructed. The current region (1216) may include a plurality of coding blocks (1221) to (1229). The current region (1216) has a collocated region in a previously reconstructed CTB (1210), namely, region (1211). The search range of one of the coding blocks (1221) to (1229) to be reconstructed may exclude the collocated region (1211). The search range may include regions (1212) to (1214) of the previously reconstructed CTB (1210) that are reconstructed after the collocated region (1211) and before the current region (1216) in decoding order.

[0134] Reference Fig. 12A , the position of the collocated region (1211) is offset from the position of the current region (1216) by the CTB width, for example, 128 samples. For example, the position of the collocated region (1211) is shifted to the left by 128 samples from the position of the current region (1216).

[0135] Refer again Fig. 12A , when the current region (1216) is the upper left region of the current CTB (1215), the juxtaposed region (1211) is the upper left region of the previously reconstructed CTB (1210), and the search region excludes the upper left region of the previously reconstructed CTB.

[0136] Reference Fig. 12B , the current region (1217) is being reconstructed. The current region (1217) may include multiple coding blocks (1241) to (1249). The current region (1217) has a collocated region (i.e., region (1212) in the previously reconstructed CTB (1210)). The search range of one of the multiple coding blocks (1241) to (1249) may exclude the collocated region (1212). The search range includes regions (1213) to (1214) of the previously reconstructed CTB (1210) and a region (1216) in the current CTB (1215) that is reconstructed after the collocated region (1212) and before the current region (1217). Due to the limitation of the reference memory size (i.e., one CTB size), the search range further excludes region (1211). Similarly, the position of the collocated region (1212) is offset from the position of the current region (1217) by the CTB width, for example, 128 samples.

[0137] exist Fig. 12B In the example of , the current region (1217) is the upper right region of the current CTB (1215), the juxtaposed region (1212) is also the upper right region of the previously reconstructed CTB (1210), and the search region excludes the upper right region of the previously reconstructed CTB (1210).

[0138] Reference Fig. 12C, the current region (1218) is being reconstructed. The current region (1218) may include multiple coding blocks (1261) to (1269). The current region (1218) has a collocated region (i.e., region (1213)) in a previously reconstructed CTB (1210). The search range of one of the multiple coding blocks (1261) to (1269) may exclude the collocated region (1213). The search range includes region (1214) of the previously reconstructed CTB (1210) and regions (1216) to (1217) in the current CTB (1215) that are reconstructed after the collocated region (1213) and before the current region (1218). Similarly, due to the limitation of the reference memory size, the search range further excludes regions (1211) to (1212). The position of the collocated region (1213) is offset from the position of the current region (1218) by the CTB width, for example, 128 samples. Fig. 12C In the example, when the current region (1218) is the lower left region of the current CTB (1215), the juxtaposed region (1213) is also the lower left region of the previously reconstructed CTB (1210), and the search region excludes the lower left region of the previously reconstructed CTB (1210).

[0139] Reference Fig.12D , the current region (1219) is being reconstructed. The current region (1219) may include a plurality of coding blocks (1281) to (1289). The current region (1219) has a collocated region (i.e., region (1214)) in a previously reconstructed CTB (1210). The search range of one of the plurality of coding blocks (1281) to (1289) may exclude the collocated region (1214). The search range includes regions (1216) to (1218) in the current CTB (1215) that are reconstructed after the collocated region (1214) and before the current region (1219) in decoding order. Due to the limitation of the reference memory size, the search range excludes regions (1211) to (1213), and therefore, the search range excludes the previously reconstructed CTB (1210). Similarly, the position of the collocated region (1214) is offset from the position of the current region (1219) by the CTB width, for example, 128 samples. In Fig.12D In the example, when the current region (1219) is the lower right region of the current CTB (1215), the juxtaposed region (1214) is also the lower right region of the previously reconstructed CTB (1210), and the search region excludes the lower right region of the previously reconstructed CTB (1210).

[0140] Fig.13An example of a string copy mode according to an embodiment of the present disclosure is shown. The string copy mode may also be referred to as a string matching mode, an intra-string copy mode, or a string prediction. The current picture (1310) includes a reconstructed area (gray area) (1320) and an area (1321) in reconstruction. The current block (1335) in the area (1321) is in reconstruction. The current block (1335) may be a CB, a CU, etc. The current block (1335) may include multiple strings (e.g., strings (1330) and (1331)). In the example, the current block (1335) is divided into multiple continuous strings, where one string follows another along the scanning order. The scanning order may be any suitable scanning order, such as a raster scanning order, a horizontal scanning order, or other predefined scanning order.

[0141] The reconstruction area (1320) may be used as a reference area for reconstructing the strings (1330) and (1331).

[0142] For each of the plurality of strings, a string offset vector (referred to as SV) and / or a length of the string (referred to as string length) may be signaled or inferred. An SV (e.g., SV0) may be a displacement vector indicating a displacement between a string to be reconstructed (e.g., string (1330)) and a corresponding reference string (e.g., reference string (1300)) located in an already reconstructed reference region (1320). The reference string may be used to reconstruct the string to be reconstructed. Thus, the SV may indicate where the corresponding reference string is located in the reference region (1320). The string length may also correspond to the length of the reference string. Fig.13 , the current block (1335) is an 8×8 CB including 64 samples, and is divided into two strings (e.g., strings (1330) and (1331)) using a raster scan order. String (1330) includes the first 29 samples of the current block (1335), and string (1331) includes the remaining 35 samples of the current block (1335). The reference string (1300) used to reconstruct the string (1330) can be indicated by the corresponding string vector SV0, and the reference string (1301) used to reconstruct the string (1331) can be indicated by the corresponding string vector SV1.

[0143] In general, string size (also called string length) can refer to the length of the string or the number of samples in the string. Fig.13 , string (1330) includes 29 samples, so the string size or string length of string (1330) is 29. String (1331) includes 35 samples, so the string size or string length of string (1331) is 35. The string location (or string position) can be represented by the sample position of a sample in the string (e.g., the first sample in decoding order).

[0144] The above description can be appropriately adapted to reconstruct a current block comprising any suitable number of strings. In an example, when a sample in the current block does not have a matching sample in a reference region, an escaped sample (or an escaped pixel) is indicated by a signal, and the value of the escaped sample can be directly encoded without reference to the reconstructed sample in the reference region. In an example, a block includes multiple strings and one or more escaped samples, wherein multiple strings are reconstructed using a string copy mode, and one or more escaped samples are directly encoded and not predicted using the string copy mode. One or more escaped samples can be located at any suitable position within the block. In an example, one or more escaped samples in a block are located outside multiple strings.

[0145] High level syntax (HLS) can specify parameters that can be shared by lower level coding layers. For example, the CTU size (also the maximum size of the CB) is specified at the sequence level or in the sequence parameter set (SPS), and does not change from one picture to another. HLS may correspond to a high-level. The high level (or high level) may be higher than the block level. The high level may correspond to a video sequence (or sequence level), one or more pictures, pictures (or picture level), slices (or slice levels), tile groups (or tile group levels), tiles (or tile levels), CTUs (or CTU levels), etc. Typically, HLS may include SPS, picture parameter sets (PPS), picture headers, slice headers, adaptive parameter sets (APS), etc. In the example, HLS corresponds to a tile, tile group, or similar sub-picture level. In the example, HLS corresponds to a CTU.

[0146] Each HLS may have a specific coverage range. For example, a PPS may specify common syntax elements that may be shared by one or more pictures. A picture header may specify common syntax elements used within a picture. A first HLS corresponding to a first level may overwrite syntax elements provided in a second HLS corresponding to a second level, wherein the second level is higher than the first level, and the second range covered by the second HLS includes the first range covered by the first HLS. For example, an HLS for a picture header (also referred to as a picture header HLS) may overwrite syntax elements in a PPS referenced by a current picture, wherein the first HLS is a picture header HLS and the second HLS is a PPS. In an example, a slice header belonging to a current picture may overwrite syntax elements or parameters allocated at a picture header of the same current picture.

[0147] Fig.14An exemplary bitstream structure (1400) according to an embodiment of the present disclosure is shown, the exemplary bitstream structure (1400) including an SPS (1401), picture headers (1411) and (1421), and block-level syntax (e.g., block data (1412(1) to 1412(n)) and block data (1422(1) to 1422(m)). Parameters n and m may be any suitable positive integers. Parameters n and m may be the same or different. Block data (1412(1) to 1412(n)) and block data (1422(1) to 1422(m)) may include data such as a residual of a block and side information (or control information) for encoding the block. Referring to Fig.14 , the SPS (1401) corresponds to a video sequence including a plurality of pictures, for example, a first picture and a second picture. The first picture may include blocks 1 to n and the second picture may include blocks 1 to m.

[0148] The syntax elements for encoding blocks 1 to n and blocks 1 to m may be included in various HLSs such as SPS (1401) and picture headers (1411) and (1421). The syntax elements for encoding blocks 1 to n and blocks 1 to m may also be included in block-level syntax, for example, in block data (1412 (1) to 1412 (n)) and block data (1422 (1) to 1422 (m)), respectively. Therefore, the picture data (1410) of the first picture includes a picture header (1411) followed by block data (1412 (1) to 1412 (n)). The picture data (1420) of the second picture may include a picture header (1421) followed by block data (1422 (1) to 1422 (m)).

[0149] The SPS (1401) may specify common syntax elements that may be shared by a video sequence. The picture header (1411) may specify common syntax elements used within a first picture. The picture header (1421) may specify common syntax elements used within a second picture. In an example, syntax elements in the picture header (1411) overwrite corresponding syntax elements in the SPS (1401). In an example, syntax elements in block data (e.g., (1412(2))) overwrite corresponding syntax elements in the picture header (1411).

[0150] FIG. 15A to FIG. 15BAn exemplary syntax table (1500) is shown according to an embodiment of the present disclosure. The syntax table (1500) may correspond to an appropriate HLS at an appropriate high level. The syntax table (1500) may include flags (or encoding tool enable flags) for encoding tools included in certain defined profiles. One or more of these flags may be indicated by a signal. For example, the syntax table (1500) is an SPS syntax table that includes an SPS. In an example, such as in AVS3, the syntax table (1500) shows a set of SPS flags for encoding tools included in certain defined profiles.

[0151] Each of these flags can be used to enable (e.g., by setting the flag to 1) or disable (e.g., by setting the flag to 0) a coding tool in the bitstream. In the absence of the flag in the bitstream, it can be inferred that the flag is disabled (e.g., the value of the flag is 0). The variable (or coding tool variable) corresponding to the flag can be set equal to the flag. In an example, the variable corresponding to the flag is not indicated by a signal, but is derived from the flag. In an example, the variable corresponding to the flag can be modified without modifying the flag.

[0152] In an example, a first set of coding tool enable flags for encoding content captured by a camera is considered to include, but is not limited to, flags such as eipm_enable_flag, dmvr_enable_flag, bio_enable_flag, affine_umve_enable_flag, etmvp_enable_flag, subtmvp_enable_flag, st_chroma_enable_flag, ipf_chroma_enable_flag, ealf_enable_flag, sp_enable_flag, and iip_enable_flag. The first set of coding tool enable flags may correspond to the first set of coding tools described below. The first set of coding tools may include any suitable coding tools for encoding content captured by a camera. For example, the first set of encoding tools includes, but is not limited to, one or more of the following: extended intra prediction mode, decoder-side motion vector refinement, bidirectional optical flow, final motion vector expression, enhanced temporal motion vector prediction, sub-block based temporal motion vector prediction, chroma secondary transform, chroma intra prediction filtering, enhanced adaptive loop filtering, secondary prediction for affine mode, and improved intra prediction mode.

[0153] Each coding tool enable flag in the first set of coding tool enable flags may indicate whether a corresponding coding tool in the first set of coding tools may be used in the bitstream. For example, eipm_enable_flag corresponds to an extended intra prediction mode. The first set of coding tools may be used to encode content captured by a camera and may be referred to as a camera-captured content coding tool.

[0154] In an example, a second set of coding tool enable flags for encoding screen content is considered to include flags such as ibc_enable_flag, isc_enable_flag, fimc_enable_flag, and ists_enable_flag. The second set of coding tool enable flags may correspond to a second set of coding tools (also referred to as screen content coding (SCC) tools) as described below. Each coding tool enable flag in the second set of coding tool enable flags may indicate whether a corresponding coding tool in the second set of coding tools may be used in a bitstream. For example, ibc_enable_flag corresponds to an intra-block copy mode. The second set of coding tools (or SCC tools) may be used to encode screen content. In an example, the first set of coding tools may be used to encode content captured by a camera more efficiently than encoding screen content.

[0155] The semantics of the subset of flags in the syntax table (1500) are listed as follows, wherein the subset of flags in the syntax table (1500) includes a first set of encoding tool enable flags and a second set of encoding tool enable flags.

[0156] eipm_enable_flag may indicate whether an extended intra prediction mode may be used in a bitstream, for example, at a high level (e.g., at the SPS level of a video sequence). eipm_enable_flag equal to 1 may indicate that an extended intra prediction mode may be used in a bitstream, for example, at a high level (e.g., at the SPS level of a video sequence). eipm_enable_flag equal to 0 may indicate that an extended intra prediction mode may not be used in a bitstream, for example, at a high level. In the case where eipm_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of eipm_enable_flag may be inferred to be 0. The variable (or coding tool variable) EipmEnableFlag may be set equal to eipm_enable_flag. In an example, the variable EipmEnableFlag is not signaled, but derived from eipm_enable_flag. In an example, the variable EipmEnableFlag may be modified without modifying eipm_enable_flag.

[0157] The dmvr_enable_flag may indicate whether decoder-side motion vector refinement may be used in the bitstream, for example, at a high level (e.g., SPS level of a video sequence). The dmvr_enable_flag being equal to 1 may indicate that decoder-side motion vector refinement may be used in the bitstream, for example, at a high level. The dmvr_enable_flag being equal to 0 may indicate that decoder-side motion vector refinement may not be used in the bitstream, for example, at a high level. In the event that the dmvr_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of the dmvr_enable_flag may be inferred to be 0. The variable (or coding tool variable) DmvrEnableFlag may be set equal to dmvr_enable_flag. In an example, the variable DmvrEnableFlag is not signaled, but is derived from the dmvr_enable_flag. In the example, the variable DmvrEnableFlag may be modified without modifying dmvr_enable_flag.

[0158] bio_enable_flag may indicate whether bidirectional optical flow may be used in the bitstream, for example, at a high level (e.g., SPS level of a video sequence). bio_enable_flag equal to 1 may indicate that bidirectional optical flow may be used in the bitstream, for example, at a high level. bio_enable_flag equal to 0 may indicate that bidirectional optical flow may not be used in the bitstream, for example, at a high level. In the case where bio_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of bio_enable_flag may be inferred to be 0. The variable (or coding tool variable) BioEnableFlag may be set equal to bio_enable_flag. In an example, the variable BioEnableFlag is not signaled, but is derived from Bio_Enable_Flag. In an example, the variable BioEnableFlag may be modified without modifying bio_enable_flag.

[0159] affine_umve_enable_flag may indicate whether the final motion vector expression for the affine mode may be used in the bitstream, for example, at a high level (e.g., SPS level of a video sequence). affine_umve_enable_flag equal to 1 may indicate that the final motion vector expression for the affine mode may be used in the bitstream, for example, at a high level. affine_umve_enable_flag equal to 0 may indicate that the final motion vector expression for the affine mode may not be used in the bitstream, for example, at a high level. In the case where affine_umve_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of affine_umve_enable_flag may be inferred to be 0. The variable (or coding tool variable) AffineUmveEnableFlag may be set equal to affine_umve_enable_flag. In an example, the variable AffineUmveEnableFlag is not signaled, but is derived from affine_umve_enable_flag. In the example, the variable AffineUmveEnableFlag can be modified without modifying affine_umve_enable_flag.

[0160] etmvp_enable_flag may indicate whether enhanced temporal motion vector prediction may be used in the bitstream, for example, at a high level (e.g., at the SPS level of a video sequence). etmvp_enable_flag being equal to 1 may indicate that enhanced temporal motion vector prediction for an affine mode may be used in the bitstream, for example, at a high level. etmvp_enable_flag being equal to 0 may indicate that enhanced temporal motion vector prediction may not be used in the bitstream, for example, at a high level. In the event that etmvp_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of etmvp_enable_flag may be inferred to be 0. The variable (or coding tool variable) EtmvpEnableFlag may be set equal to etmvp_enable_flag. In the example, the variable EtmvpEnableFlag is not signaled, but is derived from etmvp_enable_flag. In the example, the variable EtmvpEnableFlag may be modified without modifying etmvp_enable_flag.

[0161] subtmvp_enable_flag may indicate whether sub-block-based temporal motion vector prediction may be used in the bitstream, for example, at a high level (e.g., at the SPS level of a video sequence). subtmvp_enable_flag being equal to 1 may indicate that sub-block-based temporal motion vector prediction for an affine mode may be used in the bitstream, for example, at a high level. subtmvp_enable_flag being equal to 0 may indicate that sub-block-based temporal motion vector prediction may not be used in the bitstream, for example, at a high level. In the event that subtmvp_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of subtmvp_enable_flag may be inferred to be 0. The variable (or coding tool variable) SubTmvpEnableFlag may be set equal to subtmvp_enable_flag. In the example, the variable SubTmvpEnableFlag is not signaled, but is derived from subtmvp_enable_flag. In the example, the variable SubTmvpEnableFlag can be modified without modifying subtmvp_enable_flag.

[0162] st_chroma_enable_flag may indicate whether a chroma secondary transform may be used in the bitstream, for example, at a high level (e.g., at the SPS level of a video sequence). st_chroma_enable_flag being equal to 1 may indicate that a chroma secondary transform may be used in the bitstream, for example, at a high level. st_chroma_enable_flag being equal to 0 may indicate that a chroma secondary transform may not be used in the bitstream, for example, at a high level. In the event that st_chroma_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of st_chroma_enable_flag may be inferred to be 0. The variable (or coding tool variable) StChromaEnableFlag may be set equal to st_chroma_enable_flag. In an example, the variable StChromaEnableFlag is not signaled, but derived from st_chroma_enable_flag. In an example, the variable StChromaEnableFlag may be modified without modifying st_chroma_enable_flag.

[0163] ipf_chroma_enable_flag may indicate whether chroma intra prediction filtering may be used in the bitstream, for example, at a high level (e.g., SPS level of a video sequence). ipf_chroma_enable_flag equal to 1 may indicate that chroma intra prediction filtering may be used in the bitstream, for example, at a high level. ipf_chroma_enable_flag equal to 0 may indicate that chroma intra prediction filtering may not be used in the bitstream, for example, at a high level. In the event that ipf_chroma_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of ipf_chroma_enable_flag may be inferred to be 0. The variable (or coding tool variable) IpfChromaEnableFlag may be set equal to ipf_chroma_enable_flag. In an example, the variable IpfChromaEnableFlag is not signaled, but is derived from ipf_chroma_enable_flag. In the example, the variable IpfChromaEnableFlag can be modified without modifying ipf_chroma_enable_flag.

[0164] ealf_enable_flag may indicate whether enhanced adaptive loop filtering may be used in the bitstream, for example, at a high level (e.g., SPS level of a video sequence). ealf_enable_flag equal to 1 may indicate that enhanced adaptive loop filtering may be used in the bitstream, for example, at a high level. ealf_enable_flag equal to 0 may indicate that enhanced adaptive loop filtering may not be used in the bitstream, for example, at a high level. In the case where ealf_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of ealf_enable_flag (or coding tool variable) may be inferred to be 0. The variable EalfEnableFlag may be set equal to ealf_enable_flag. In an example, the variable EalfEnableFlag is not signaled, but is derived from ealf_enable_flag. In an example, the variable EalfEnableFlag may be modified without modifying ealf_enable_flag.

[0165] sp_enable_flag may indicate whether secondary prediction for affine mode may be used in the bitstream, for example, at a high level (e.g., SPS level of a video sequence). sp_enable_flag being equal to 1 may indicate that secondary prediction for affine mode may be used in the bitstream, for example, at a high level. sp_enable_flag being equal to 0 may indicate that secondary prediction for affine mode may not be used in the bitstream, for example, at a high level. In the event that sp_enable_flag does not exist (e.g., not signaled in the syntax table (1500)), the value of sp_enable_flag may be inferred to be 0. The variable (or coding tool variable) SpEnableFlag may be set equal to sp_enable_flag. In an example, the variable SpEnableFlag is not signaled, but is derived from sp_enable_flag. In an example, the variable SpEnableFlag may be modified without modifying sp_enable_flag.

[0166] iip_enable_flag may indicate whether an improved intra prediction mode may be used in a bitstream, for example, at a high level (e.g., an SPS level of a video sequence). iip_enable_flag being equal to 1 may indicate that an improved intra prediction mode may be used in a bitstream, for example, at a high level. iip_enable_flag being equal to 0 may indicate that an improved intra prediction mode may not be used in a bitstream, for example, at a high level. In the event that iip_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of iip_enable_flag may be inferred to be 0. The variable (or coding tool variable) ipEnableFlag may be set equal to iip_enable_flag. In an example, the variable ipEnableFlag is not signaled, but derived from iip_enable_flag. In an example, the variable ipEnableFlag may be modified without modifying iip_enable_flag.

[0167] ibc_enable_flag may indicate whether intra block copy (IBC) mode may be used in the bitstream, for example, at a high level (e.g., SPS level of a video sequence). ibc_enable_flag equal to 1 may indicate that the IBC mode may be used in the bitstream, for example, at a high level. ibc_enable_flag equal to 0 may indicate that the IBC mode may not be used in the bitstream, for example, at a high level. In the case where ibc_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of ibc_enable_flag may be inferred to be 0. The variable (or coding tool variable) IbcEnableFlag may be set equal to ibc_enable_flag. In an example, the variable IbcEnableFlag is not signaled, but derived from ibc_enable_flag. In an example, the variable IbcEnableFlag may be modified without modifying ibc_enable_flag.

[0168] isc_enable_flag may indicate whether the intra-string copy mode may be used in the bitstream, for example, at a high level (e.g., the SPS level of a video sequence). isc_enable_flag being equal to 1 may indicate that the intra-string copy mode may be used in the bitstream, for example, at a high level. isc_enable_flag being equal to 0 may indicate that the intra-string copy mode may not be used in the bitstream, for example, at a high level. In the case where isc_enable_flag does not exist (e.g., not indicated by a signal in the syntax table (1500)), the value of isc_enable_flag may be inferred to be 0. The variable (or coding tool variable) IscEnableFlag may be set equal to isc_enable_flag. In an example, the variable IscEnableFlag is not indicated by a signal, but is derived from isc_enable_flag. In an example, the variable IscEnableFlag may be modified without modifying isc_enable_flag.

[0169] fimc_enable_flag may indicate whether frequency-based intra-mode coding may be used in the bitstream, for example, at a high level (e.g., SPS level of a video sequence). fimc_enable_flag equal to 1 may indicate that frequency-based intra-mode coding may be used in the bitstream, for example, at a high level. fimc_enable_flag equal to 0 may indicate that frequency-based intra-mode coding may not be used in the bitstream, for example, at a high level. In the case where fimc_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of fimc_enable_flag may be inferred to be 0. The variable (or coding tool variable) FimcEnableFlag may be set equal to fimc_enable_flag. In an example, the variable FimcEnableFlag is not signaled, but is derived from fimc_enable_flag. In an example, the variable FimcEnableFlag may be modified without modifying fimc_enable_flag.

[0170] ists_enable_flag may indicate whether an implicitly signaled transform skip mode may be used in the bitstream, e.g., at a high level (e.g., SPS level of a video sequence). ists_enable_flag equal to 1 may indicate that an implicitly signaled transform skip mode may be used in the bitstream, e.g., at a high level. ists_enable_flag equal to 0 may indicate that an implicitly signaled transform skip mode may not be used in the bitstream, e.g., at a high level. In the case that ists_enable_flag is not present (e.g., not signaled in the syntax table (1500)), the value of ists_enable_flag may be inferred to be 0. The variable (or coding tool variable) IstsEnableFlag may be set equal to ists_enable_flag. In an example, the variable IstsEnableFlag is not signaled, but is derived from ists_enable_flag. In an example, the variable IstsEnableFlag may be modified without modifying ists_enable_flag.

[0171] In some examples, whether to signal an enable flag (eg, affine_umve_enable_flag) of a corresponding coding tool may depend on one or more additional conditions. The one or more additional conditions may include whether other coding tools may be used in the bitstream. Fig.15A The box (1510) in shows whether affine_umve_enable_flag is signaled depends on two variables AffineEnableFlag and UmveEnableFlag. In an example, the variable AffineEnableFlag can indicate whether the affine mode can be used in the bitstream, while the variable UmveEnableFlag can indicate whether the final motion vector expression can be used in the bitstream.

[0172] As described above, the syntax table (1500) can correspond to an appropriate HLS at an appropriate high level. For example, the syntax table (1500) can be an SPS syntax table. Similarly, a picture-level enable flag can be signaled in a picture header (or slice header) of a picture to determine whether the coding tools associated with the picture-level enable flag can be used for the picture.

[0173] The value of a lower-level enable flag (e.g., a picture-level enable flag) of a coding tool may depend on the value of a higher-level enable flag (e.g., a sequence-level enable flag or an SPS flag) of the same coding tool. In some examples, the value of a lower-level enable flag (e.g., a picture-level enable flag) of a coding tool may depend on the value of a higher-level enable flag (e.g., a sequence-level enable flag or an SPS flag) of the same coding tool and one or more additional conditions. As described above with reference to Fig.15A As described in block (1510) of , the one or more additional conditions may include whether other coding tools may be used in the bitstream.

[0174] Therefore, if a higher level enable flag (e.g., a sequence level enable flag or an SPS flag) indicates that the coding tool cannot be used in the bitstream, the value of the corresponding picture level enable flag is not signaled, and the value of the corresponding picture level enable flag can be inferred to be 0. Alternatively, the corresponding picture level enable flag can be signaled, and the value of the signaled corresponding picture level enable flag is 0.

[0175] The above description may be appropriately adapted to any suitable higher level, such as a PPS, slice, tile group, tile, CTU, or any suitable sub-picture level higher than the block level. In an example, a slice level enable flag may be signaled in a slice header of a slice in a picture to determine whether a coding tool associated with the slice level enable flag may be used for the slice.

[0176] In some examples, such as in most application scenarios, certain common syntax elements signaled at each slice header of each slice in a picture may be placed in the picture header of the picture if these common syntax elements do not change from one slice to another.

[0177] When multiple encoding tools are not effective for encoding a certain type of video content (e.g., video, picture), the multiple encoding tools can be disabled. In an example, each of the multiple encoding tools is disabled individually. On the other hand, it is challenging to identify the relationship between the availability of encoding tools and a specific video or type of video.

[0178] Screen content may refer to the following types of content (e.g., video content): the content includes computer-generated content, such as computer-generated text, graphics, animations, and / or similar content. Screen content video may include the screen content described above, such as computer-generated text, graphics, animations, and / or similar content. In some examples, screen content may refer to a mixture of computer-generated content (e.g., computer-generated text, graphics, animations, and / or similar content) and camera-captured content (e.g., camera-captured video). In some examples, in addition to computer-generated content, screen content video may also include camera-captured content (e.g., camera-captured video).

[0179] Screen content video may exhibit different characteristics compared to video captured by a camera. In various embodiments, screen content video (e.g., including computer-generated content) may be non-noisy, have sharp edges, and / or have multiple repeating patterns. For example, repeating patterns in text- and graphics-rich content frequently appear in the same picture of screen content video. Using previously reconstructed blocks in a picture with the same or similar patterns as predictors to predict blocks in a picture may effectively reduce prediction errors and thereby improve the coding efficiency of the screen content video.

[0180] Since screen content may have characteristics different from those of content captured by a camera, video coding tools developed primarily for video captured by a camera may not be as effective when applied to screen content video. In addition, SCC tools may be developed to encode screen content video more efficiently than video coding tools developed primarily for video captured by a camera. For a typical screen content video with text and graphics, the same picture may include repeated patterns. Therefore, referring to Figures 9 to 11 as well as FIG. 12A to FIG. 12D The IBC model described and referenced Fig.13 The string replication mode described may be effective.

[0181] SCC tools may include frequency-based intra-mode coding and the transform skip (TS) mode suitable for screen content.

[0182] In the recently developed video coding standards including AVS3, SCC tools supporting screen content are added. SCC tools may include any suitable coding tools for encoding screen content. For example, SCC tools include a combination of IBC mode, string copy mode, frequency-based intra-frame mode coding, TS mode, etc. SCC tools can effectively encode screen content videos. As described above, coding tools designed to process camera-captured content or video may be invalid for compressing screen content videos. In the case where the coding tools designed to process camera-captured content or video are invalid when encoding screen content videos, the coding tools designed to process camera-captured content can be turned off, for example, to save encoder running time. In various examples, the advanced enable flag of the coding tool will be disabled separately. Therefore, it is advantageous to design the syntax structure in HLS to simultaneously disable multiple coding tools, for example, by disabling the syntax elements of multiple coding tools. Disabling multiple coding tools simultaneously rather than individually can save encoder running time and thereby improve coding efficiency. For example, in the case where multiple coding tools are turned off at the same time, the encoder does not need to check the usefulness of multiple coding tools, and thus can save encoder running time. Although examples are described for screen content video and camera captured video, it should be noted that in other implementations of the present disclosure, multiple encoding tools for other types of video or other criteria may be turned off simultaneously.

[0183] As described above, the high level may correspond to a video sequence, one or more pictures, pictures, slices, tile groups, tiles, CTUs, etc. According to aspects of the present disclosure, the high level may refer to a sequence level, a picture level, or a sub-picture level. The sub-picture level may refer to a slice level, a tile level, a tile group level, a CTU level, etc. In an example, a high level is a level higher than a block level (or block level). The high-level control flag may include, but is not limited to, a flag signaled in one of the following high levels or a combination of the following high levels: SPS, PPS, picture header, slice header, tile, tile group, sub-picture level, CTU level, etc.

[0184] In general, it is possible to compare the FIG. 15A to FIG. 15BThe coding tool enable flag (e.g., eipm_enable_flag) may be indicated (e.g., signaled or inferred) in any suitable syntax structure (e.g., SPS syntax table (1500)) corresponding to the sequence level (e.g., the SPS syntax table (1500)). The coding tool enable flag may indicate whether the corresponding coding tool can be used for at least one block of the appropriate level in the bitstream. Each of the coding tool enable flags of the corresponding coding tool may have different syntax at different levels. For example, for the extended intra-frame prediction mode, the syntax elements ph_eipm_enable_flag and sps_eipm_enable_flag may be used at the picture level and the sequence level, respectively. Alternatively, each of the coding tool enable flags may have the same syntax at different levels. The coding tool enable flags may include a first set of coding tool enable flags for encoding content captured by a camera and a second set of coding tool enable flags for encoding screen content.

[0185] According to various aspects of the present disclosure, a first set of encoding tool enable flags for encoding content captured by a camera may include but are not limited to the following flags: eipm_enable_flag, dmvr_enable_flag, bio_enable_flag, affine_umve_enable_flag, etmvp_enable_flag, subtmvp_enable_flag, st_chroma_enable_flag, ipf_chroma_enable_flag, ealf_enable_flag, sp_enable_flag, iip_enable_flag, etc.

[0186] In embodiments, the . Figures 15A to 15BA first set of coding tool enable flags may be indicated (e.g., signaled or inferred) in any suitable HLS (e.g., sequence level described). One of the first set of coding tool enable flags may be signaled in the HLS (e.g., SPS). Alternatively, one of the first set of coding tool enable flags may be inferred instead of being signaled. Each of the first set of coding tool enable flags for the corresponding coding tool may have different syntax at different levels. For example, for an extended intra prediction mode, syntax elements ph_eipm_enable_flag and sps_eipm_enable_flag may be used at a picture level and a sequence level, respectively. Alternatively, each of the first set of coding tool enable flags for the corresponding coding tool may have the same syntax at different levels. For example, for an extended intra prediction mode, the same syntax element eipm_enable_flag may be used at a picture level and a sequence level. In some embodiments, the first set of coding tool enable flags may be signaled as an SPS flag in the SPS, or the first set of coding tool enable flags may be signaled as a picture header flag in a picture header.

[0187] The SCC tools may include but are not limited to a combination of IBC mode, string copy mode, frequency-based intra-mode coding, implicit signaling TS mode, etc. According to aspects of the present disclosure, the second set of coding tool enable flags for encoding screen content may include but are not limited to the following flags: ibc_enable_flag, isc_enable_flag, fimc_enable_flag, ists_enable_flag, etc.

[0188] In embodiments, the . Figures 15A to 15B The second set of coding tool enable flags may be indicated (e.g., signaled or inferred) in any suitable HLS (e.g., sequence level described). One of the second set of coding tool enable flags may be signaled in the HLS (e.g., SPS). Alternatively, one of the second set of coding tool enable flags may be inferred instead of being signaled. Each of the second set of coding tool enable flags for the corresponding coding tool may have different syntax at different levels. Alternatively, each of the second set of coding tool enable flags for the corresponding coding tool may have the same syntax at different levels. In some embodiments, the second set of coding tool enable flags may be signaled in the SPS as an SPS flag, or the second set of coding tool enable flags may be signaled in a picture header as a picture header flag.

[0189] A first coding tool enable flag (e.g., sps_eipm_enable_flag) of a coding tool (e.g., an extended intra prediction mode) at a higher level (e.g., SPS) may be used to determine whether to signal a second coding tool enable flag (e.g., ph_eipm_enable_flag) of a coding tool at a lower level (e.g., picture header). The higher level is higher than the lower level. If the first coding tool enable flag is a first value (e.g., value 0), then (i) the second coding tool enable flag is not signaled and the second coding tool enable flag is inferred to be the first value, or (ii) the second coding tool enable flag is signaled as the first value. If the first coding tool enable flag is a second value (e.g., value 1), then (i) the second coding tool enable flag may be signaled as the first value or the second value, or (ii) the second coding tool enable flag may be inferred to be the first value.

[0190] In an example, a first coding tool enable flag (e.g., sps_eipm_enable_flag) in the SPS is 0, so a second coding tool enable flag (e.g., ph_eipm_enable_flag) in the picture header is inferred to be 0. If the second coding tool enable flag (e.g., ph_eipm_enable_flag) in the picture header is 0, the corresponding coding tool flag at the block level is inferred to be 0.

[0191] In an example, a first coding tool enable flag (e.g., sps_eipm_enable_flag) in the SPS is 1, and thus a second coding tool enable flag (e.g., ph_eipm_enable_flag) in the picture header is signaled as 0 or 1. If the second coding tool enable flag (e.g., ph_eipm_enable_flag) in the picture header is 1, the corresponding coding tool flag at the block level is signaled as 0 or 1.

[0192] According to aspects of the present disclosure, encoding information of multiple blocks can be decoded from an encoded video bitstream. The encoding information may include advanced control flags for the multiple blocks. The advanced control flag may indicate whether multiple encoding tools are disabled for at least one of the multiple blocks. At least one of the multiple blocks includes a current block. It can be determined based on the advanced control flag whether multiple encoding tools are disabled for at least one of the multiple blocks. Subsequently, the current block can be reconstructed without the multiple encoding tools based on the multiple encoding tools being determined to be disabled.

[0193] In an embodiment, the plurality of encoding tools include a first set of encoding tools other than SCC tools or content encoding tools captured by the camera. The advanced control flag may be a scc_only_enable_flag. A first value of the advanced control flag (e.g., true or having a value of 1) may indicate that the plurality of encoding tools (e.g., content encoding tools captured by the camera) are disabled for at least one of the plurality of encoding blocks. In an example, the first value of the advanced control flag indicates that only the SCC tool is enabled for at least one of the plurality of encoding blocks. A second value of the advanced control flag (e.g., false or having a value of 0) may indicate that the plurality of encoding tools (e.g., content encoding tools captured by the camera) are not disabled (e.g., may be enabled) for at least one of the plurality of encoding blocks.

[0194] Although the first set of encoding tools or camera-captured content encoding tools are described as examples of multiple encoding tools, the multiple encoding tools may include any suitable set of encoding tools or encoding modes, and the multiple encoding tools may be turned off simultaneously for at least one of the multiple blocks via high-level syntax elements (e.g., suitable high-level control flags according to aspects of the present disclosure). In an embodiment, the SCC tools are turned off simultaneously via high-level syntax elements. For example, if at least one of the multiple blocks includes camera-captured content and the SCC tools may be invalid for encoding the camera-captured content, the SCC tools may be disabled simultaneously for at least one of the multiple blocks, and the camera-captured content encoding tools may be used to encode at least one of the multiple blocks. Other sets of tools may be turned off simultaneously in a similar manner.

[0195] The advanced control flag may be signaled in one of the SPS, PPS, picture header, slice header, tile group level, and tile level.

[0196] In an embodiment, the advanced control flag indicates that the plurality of encoding tools are disabled for at least one of the plurality of blocks, and the plurality of encoding tools may be determined to be disabled for at least one of the plurality of blocks.

[0197] In an embodiment, the advanced control flag indicates that multiple encoding tools can be enabled for at least one of the multiple blocks. FIG. 15A to FIG. 15B For each of the coding tools described in , it can be determined whether the coding tool (e.g., extended intra prediction mode) is enabled for at least one of the plurality of blocks based on a corresponding indicator of the coding tool. The corresponding indicator can refer to a coding tool enabling flag (e.g., eipm_enable_flag) or a corresponding coding tool variable (e.g., EipmEnableFlag), such as referring to FIG. 15A to FIG. 15BThe encoding tool variables described above. A corresponding indicator for each of the plurality of encoding tools may indicate whether the encoding tool is enabled for at least one of the plurality of blocks.

[0198] In an embodiment, an advanced control flag in an SPS (or sequence header) of a video sequence may indicate whether multiple coding tools are disabled for multiple blocks in the video sequence. If the advanced control flag indicates that multiple coding tools are disabled for multiple blocks in the video sequence, it may be determined that multiple coding tools are disabled for multiple blocks in the video sequence. The advanced control flag may be signaled in an SPS of a video sequence.

[0199] A high-level control flag (e.g., scc_only_enable_flag) may be in a sequence header (or SPS), indicating that multiple coding tools are not required for the video sequence and the lower-level coding layers covered by the sequence header or SPS (e.g., the current picture in the video sequence, at least one of the multiple blocks).

[0200] In an example, when the advanced control flag (e.g., scc_only_enable_flag) is equal to 1, multiple coding tools are not required to encode the video sequence. Therefore, the coding tool enable flags or coding tool variables of multiple coding tools will be set or inferred to a value of 0 to encode the video sequence. In an example, the encoder and the decoder can set the coding tool enable flag or coding tool variable. When the advanced control flag (e.g., scc_only_enable_flag) is equal to 0, the use of multiple coding tools can be determined based on the corresponding coding tool enable flag and / or other conditions. As described above with reference to Fig.15A As described in block (1510) of , in an example, other conditions may include whether other coding tools are used. An advanced control flag (e.g., scc_only_enable_flag) may be signaled at a sequence header or SPS, for example, before any of the coding tool enable flags of the plurality of coding tools.

[0201] In the example, the advanced control flag is called scc_only_enable_flag. Fig.16 The syntax table (1600) in an SPS according to an embodiment of the present disclosure is shown, and the semantics can be described as follows. For illustration purposes, Fig.16 A first set of coding tool enable flags for the first set of coding tools is shown. As indicated by blocks (1610) to (1611), whether the first set of coding tool enable flags are signaled may depend on an advanced control flag (eg, scc_only_enable_flag).

[0202] A variable (or advanced control variable) (e.g., SccOnlyEnableFlag) may be set equal to an advanced control flag (e.g., scc_only_enable_flag). In an example, an advanced control flag (e.g., scc_only_enable_flag) equal to 1 (or true) may indicate that related syntax elements (e.g., a first set of coding tool enable flags) in an HLS (e.g., SPS) are not signaled in the bitstream. For example, scc_only_enable_flag is 1, and the advanced control variable (e.g., SccOnlyEnableFlag) is set to scc_only_enable_flag. Therefore, SccOnlyEnableFlag is 1. Then, !SccOnlyEnableFlag is 0. Therefore, related syntax elements (e.g., a first set of coding tool enable flags) between blocks (1610)-(1611) are not signaled.

[0203] An advanced control flag (eg, scc_only_enable_flag) equal to 0 may indicate that the related syntax elements (eg, the first set of coding tool enable flags) may be signaled in the bitstream. Fig.16 , when scc_only_enable_flag is 0, eipm_enable_flag, dmvr_enable_flag, bio_enable_flag, etmvp_enable_flag, subtmvp_enable_flag and iip_enable_flag are indicated by a signal.

[0204] Based on other conditions indicated by blocks (1621) to (1622), it may also be determined whether other flags of the first set of coding tool enable flags, such as affine_umve_enable_flag, st_chroma_enable_flag, ipf_chroma_enable_flag, ealf_enable_flag, and sp_enable_flag, are signaled. As described above with reference to block (1510), other conditions may include whether other coding tools may be used. For the sake of brevity, a detailed description is omitted. With reference to block (1622), if the variable SecondaryTransformEnableFlag is 1, indicating that the secondary transform may be used, then st_chroma_enable_flag is signaled. Otherwise, if the variable SecondaryTransformEnableFlag is 0, indicating that the secondary transform is disabled, then st_chroma_enable_flag is not signaled.

[0205] When the advanced control flag (eg, scc_only_enable_flag) is not present in the bitstream, the value of the advanced control flag (eg, scc_only_enable_flag) may be inferred to be 0.

[0206] In an embodiment, a corresponding indicator of each of the plurality of coding tools may be determined based on an advanced control flag. In an example, the corresponding indicator is a coding tool variable (e.g., EipmEnableFlag), and for each of the plurality of coding tools, the coding tool variable may be determined based on the advanced control flag and a corresponding coding tool enable flag of the corresponding coding tool, wherein the corresponding coding tool enable flag may be used for at least one of the plurality of blocks. For example, the coding tool variable (e.g., EipmEnableFlag) is determined based on a variable (e.g., SccOnlyEnableFlag) of a corresponding coding tool enable flag (e.g., eipm_enable_flag) and an advanced control flag (e.g., scc_only_enable_flag). In an example, the coding tool enable flag (e.g., eipm_enable_flag) is in a picture header, so the coding tool variable (e.g., EipmEnableFlag) is a picture header variable.

[0207] In an embodiment, based on the advanced control flag indicating that the multiple coding tools are disabled for at least one of the multiple blocks, the corresponding indicator of each of the multiple coding tools can be set to a value of 0, indicating that the corresponding coding tool is disabled for at least one of the multiple blocks.

[0208] In an embodiment, an advanced control flag is indicated in an SPS or a picture header, and a corresponding indicator of each of a plurality of coding tools may be used for a current picture including at least one of a plurality of blocks.

[0209] In an embodiment, an advanced control flag in a picture header of the current picture may indicate whether multiple coding tools are disabled for multiple blocks in the current picture. Disabling multiple coding tools for multiple blocks in the current picture may be determined based on the advanced control flag indicating that multiple coding tools are disabled for multiple blocks in the current picture. The advanced control flag may be signaled in a picture header of the current picture.

[0210] An advanced control flag (e.g., scc_only_enable_flag) in a picture header may be used to indicate that multiple coding tools are not required for the current picture and the lower-level coding layers covered by the picture header. An advanced control flag (e.g., scc_only_enable_flag) equal to 1 may indicate that multiple coding tools are not required to encode the current picture. Therefore, the coding tool enable flags or corresponding coding tool variables of the multiple coding tools are set to zero to encode the current picture. An advanced control flag (e.g., scc_only_enable_flag) equal to 0 may indicate that the use of multiple coding tools may be determined based on the corresponding coding tool enable flags and / or other conditions. As described above with reference to Fig.15A As described in block (1510) of , in an example, other conditions may include whether other coding tools are used. An advanced control flag (e.g., scc_only_enable_flag) may be signaled in a picture header, for example, before any of the coding tool enable flags of the plurality of coding tools.

[0211] In the example, the advanced control flag is called scc_only_enable_flag. The semantics of the advanced control flag can be described as follows.

[0212] In an example, an advanced control flag (e.g., scc_only_enable_flag) may be used to determine a coding tool enable flag (e.g., a first set of coding tool enable flags) or a coding tool variable (e.g., a coding tool variable corresponding to the first set of coding tool enable flags) of a corresponding coding tool or mode (e.g., a first set of coding tools). An advanced control flag (e.g., scc_only_enable_flag) equal to 1 may indicate that the corresponding coding tool cannot be used, for example, for at least one of a plurality of blocks in a current picture. An advanced control flag (e.g., scc_only_enable_flag) equal to 0 may indicate that the corresponding coding tool may be used, for example, for at least one of a plurality of blocks in a current picture. When the advanced control flag (e.g., scc_only_enable_flag) is not present, the value of the advanced control flag (e.g., scc_only_enable_flag) may be inferred to be 0. An advanced control variable (e.g., SccOnlyEnableFlag) may be set to the advanced control flag (e.g., scc_only_enable_flag).

[0213] When a high-level control flag (e.g., scc_only_enable_flag) in a picture header can be used to indicate whether to disable multiple coding tools for the current picture and the lower-level coding layers covered by the picture header, the above reference can be adjusted appropriately. Fig.16 Description.

[0214] In some embodiments, the plurality of encoding tools include a first set of encoding tools (or camera-captured encoding tools). As described above, encoding tool variables of camera-captured encoding tools may be derived based on advanced variables (e.g., SccOnlyEnableFlag). Encoding tool variables may be determined based on corresponding encoding tool enable flags and advanced control flags. Encoding tool variables may be determined based on corresponding encoding tool enable flags and advanced variables corresponding to advanced control flags (e.g., SccOnlyEnableFlag).

[0215] In an example, the coding tool variable EipmEnableFlag is determined based on eipm_enable_flag and an advanced variable (e.g., SccOnlyEnableFlag) as shown in Formula 1. If the advanced variable is 1 and / or eipm_enable_flag is 0, the coding tool variable EipmEnableFlag is 0. If the advanced variable is 0 and eipm_enable_flag is 1, the coding tool variable EipmEnableFlag is 1.

[0216] EipmEnableFlag=eipm_enable_flag&!SccOnlyEnableFlag Formula 1.

[0217] As shown in Equations 2 to 11, the above description can be appropriately adapted to the relationship between encoding tool variables and corresponding encoding tool enablement flags and advanced variables.

[0218] DmvrEnableFlag=dmvr_enable_flag&!SccOnlyEnableFlag (Formula 2)

[0219] BioEnableFlag=bio_enable_flag&!SccOnlyEnableFlag (Formula 3)

[0220] AffineUmveEnableFlag=affine_umve_enable_flag&!SccOnlyEnableFlag (Formula 4)

[0221] EtmvpEnableFlag=etmvp_enable_flag&! SccOnlyEnableFlag(Formula 5)

[0222] SubTmvpEnableFlag=subtmvp_enable_flag&! SccOnlyEnableFlag(Formula 6)

[0223] StChromaEnableFlag=st_chroma_enable_flag&! SccOnlyEnableFlag(Formula 7)

[0224] IpfChromaEnableFlag=ipf_chroma_enable_flag&! SccOnlyEnableFlag(Formula 8)

[0225] EalfEnaleFlag=ealf_enable__flag&! SccOnlyEnableFlag(Formula 9)

[0226] SpEnableFlag=sp_enable_flag&! SccOnlyEnableFlag(Formula 10)

[0227] IipEnableFlag=iip_enable_flag&!SccOnlyEnableFlag (Formula 11)

[0228] In an embodiment, a coding tool enable flag or a corresponding coding tool variable of a coding tool at a picture level may be derived using an advanced control flag (eg, scc_only_enable_flag) in a picture header or SPS of the current picture.

[0229] An advanced control flag (e.g., scc_only_enable_flag) equal to 1 may indicate that multiple coding tools are not required to encode the current picture. Therefore, the coding tool enable flags or corresponding coding tool variables of multiple coding tools will be set to zero to encode the current picture. An advanced control flag (e.g., scc_only_enable_flag) equal to 0 may indicate that the use of multiple coding tools may be determined based on the corresponding coding tool enable flags and / or other conditions. As described above with reference to Fig.15A As described in block (1510) of , in an example, other conditions may include whether other coding tools are used. An advanced control flag (e.g., scc_only_enable_flag) may be signaled in a picture header, for example, before any of the coding tool enable flags of the plurality of coding tools (if present).

[0230] In the example, the advanced control flag is called scc_only_enable_flag. The semantics of the advanced control flag can be described as follows.

[0231] In an example, an advanced control flag (e.g., scc_only_enable_flag) may be used to determine a coding tool enable flag (e.g., a first set of coding tool enable flags) or a coding tool variable (e.g., a coding tool variable corresponding to the first set of coding tool enable flags) of a corresponding coding tool or mode (e.g., a first set of coding tools). An advanced control flag (e.g., scc_only_enable_flag) equal to 1 may indicate that the corresponding coding tool cannot be used, for example, for at least one of a plurality of blocks in a current picture. An advanced control flag (e.g., scc_only_enable_flag) equal to 0 may indicate that the corresponding coding tool may be used, for example, for at least one of a plurality of blocks in a current picture. When the advanced control flag (e.g., scc_only_enable_flag) is not present, the value of the advanced control flag (e.g., scc_only_enable_flag) may be inferred to be 0. An advanced control variable (e.g., SccOnlyEnableFlag) may be set to the advanced control flag (e.g., scc_only_enable_flag).

[0232] When an advanced control flag (e.g., scc_only_enable_flag) in the SPS or picture header can be used to indicate whether to disable multiple coding tools for the current picture, the above reference can be adjusted appropriately. Fig.16 Description.

[0233] In an embodiment, if the advanced control variable (eg, SccOnlyEnableFlag) is equal to 1, the coding tool enable flags of the plurality of coding tools at the picture header need not be signaled. Instead, the coding tool enable flags may be inferred to be 0.

[0234] In an embodiment, one of the plurality of coding tools has a first coding tool enable flag (also referred to as a picture header enable flag) in a picture header and a second coding tool enable flag (also referred to as an SPS enable flag) in an SPS (or SPS level). A coding tool variable of the one of the plurality of coding tools at a picture level may be derived based on an existing condition (e.g., a picture header enable flag) and a value of a high-level variable (e.g., SccOnlyEnableFlag).

[0235] For example, for an extended intra prediction mode, the picture header enable flag for the extended intra prediction mode is ph_eipm_enable_flag. A picture header enable flag (e.g., ph_eipm_enable_flag) equal to 1 may indicate that the extended intra prediction mode may be used in the current picture. A picture header enable flag (e.g., ph_eipm_enable_flag) equal to 0 may indicate that the extended intra prediction mode may not be used in the current picture. When the picture header enable flag (e.g., ph_eipm_enable_flag) does not exist, it may be inferred that the value of the picture header enable flag (e.g., ph_eipm_enable_flag) is 0. A coding tool variable (e.g., EipmEnableFlag) may be set equal to ph_eipm_enable_flag &&! SccOnlyEnableFlag.

[0236] According to aspects of the present disclosure, an advanced control flag may be indicated in an SPS or a picture header. At least one of the plurality of blocks may be a current block. In an embodiment, it may be determined based on the advanced control flag whether to signal a corresponding block-level enable flag for each of the plurality of coding tools, the corresponding block-level enable flag indicating whether the coding tool is enabled for the current block.

[0237] In an embodiment, an advanced control flag may be used to indicate whether multiple encoding tools are disabled for the current block. In an example, an advanced control flag is used to indicate that multiple encoding tools are not required for block-level syntax.

[0238] The advanced control flag equal to 1 may indicate that multiple coding tools are not required to encode the current block (e.g., CB, PB, or TB). Therefore, the block-level coding tool flag indicating multiple coding tools is not signaled. The advanced control flag equal to 0 may indicate that the use of multiple coding tools may be determined based on the coding tool enable flag and / or other conditions. In an example, the advanced control flag is signaled at the picture header.

[0239] In an embodiment, the advanced control flag is scc_only_enable_flag. The semantics of the advanced control flag (e.g., scc_only_enable_flag) can be described as follows. The advanced control flag (e.g., scc_only_enable_flag) can be used to determine the signaling of the usage flag (or block-level flag) of the coding tool at the block level of the current block. In an embodiment, the usage flag or block-level flag of the coding tool indicates whether the coding tool is used for the current block. The advanced control flag (e.g., scc_only_enable_flag) equal to 1 can indicate that the coding tool cannot be used for the current block. The advanced control flag (e.g., scc_only_enable_flag) equal to 0 can indicate that the coding tool can be used for the current block. When the advanced control flag (e.g., scc_only_enable_flag) does not exist, the value of the advanced control flag (e.g., scc_only_enable_flag) can be inferred to be 0. An advanced control variable (eg, SccOnlyEnableFlag) corresponding to an advanced control flag (eg, scc_only_enable_flag) may be set equal to the advanced control flag (eg, scc_only_enable_flag).

[0240] Figures 17A to 17B FIG. 1 shows an exemplary syntax of a block-level tag according to an embodiment of the present disclosure. Fig.17A , at the block level, it is determined whether to signal a block-level flag (e.g., eipm_pu_flag) for the extended intra-frame prediction mode based on the coding tool variable for the extended intra-frame prediction mode (e.g., EipmEnableFlag) and the condition (e.g., IntraLumaPedModeIndex is greater than 1).

[0241] According to various aspects of the present disclosure, Fig. 17B Modifications shown Fig.17A In the syntax of Fig. 17BIn the example, at the block level, it is determined whether to signal a block-level flag (e.g., eipm_pu_flag) for the extended intra-frame prediction mode based on a coding tool variable (e.g., EipmEnableFlag) for the extended intra-frame prediction mode, a condition (e.g., IntraLumaPedModeIndex is greater than 1), and an advanced control variable (e.g., SccOnlyEnableFlag). In the example, the block-level flag (e.g., eipm_pu_flag) indicates whether the extended intra-frame prediction mode is used for the block. If the advanced control variable (e.g., SccOnlyEnableFlag) is 0, it is determined whether to signal a block-level flag (e.g., eipm_pu_flag) for the extended intra-frame prediction mode based on a coding tool variable (e.g., EipmEnableFlag) for the extended intra-frame prediction mode and a condition (e.g., IntraLumaPedModeIndex is greater than 1), which is the same as Fig.17A If the advanced control variable (e.g., SccOnlyEnableFlag) is 1, regardless of the coding tool variable (e.g., EipmEnableFlag) and conditions (e.g., IntraLumaPedModeIndex is greater than 1) for the extended intra prediction mode, it is determined not to signal the block-level flag (e.g., eipm_pu_flag) for the extended intra prediction mode, which is the same as Fig.17A Different from what is described in .

[0242] Fig.18 A flowchart outlining a process (1800) according to an embodiment of the present disclosure is shown. The method (1800) can be used for reconstruction of blocks such as CBs, PBs, PUs, CUs, TBs, TUs, etc. In various embodiments, the process (1800) is performed by a processing circuit, such as a processing circuit in a terminal device (310), (320), (330), and (340), a processing circuit that performs the functions of a video encoder (403), a processing circuit that performs the functions of a video decoder (410), a processing circuit that performs the functions of a video decoder (510), a processing circuit that performs the functions of a video encoder (603), etc. In some embodiments, the process (1800) is implemented in software instructions, so that when the processing circuit executes the software instructions, the processing circuit performs the process (1800). The process starts at (S1801) and proceeds to (S1810).

[0243] At (S1810), encoding information of a plurality of blocks may be decoded from an encoded video bitstream. The encoding information may indicate advanced control flags of the plurality of blocks. The advanced control flags may indicate whether a plurality of encoding tools are disabled for at least one of the plurality of blocks. At least one of the plurality of blocks may include a current block.

[0244] The advanced control flag may be signaled in one of the SPS, PPS, picture header, slice header, tile group level, and tile level.

[0245] The plurality of coding tools may include a camera-captured content coding tool different from a screen content coding (SCC) tool. The advanced control flag may be a first value that indicates that only the SCC tool is enabled for at least one of the plurality of coding blocks, and the camera-captured content coding tool is disabled for at least one of the plurality of coding blocks. The advanced control flag may be a second value that indicates that both the camera-captured content coding tool and the SCC tool are enabled for at least one of the plurality of coding blocks.

[0246] At (S1820), it may be determined whether to disable a plurality of encoding tools for at least one of the plurality of blocks based on the advanced control flag.

[0247] In an example, the advanced control flag indicates that the plurality of encoding tools are disabled for at least one of the plurality of blocks, and it may be determined that the plurality of encoding tools are disabled for at least one of the plurality of blocks.

[0248] In an example, an advanced control flag indicates that multiple coding tools are enabled for at least one of the multiple blocks, and for each of the multiple coding tools, whether the coding tool is enabled for at least one of the multiple blocks can be determined based on a corresponding indicator of the coding tool.

[0249] In an embodiment, a corresponding indicator for each of a plurality of coding tools may be determined based on an advanced control flag, the corresponding indicator indicating whether the corresponding coding tool is enabled for at least one of the plurality of blocks. In an example, the corresponding indicator is a variable. For each of the plurality of coding tools, a corresponding variable may be determined based on the advanced control flag and a corresponding enable flag of the corresponding coding tool. The corresponding enable flag may be used for at least one of the plurality of blocks. In an example, based on the advanced control flag indicating that the plurality of coding tools are disabled for at least one of the plurality of blocks, the corresponding indicator for each of the plurality of coding tools is set to a value of 0 to indicate that the corresponding coding tool is disabled for at least one of the plurality of blocks. In an example, the advanced control flag is indicated in an SPS or a picture header, and the corresponding indicator for each of the plurality of coding tools is for a current picture including at least one of the plurality of blocks.

[0250] In an embodiment, an advanced control flag is signaled in an SPS of a video sequence, and the advanced control flag indicates whether to disable multiple coding tools for multiple blocks in the video sequence. Based on the advanced control flag indicating that multiple coding tools are disabled for multiple blocks in the video sequence, it can be determined to disable multiple coding tools for multiple blocks in the video sequence.

[0251] In an embodiment, an advanced control flag is signaled in a picture header of the current picture, and the advanced control flag indicates whether multiple coding tools are disabled for multiple blocks in the current picture. Disabling multiple coding tools for multiple blocks in the current picture may be determined based on the advanced control flag indicating that multiple coding tools are disabled for multiple blocks in the current picture.

[0252] In an embodiment, an advanced control flag is indicated in an SPS or a picture header, and at least one of the plurality of blocks is a current block. In an example, it may be determined whether to signal a corresponding block-level enable flag for each of the plurality of coding tools based on the advanced control flag, the corresponding block-level enable flag indicating whether the coding tool is enabled for the current block.

[0253] At (S1830), reconstruction of the current block without the plurality of encoding tools may be determined to be disabled based on the plurality of encoding tools.

[0254] The process (1800) may be adjusted appropriately. Steps in the process (1800) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0255] The embodiments of the present disclosure may be used alone or in any order. In addition, each method (or embodiment), encoder and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer readable medium. The embodiments of the present disclosure may be applied to luminance blocks or chrominance blocks.

[0256] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Fig.19 A computer system (1900) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0257] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to mechanisms such as assembly, compilation, linking, etc. to create code comprising instructions that may be executed directly by one or more computer central processing units (CPU), graphics processing units (GPU), etc., or through interpretation, microcode execution, etc.

[0258] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0259] Fig.19 The components for the computer system (1900) shown in the example are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement relating to any one or combination of components shown in the exemplary embodiment of the computer system (1900).

[0260] Computer system 1900 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, tapping), visual input (e.g., gestures), olfactory input (not shown). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0261] The input human-machine interface device may include one or more of the following (only one of each item is drawn): keyboard (1901), mouse (1902), touch pad (1903), touch screen (1910), data gloves (not shown), joystick (1905), microphone (1906), scanner (1907), camera (1908).

[0262] The computer system (1900) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more senses of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include: tactile output devices (e.g., tactile feedback through a touch screen (1910), a data glove (not shown), or a joystick (1905), but there may also be tactile feedback devices that are not used as input devices); audio output devices (e.g., speakers (1909), headphones (not depicted)); visual output devices (e.g., screens (1910), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may be able to output two-dimensional visual output or more than three-dimensional output by means such as stereoscopic image output; virtual reality glasses (not depicted); holographic displays and cigarette cans (not depicted)); and printers (not depicted).

[0263] The computer system (1900) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1920) with CD / DVD etc. media (1921), thumb drives (1922), removable hard drives or solid-state drives (1923), traditional magnetic media such as magnetic tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD based devices such as security dongles (not depicted), and the like.

[0264] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.

[0265] The computer system (1900) may also include an interface (1954) to one or more communication networks (1955). The network may be, for example, a wireless network, a wired network, an optical network. The network may also be a local area network, a wide area network, a metropolitan area network, an in-vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include: local area networks (e.g., Ethernet, wireless LAN), cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired connections or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CAN buses, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1949) (e.g., a USB port of the computer system (1900)); other networks are typically integrated into the core of the computer system (1900) by attaching to a system bus as described below (e.g., integrated into a PC computer system via an Ethernet interface, or integrated into a smartphone computer system via a cellular network interface). The computer system (1900) may communicate with other entities using any of these networks. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., a CAN bus to certain CAN bus devices), or two-way (e.g., using a local area digital network or a wide area digital network to other computer systems). Certain protocols and protocol stacks may be used on each of these networks and network interfaces as described above.

[0266] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the core (1940) of the computer system (1900).

[0267] The core (1940) may include one or more central processing units (CPUs) (1941), graphics processing units (GPUs) (1942), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1943), hardware accelerators for certain tasks (1944), graphics adapters (1950), etc. These devices, as well as read-only memory (ROM) (1945), random access memory (1946), internal mass storage devices (e.g., internal non-user accessible hard drives, SSDs, etc.) (1947) may be connected via a system bus (1948). In some computer systems, the system bus (1948) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus (2148) directly or via a peripheral bus (1949). In an example, a screen (1910) may be connected to a graphics adapter (2150). Peripheral bus architectures include PCI, USB, etc.

[0268] The CPU (1941), GPU (1942), FPGA (1943) and accelerator (1944) can execute certain instructions, which in combination can constitute the computer code mentioned above. The computer code can be stored in ROM (1945) or RAM (1946). Transient data can also be stored in RAM (1946), while permanent data can be stored in, for example, an internal mass storage device (1947). Fast storage and retrieval of any storage device in the storage device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1941), GPUs (1942), mass storage devices (1947), ROMs (1945), RAMs (1946), etc.

[0269] The computer readable medium may have thereon computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of a type well known and available to those skilled in the art of computer software.

[0270] As an example and not limitation, a computer system (1900) having an architecture and in particular a core (1940) can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software contained in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with a user-accessible mass storage device as described above, as well as certain storage devices of the core (1940) having a non-transitory nature, such as a core mass storage device (1947) or a ROM (1945). Software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (1940). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can enable the core (1940) - and in particular the processor therein (including a CPU, GPU, FPGA, etc.) - to perform specific processing or specific parts of specific processing described herein, including defining data structures stored in RAM (1946) and modifying such data structures according to processing defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise contained in circuitry (e.g., accelerator (1944)) that may operate in place of or in conjunction with software to perform specific processing or specific portions of specific processing described herein. Where appropriate, reference to software may include logic, and vice versa. Where appropriate, reference to a computer-readable medium may include circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry implementing logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0271] Appendix A: Acronyms

[0272] JEM: Joint Development Model

[0273] VVC: Versatile Video Coding

[0274] BMS: Benchmark Set

[0275] MV: Motion Vector

[0276] HEVC: High Efficiency Video Coding

[0277] SEI: Supplemental Enhancement Information

[0278] VUI: Video Availability Information

[0279] GOP: Group of Pictures

[0280] TU: Transform Unit

[0281] PU: prediction unit

[0282] CTU: Coding Tree Unit

[0283] CTB: Coding Tree Block

[0284] PB: prediction block

[0285] HRD: Hypothesized Reference Decoder

[0286] SNR: Signal to Noise Ratio

[0287] CPU: Central Processing Unit

[0288] GPU: Graphics Processing Unit

[0289] CRT: cathode ray tube

[0290] LCD: Liquid Crystal Display

[0291] OLED: Organic Light Emitting Diode

[0292] CD: Compact Disc

[0293] DVD: Digital Video Disc

[0294] ROM: Read Only Memory

[0295] RAM: Random Access Memory

[0296] ASIC: Application-Specific Integrated Circuit

[0297] PLD: Programmable Logic Device

[0298] LAN: Local Area Network

[0299] GSM: Global System for Mobile Communications

[0300] LTE: Long Term Evolution

[0301] CANBus: Controller Area Network Bus

[0302] USB: Universal Serial Bus

[0303] PCI: Peripheral Component Interconnect

[0304] FPGA: Field Programmable Gate Array

[0305] SSD: Solid State Drive

[0306] IC: Integrated Circuit

[0307] CU: Coding Unit

[0308] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. It will therefore be appreciated that, although not explicitly shown or described herein, those skilled in the art will be able to conceive of many systems and methods that implement the principles of the present disclosure and are therefore within its spirit and scope.

Claims

1. A method for video decoding, characterized in that: The method comprises: decoding encoding information for a plurality of blocks from an encoded video bitstream, the encoding information comprising advanced control flags for the plurality of blocks, the advanced control flags indicating whether a plurality of encoding tools are disabled for at least one block of the plurality of blocks, the at least one block of the plurality of blocks including the current block; determining, based on the advanced control flag, whether to disable the plurality of encoding tools for at least one block of the plurality of blocks; In response to determining that the advanced control flag indicates that the plurality of encoding tools are enabled, decoding an encoding tool enable flag from the encoding information, the encoding tool enable flag indicating whether a corresponding different one of the plurality of encoding tools is enabled; determining whether each of the plurality of encoding tools is disabled for at least one block of the plurality of blocks based on the encoding tool enable flag; In response to determining that the advanced control flag indicates that the plurality of encoding tools are disabled, determining that all of the plurality of encoding tools are disabled without relying on the encoding tool enable flag; Based on the plurality of coding tools being determined to be disabled, the current block is reconstructed without the plurality of coding tools.

2. The method according to claim 1, characterized in that The advanced control flag indicates that the plurality of encoding tools are disabled for at least one block among the plurality of blocks, and it is determined that the plurality of encoding tools are disabled for at least one block among the plurality of blocks.

3. The method according to claim 1, characterized in that The advanced control flag indicates that the multiple coding tools are enabled for at least one block among the multiple blocks, and for each of the multiple coding tools, based on a corresponding indicator of the coding tool, it is determined whether the coding tool is enabled for at least one block among the multiple blocks.

4. The method according to any one of claims 1 to 3, characterized in that The plurality of encoding tools includes a camera captured content encoding tool different from a screen content encoding (SCC) tool, The advanced control flag indicates, as a first value, that only the SCC tool is enabled for at least one of the plurality of blocks and that a content encoding tool captured by the camera is disabled for at least one of the plurality of blocks, and The advanced control flag indicates, as a second value, that the camera-captured content encoding tool and the SCC tool are enabled for at least one block of the plurality of blocks.

5. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Based on the advanced control flag, a corresponding indicator for each of the plurality of encoding tools is determined, the corresponding indicator indicating whether the corresponding encoding tool is enabled for at least one block of the plurality of blocks.

6. The method according to claim 5, characterized in that The corresponding indicator is a variable, and Determining the respective indicator further comprises determining, for each encoding tool of the plurality of encoding tools, a respective variable based on the advanced control flag and a respective enablement flag of the respective encoding tool, the respective enablement flag being used for at least one block of the plurality of blocks.

7. The method according to claim 5, characterized in that Based on the advanced control flag indicating that the plurality of encoding tools are disabled for at least one of the plurality of blocks, the corresponding indicator of each of the plurality of encoding tools is set to a value of 0 to indicate that the corresponding encoding tool is disabled for at least one of the plurality of blocks.

8. The method according to claim 5, characterized in that The advanced control flag is indicated in a sequence parameter set (SPS) or a picture header, and The respective indicator of each of the plurality of encoding tools is for a current picture including at least one block of the plurality of blocks.

9. The method according to any one of claims 1 to 3, characterized in that The advanced control flag is signaled in a sequence parameter set (SPS) of a video sequence, and the advanced control flag indicates whether the multiple coding tools are disabled for the multiple blocks in the video sequence. Based on the advanced control flag indicating that the multiple coding tools are disabled for the multiple blocks in the video sequence, it is determined to disable the multiple coding tools for the multiple blocks in the video sequence.

10. The method according to any one of claims 1 to 3, characterized in that The advanced control flag is indicated by a signal in a picture header of the current picture, and the advanced control flag indicates whether the multiple coding tools are disabled for the multiple blocks in the current picture. Based on the indication by the advanced control flag that the multiple coding tools are disabled for the multiple blocks in the current picture, it is determined to disable the multiple coding tools for the multiple blocks in the current picture.

11. The method according to any one of claims 1 to 3, characterized in that The advanced control flag is indicated in a sequence parameter set (SPS) or a picture header, and At least one block among the plurality of blocks is the current block.

12. The method according to claim 11, characterized in that Also includes: Based on the advanced control flag, it is determined whether to signal a corresponding block-level enable flag for each of the plurality of encoding tools, the corresponding block-level enable flag indicating whether the encoding tool is enabled for the current block.

13. The method according to any one of claims 1 to 3, characterized in that The advanced control flag is signaled in one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, a tile group level, and a tile level.

14. A video encoding method, characterized in that: The method comprises: determining coding information for a plurality of blocks, the coding information comprising advanced control flags for the plurality of blocks, the advanced control flags indicating whether a plurality of coding tools are disabled for at least one block of the plurality of blocks, the at least one block of the plurality of blocks comprising the current block; determining whether at least one of the plurality of blocks has a plurality of encoding tools disabled; In response to determining that the advanced control flag indicates that the plurality of encoding tools are enabled, determining an encoding tool enable flag indicating whether a corresponding different one of the plurality of encoding tools is enabled; determining whether each of the plurality of encoding tools is disabled for at least one block of the plurality of blocks based on the encoding tool enable flag; In response to determining that the advanced control flag indicates that the plurality of encoding tools are disabled, determining that all of the plurality of encoding tools are disabled without relying on the encoding tool enable flag; Based on the plurality of coding tools being determined to be disabled, the current block is reconstructed without the plurality of coding tools.

15. A computer device, characterized in that: The computer device comprises: one or more computer-readable non-transitory storage media configured to store computer program code; and One or more computer processors configured to access the computer program code and execute the method according to any one of claims 1 to 14 as instructed by the computer program code.

16. A device for video decoding, comprising: Processing circuitry configured to perform the method according to any one of claims 1 to 14.

17. A non-transitory computer readable medium storing instructions, characterized in that: When the instructions are executed by a computer for video decoding, the computer is caused to perform the method according to any one of claims 1 to 13.

18. A computer storage medium, characterized in that: Instructions are stored, and the instructions can be executed by at least one processor to execute the method of claim 14, generate a code stream, and store it.

Citation Information

Patent Citations

  • Method and apparatus for video coding

    US20210021841A1