Different ranges of clipping processes for video and image compression

JP2026530272APending Publication Date: 2026-09-08TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025547815
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-04
Filing Date
2023-11-09
Publication Date
2026-09-08

Smart Images

  • Figure 2026530272000001_ABST
    Figure 2026530272000001_ABST
Patent Text Reader

Abstract

The various implementations described herein include methods and systems for coding video. In one embodiment, the method includes: receiving a video bitstream containing encoded video data; obtaining information indicating two or more different clipping ranges of the encoded video data; deriving two or more different clipping ranges based on the obtained information; performing a first clipping operation on a first portion of the encoded video data based on a first clipping range among the two or more different clipping ranges; and performing a second clipping operation on a second portion of the encoded video data based on a second clipping range different from the first clipping range.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 532,661, filed on 14 August 2023, entitled “Different Scopes of Clipping Processes for Video and Image Compression,” and is a continuation of and claims priority to U.S. Patent Application No. 18 / 502,002, filed on 4 November 2023, entitled “Different Scopes of Clipping Processes for Video and Image Compression.”

[0002] The disclosed embodiments relate to video coding, including, but not limited to, systems and methods that generally provide a different range of clipping processes relating to video and image compression. [Background technology]

[0003] Digital video is supported by a variety of electronic devices, including digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconference devices, and video streaming devices. These electronic devices transmit, receive, or communicate digital video data over communication networks and / or store digital video data in storage devices. Because of the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding may be used to compress video data according to one or more video coding standards before it is transmitted or stored. Video coding may be performed by hardware and / or software on electronic / client devices or servers providing cloud services.

[0004] Video coding generally utilizes prediction methods (e.g., interpretation, intrapretation) that take advantage of the inherent redundancy in video data. The goal of video coding is to compress video data into a format that uses a lower bitrate while avoiding or minimizing a decrease in video quality. Several video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Multipurpose Video Coding (VVC / H.266) is a video compression standard intended as a successor to HEVC. The ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. A valid version 1.0.0, including Errata 1 of this specification, was released on January 8, 2019. [Overview of the project] [Problems that the invention aims to solve]

[0005] This disclosure describes a method for using different ranges for clipping operations in different parts of video and image processing, for example, to improve compression efficiency. For example, a codec performs conversions between a visual media file and a bitstream of visual media data. The use of different ranges can improve the coding process by reducing noise from the clipped portion of the decoded video data and / or areas where useful data is not encoded, which may be introduced by the codec itself (for example, as quantization noise) during the quantization step (for example, such codec noise may affect prediction accuracy). Furthermore, the coding process and techniques can operate with 2 bytes (or 16 bits) in filtering or other precision modification processes within the codec, scaling up the input signal in several intermediate processing stages (for example, using an expanded range over the entire original unmodified range), even if the codec has a bit depth smaller than 16 bits. By using different ranges for clipping operations, useful data in the expanded range can be captured during intermediate processing (stages) to improve the accuracy of the final output. [Means for solving the problem]

[0006] According to several embodiments, a method for video decoding is provided. The method includes: (i) receiving a video bitstream (e.g., a coded video sequence) containing encoded video data; (ii) obtaining information indicating two or more different clipping ranges for the encoded video data; (iii) deriving two or more different clipping ranges based on the obtained information; (iv) performing a first clipping operation on a first portion of the encoded video data based on a first clipping range among the two or more different clipping ranges; and (v) performing a second clipping operation on a second portion of the encoded video data based on a second clipping range different from the first clipping range.

[0007] According to several embodiments, a video encoding method is provided. The method includes (i) receiving a current image; (ii) determining two or more clipping ranges from the current image; and (iii) encoding information indicating the two or more clipping ranges into a video bitstream containing encoded video data.

[0008] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic device. The computing system includes a control circuit and a memory for storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes encoder components and decoder components (e.g., a transcoder).

[0009] According to some embodiments, a non-temporary computer-readable storage medium is provided. The non-temporary computer-readable storage medium stores one or more instruction sets for execution by a computing system. One or more instruction sets include instructions for executing any of the methods described herein.

[0010] Accordingly, devices and systems are disclosed along with methods for encoding and decoding video. Such methods, devices, and systems may complement or replace conventional methods, devices, and systems for video encoding / decoding. The features and advantages described herein are not necessarily exhaustive, and in particular, several additional features and advantages will be apparent to those skilled in the art in consideration of the drawings, specification, and claims provided herein. Furthermore, it should be noted that the language used herein has been chosen primarily for readability and explanatory purposes and is not necessarily chosen to describe or limit the subject matter described herein.

[0011] To enable a more detailed understanding of this disclosure, a more detailed description can be provided by referring to the features of various embodiments, some of which are shown in the accompanying drawings. However, as a person skilled in the art will understand from reading this disclosure, the description may be open to other valid features, and therefore the accompanying drawings are merely illustrative of the relevant features of this disclosure and should not be considered necessarily limiting. [Brief explanation of the drawing]

[0012] [Figure 1] This block diagram shows an exemplary communication system in several embodiments. [Figure 2A] This is a diagram showing exemplary elements of encoder components according to several embodiments. [Figure 2B] A directory diagram showing exemplary elements of decoder components according to several embodiments. [Figure 3]It is a block diagram illustrating an exemplary server system according to some embodiments. [Figure 4A] It is a flow diagram illustrating an exemplary method for encoding video according to some embodiments. [Figure 4B] It is a flow diagram illustrating an exemplary method for decoding video according to some embodiments. DETAILED DESCRIPTION OF EMBODIMENTS

[0013] By convention, the various features illustrated in the drawings are not necessarily drawn to scale, and like reference numerals may be used throughout the specification and drawings to indicate like features.

[0014] The present disclosure describes methods that use different ranges for clipping operations in different portions of video and image data, for example to improve compression efficiency. For example, video data may be clipped to remove noise, such as quantization noise, and / or to improve decoding efficiency.

[0015] Exemplary Systems and Devices FIG. 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic device 120-1 to electronic device 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system for use in video-enabled applications, such as video conferencing applications, digital TV applications, and media storage and / or distribution applications, for example.

[0016] Source device 102 includes a video source 104 (e.g., a camera component or media storage) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to generate an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from the video source 104 may have a larger data size compared to the encoded video bitstream 108 generated by the encoder component 106. Because the encoded video bitstream 108 has a smaller data size (smaller data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth for transmission and less storage space for storage compared to the video stream from the video source 104. In some embodiments, source device 102 does not include an encoder component 106 (e.g., configured to transmit uncompressed video data to a network 110(or more)).

[0017] One or more networks 110 represent any number of networks that transmit information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wireline and / or wireless communication networks. One or more networks 110 may exchange data over circuit-switched channels and / or packet-switched channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0018] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (the streaming server is configured to store and / or deliver video content, such as an encoded video stream from a source device 102). The server system 112 includes a coder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the coder component 114 includes an encoder component and / or a decoder component. In various embodiments, the coder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the coder component 114 is configured to decode an encoded video bitstream 108 and re-encode the video data using different encoding standards and / or methods to produce encoded video data 116. In some embodiments, the server system 112 is configured to produce multiple video formats and / or encodings from the encoded video bitstream 108. In some embodiments, the server system 112 functions as a media-enabled network element (MANE). For example, the server system 112 may be configured to prune the encoded video bitstream 108 to adapt a potentially different bitstream to one or more of the electronic devices 120. In some embodiments, a MANE is provided separately from the server system 112.

[0019] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode encoded video data 116 to produce an output video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., one that is communicably coupled to an external display device and / or includes media storage). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access a server system 112 to retrieve the encoded video data 116.

[0020] The source device and / or multiple electronic devices 120 may also be referred to as “terminal devices” or “user devices.” In some embodiments, one or more of the source device 102 and / or electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smartphone, tablet, or laptop), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0021] In an exemplary operation of the communication system 100, source device 102 transmits an encoded video bitstream 108 to server system 112. For example, source device 102 may encode a stream of pictures captured by the source device. Server system 112 receives the encoded video bitstream 108 and may decode and / or encode the encoded video bitstream 108 using coder components 114. For example, server system 112 may apply a more optimal encoding to the video data for network transmission and / or storage. Server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display a video picture.

[0022] Figure 2A is a block diagram showing exemplary elements of an encoder component 106 according to several embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source which is a component of a device different from the encoder component 106). The video source 104 can provide the source video sequence as a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCB or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores pre-captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a series of separate pictures that convey motion when viewed in sequence. The picture itself may be organized as a spatial array of pixels, and each pixel may contain one or more samples depending on the sampling structure, color space, etc., used. Those skilled in the art will readily understand the relationship between pixels and samples. The following description focuses on samples.

[0023] The encoder component 106 is configured to code and / or compress pictures from a source video sequence in real time or under other time constraints required by the application to obtain a coded video sequence 216. One function of the controller 204 is to implement an appropriate coding speed. In some embodiments, the controller 204 controls and is functionally coupled to other functional units, which are described below. Parameters set by the controller 204 may include rate control-related parameters (e.g., lambda values ​​for picture skipping, quantizer, and / or rate distortion optimization techniques), picture size, picture group (GOP) layout, and maximum motion vector search range. Those skilled in the art will readily identify other functions of the controller 204 that may be related to the encoder component 106 optimized for a particular system design.

[0024] In some embodiments, the encoder component 106 is configured to operate in a coding loop. In a simplified example, the coding loop includes a source coder 202 (which is responsible for generating symbols, such as a symbol stream, based on, for example, a coded input picture and a reference picture(s)) and a (local) decoder 210. The decoder 210 reconstructs the symbols to generate sample data in a similar manner to a (remote) decoder (if the compression between the symbols and the coded video bitstream is reversible). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since decoding the symbol stream yields bit-exact results regardless of the decoder's location (local or remote), the contents of the reference picture memory 208 are also bit-exact between the local encoder and the remote encoder. In this way, the encoder's prediction unit interprets the same sample values ​​as the reference picture samples that the decoder interprets when using predictions during decoding. This principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained due to, for example, channel errors) is known to those skilled in the art.

[0025] The operation of decoder 210 may be the same as that of a remote decoder, such as decoder component 122, which will be described in detail below in relation to Figure 2B. However, referring briefly to Figure 2B, since symbols are available and the encoding / decoding of symbols to the video sequence coded by the entropy coder 214 and parser 254 may be reversible, the entropy decoding portion of decoder component 122, including buffer memory 252 and parser 254, does not need to be fully implemented in local decoder 210.

[0026] The decoder techniques described herein, with the exception of parsing / entropy decoding, may exist in substantially the same functional form in the corresponding encoders. Therefore, the subject matter of this disclosure focuses on decoder operation. Descriptions of encoder techniques may be omitted, as they may be the reverse of decoder techniques.

[0027] As part of its operation, the source coder 202 can perform motion-compensated predictive coding, which predictively codes the input frame by referencing one or more previously coded frames from the video sequence, designated as reference frames. In this manner, the coding engine 212 codes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frames(s) that may be selected as predictive references(s) to the input frame. The controller 204 may manage the coding operation of the source coder 202, including, for example, setting parameters and subgroup parameters used to encode the video data.

[0028] The decoder 210 decodes the coded video data of a frame that may be designated as a reference frame, based on symbols created by the source coder 202. The operation of the coding engine 212 may preferably be lossy. When the coded video data is decoded by a video decoder (not shown in Figure 2A), the reconstructed video sequence may be a replica of the source video sequence with some errors. The decoder 210 can reproduce the decoding process that may be performed by the remote video decoder on the reference frame and store the reconstructed reference frame in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frame with common content, which will be acquired (without transmission errors) by the remote video decoder.

[0029] The predictor 206 can perform a predictive search on the coding engine 212. That is, for a new frame to be coded, the predictor 206 may search the reference picture memory 208 for specific metadata such as reference picture motion vectors, block shapes, etc., which can serve as sample data (as candidate reference pixel blocks) or as appropriate predictive references for the new picture. The predictor 206 may operate pixel block by pixel on a sample block basis to find appropriate predictive references. As determined by the search results obtained by the predictor 206, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory 208.

[0030] The outputs of all the aforementioned functional units can undergo entropy coding in the entropy coder 214. The entropy coder 214 converts the symbols generated by the various functional units into coded video sequences by reversibly compressing the symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable-length coding, and / or arithmetic coding).

[0031] In some embodiments, the output of the entropy coder 214 is coupled to a transmitter. The transmitter may be configured to buffer the coded video sequence(s) generated by the entropy coder 214 and prepare it for transmission over a communication channel 218, which may be a hardware / software link to a storage device that stores the coded video data. The transmitter may be configured to merge the coded video data from the source coder 202 with other data to be transmitted, such as coded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter may transmit additional data along with the coded video. The source coder 202 may include such data as part of the coded video sequence. The additional data may include time / space / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, and the like.

[0032] The controller 204 can manage the operation of the encoder component 106. During coding, the controller 204 may assign a specific coded picture type to each coded picture, which may affect the coding technique applied to each picture. For example, a picture may be assigned as an intra-picture (I-picture), a predictive picture (P-picture), or a bidirectional predictive picture (B-picture). An intra-picture can be coded and decoded without using any other frames in the sequence as a source for prediction. Some video codecs enable various types of intra-pictures, including, for example, independent decoder refresh (IDR) pictures. Those skilled in the art will recognize their variations of I-pictures and their respective uses and characteristics, so they will not be repeated here. A predictive picture can be coded and decoded using intra-prediction or inter-prediction, which uses up to one motion vector and reference index to predict the sample value of each block. A bidirectional predictive picture can be coded and decoded using intra-prediction or inter-prediction, which uses up to two motion vectors and reference indexes to predict the sample value of each block. Similarly, multiple prediction pictures can use three or more reference pictures and associated metadata to reconstruct a single block.

[0033] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each), and each block may be coded. Blocks may be coded predictively by referencing other (already coded) blocks, as determined by the coding assignment applied to each picture in the block. For example, blocks of picture I may be coded unpredictably, or they may be coded predictively by referencing already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of picture P may be coded unpredictably via spatial prediction or via temporal prediction by referencing one previously coded reference picture. Pixel blocks of picture B may be coded unpredictably via spatial prediction or via temporal prediction by referencing one or two previously coded reference pictures.

[0034] Video can be captured as multiple source pictures (video pictures) in a time series. Intra-picture prediction (often abbreviated as intra-prediction) utilizes spatial correlations within a given picture, while inter-picture prediction utilizes (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still-buffered reference picture in the video, then the block in the current picture may be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0035] The encoder component 106 can perform coding operations in accordance with a predetermined video coding technique or standard, such as any of those described herein. In these operations, the encoder component 106 can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to the syntax specified by the video coding technique or standard being used.

[0036] Figure 2B is a block diagram showing exemplary elements of a decoder component 122 according to several embodiments. The decoder component 122 in Figure 2B is coupled to channel 218 and display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to a loop filter 256 and configured to transmit data to the display 124 (for example, via a wired or wireless connection).

[0037] In some embodiments, the decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more coded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each coded video sequence is independent of other coded video sequences. Each coded video sequence may be received from channel 218, which may be a hardware / software link to a storage device that stores coded video data. The receiver receives coded video data along with other data, e.g., coded audio data and / or auxiliary data streams, and this data may be forwarded to their respective usage entities (not shown). The receiver may isolate coded video sequences from other data. In some embodiments, the receiver receives additional (redundant) data along with the coded video. The additional data may be included as part of a coded video sequence(s). The additional data may be used by the decoder component 122 to decode the data and / or to more accurately reconstruct the original video data. The additional data can take the form of, for example, a time layer, a spatial layer, or an SNR enhancement layer, redundant slices, redundant pictures, or forward error correction codes.

[0038] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes called an entropy decoder), a scaler / inverse unit 258, an intra-picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, the decoder component 122 is implemented at least partially in software.

[0039] Buffer memory 252 is coupled between channel 218 and parser 254 (for example, to counteract network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 within decoder component 122 (configured, for example, to handle playout timing), a separate buffer memory is provided outside decoder component 122 (for example, to counteract network jitter). When receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isosynchronous network, buffer memory 252 may be unnecessary or small. Buffer memory 252 may be required for use in best-effort packet networks such as the Internet, and buffer memory 252 may be relatively large, or advantageously adaptively sized, and may be at least partially implemented in an operating system or similar element (not shown) outside decoder component 122.

[0040] The parser 254 is configured to reconstruct symbols 270 from a coded video sequence. The symbols may include, for example, information used to manage the operation of decoder component 122 and / or information for controlling rendering devices such as the display 124. The control information for rendering devices may take the form of Supplementary Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not shown). The parser 254 parses (entropy decodes) the coded video sequence. The coding of the coded video sequence may follow video coding techniques or standards and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, context-dependent or non-context-dependent arithmetic coding, etc. The parser 254 may extract from the coded video sequence a set of subgroup parameters relating to at least one of the subgroups of pixels in the video decoder, based on at least one parameter corresponding to a group. Subgroups can include picture groups (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transformation units (TU), and prediction units (PU). Parser 254 can also extract information such as transformation coefficients, quantization parameter values, and motion vectors from coded video sequences.

[0041] The reconstruction of symbol 270 may involve multiple different units, depending on the type of coded video picture or part thereof (interpicture and intrapicture, interblock and intrablock, etc.) and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by parser 254. The flow of such subgroup control information between parser 254 and the following multiple units is not described for illustrative purposes.

[0042] The decoder component 122 may be conceptually subdivided into several functional units, and in some implementations, these units can interact closely with each other and integrate with each other at least partially. However, for clarity, the conceptual subdivision of functional units is maintained herein.

[0043] The scaler / inverse unit 258 receives quantized transformation coefficients and control information as symbols 270 (e.g., which transformation to use, block size, quantization coefficients, and / or quantization scaling matrix) from the parser 254. The scaler / inverse unit 258 can output a block containing sample values ​​that can be input to the aggregator 268.

[0044] In some cases, the output samples of the scaler / inverse unit 258 relate to intracoded blocks, i.e., blocks that do not use predictive information from previously reconstructed pictures but can use predictive information from previously reconstructed portions of the current picture. Such predictive information may be provided by the intrapicture predictive unit 262. The intrapicture predictive unit 262 may generate blocks of the same size and shape as the block being reconstructed by using already reconstructed surrounding information fetched from the current (partially reconstructed) picture from the current picture memory 264. The aggregator 268 may, sample by sample, add the predictive information generated by the intrapicture predictive unit 262 to the output sample information provided by the scaler / inverse unit 258.

[0045] In other cases, the output samples of the scaler / inverse unit 258 are associated with a block, which has been intercoded and potentially motion-compensated. In such cases, the motion-compensated prediction unit 260 can access the reference picture memory 266 to fetch samples to be used for prediction. After motion-compensating the fetched samples according to the symbols 270 associated with the block, these samples can be added to the output of the scaler / inverse unit 258 by the aggregator 268 to generate output sample information (in this case, called residual samples or residual signals). The address in the reference picture memory 266 from which the motion-compensated prediction unit 260 fetches the predicted samples may be controlled by a motion vector. The motion vector may be available to the motion-compensated prediction unit 260 in the form of a symbol 270 which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory 266 when the exact motion vector of the subsample is used, a motion vector prediction mechanism, etc.

[0046] The output samples of the aggregator 268 can be subjected to various loop filtering techniques in the loop filter unit 256. The video compression technique may include in-loop filtering techniques controlled by parameters contained in the coded video bitstream and made available to the loop filter unit 256 as symbols 270 from the parser 254, but may also respond to metadata obtained during decoding of earlier parts (in decoding order) of the coded picture or coded video sequence, or to previously reconstructed and loop-filtered sample values. The output of the loop filter unit 256 may be a sample stream that can be output to a rendering device such as the display 124, or it may be stored in the reference picture memory 266 for use in future interpicture prediction.

[0047] A specific coded picture, once reconstructed, may be used as a reference picture for future predictions. Once a coded picture is fully reconstructed and identified as a reference picture (for example, by parser 254), the current reference picture can become part of reference picture memory 266, and any unused current picture memory may be reallocated before initiating the reconstruction of subsequent coded pictures.

[0048] The decoder component 122 can perform decoding operations according to a predetermined video compression technique that may be documented in a standard such as one of the standards described herein. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that it is faithful to the syntax of the video compression technique or standard as specified in the video compression technique documentation or standard, specifically the profile documentation therein. Also, in order to conform to some video compression technique or standard, the complexity of the coded video sequence may be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limitations set by the level may, in some cases, be further limited by the virtual reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0049] Figure 3 is a block diagram showing a server system 112 according to several embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes one or more field-programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).

[0050] The network interface(s) 304 may be configured to interface with one or more communication networks (e.g., wireless, wired, and / or optical networks). These communication networks may be local, wide-area, metropolitan, automotive, and industrial, real-time, or latency-tolerant. Examples of communication networks include local area networks such as Ethernet, cellular networks including Wi-Fi, GSM, 3G, 4G, 5G, and LTE, wired or wireless wide-area digital networks for television including cable television, satellite television, and terrestrial television, and automotive and industrial networks including CANBus. Such communications may be unidirectional, receive-only (e.g., broadcast television), transmit-only (e.g., CANbus to a specific CANbus device), or bidirectional (e.g., with other computer systems using local or wide-area digital networks). Such communications may include communications with one or more cloud computing networks.

[0051] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 may include one or more of the following: a keyboard, mouse, trackpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. The output device 308 may include one or more of the following: an audio output device (e.g., a speaker), a visual output device (e.g., a display or monitor), etc.

[0052] Memory 314 may include high-speed random-access memory (such as DRAM, SRAM, DDR RAM, and / or other random-access solid-state memory devices) and / or non-volatile memory (such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Memory 314 optionally includes one or more storage devices located remotely from the control circuit 302. Memory 314, or the non-volatile solid-state memory device(s) within Memory 314, includes a non-temporary computer-readable storage medium. In some embodiments, Memory 314, or the non-temporary computer-readable storage medium of Memory 314, stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: ● Operating system 316 that handles various basic system services and includes procedures for performing hardware-dependent tasks. ● A network communication module 318 used to connect the server system 112 to other computing devices via one or more network interfaces 304 (for example, via wired and / or wireless connections). ● A coding module 320 for performing various functions related to encoding and / or decoding data such as video data. In some embodiments, the coding module 320 is an instance of the coder component 114. The coding module 320 includes, but is not limited to, one or more of the following: Regarding the decoder component 122, the decoding module 322 performs various functions related to decoding the encoded data as described above. ○ Encoding module 340 for performing various functions related to encoded data as described above, with respect to encoder component 106 ● For example, a picture memory 352 for storing pictures and picture data, for use with a coding module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0053] In some embodiments, the decoding module 322 includes a parsing module 324 (for example, configured to perform the various functions described above with respect to the parser 254), a transformation module 326 (for example, configured to perform the various functions described above with respect to the scaler / inverse transformation unit 258), a prediction module 328 (for example, configured to perform the various functions described above with respect to the motion compensation prediction unit 260 and / or intrapicture prediction unit 262), and a filter module 330 (for example, configured to perform the various functions described above with respect to the loop filter 256).

[0054] In some embodiments, the encoding module 340 includes a code module 342 (for example, configured to perform the various functions described above with respect to the source coder 202 and / or coding engine 212) and a prediction module 344 (for example, configured to perform the various functions described above with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include a subset of the modules shown in Figure 3. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.

[0055] Each of the identified modules stored in memory 314 corresponds to an instruction set for performing the functions described herein. The modules identified above (e.g., instruction sets) do not need to be implemented as separate software programs, procedures, or modules; therefore, various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the coding module 320 may optionally not include separate decoding and encoding modules, but rather use the same set of modules to perform both sets of functions. In some embodiments, memory 314 stores a subset of the modules and data structures identified above. In some embodiments, memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0056] Figure 3 shows server system 112 according to several embodiments, but Figure 3 is not a schematic diagram of the structure of the embodiments described herein, but is intended to be a functional description of various features that may be present in one or more server systems. In practice, as will also be recognized by those skilled in the art, items shown separately can be combined, and some items can be separated. For example, some items shown separately in Figure 3 can be implemented on a single server, and a single item can be implemented by one or more servers. The actual number of servers used to implement server system 112, and how functions are allocated among them, will vary from implementation to implementation and will optionally depend in part on the amount of data traffic the server system will handle during peak and average usage periods.

[0057] Examples of coding processes and techniques The coding processes and techniques described below can be performed on the devices and systems described above (e.g., source device 102, server system 112, and / or electronic device 120). A hybrid video codec may include two or more of the following: intra-prediction, inter-prediction, transform coding, quantization, entropy coding, and post-in-loop filtering. The coding processes and techniques described herein enable the use of different ranges in the clipping process of different parts of the codec. In some embodiments, the use of different ranges improves the compression efficiency of the video codec.

[0058] The clipping range (x, y) of the clipping process determines the minimum and maximum values ​​of the sample after clipping. In some embodiments, the clipping range (x, y) includes upper and lower limits, or minimum and maximum values ​​x and y. For example, if the internal bit depth is set to 8 bits and each sample of the target signal is within the range of (0, 255), the clipping process for each sample z may be set as shown in Equation 1 below.

number

[0059] Similarly, if the internal bit depth is set to 10 bits, the clipping process for each sample z in the entire coding / decoding process may be as shown in Equation 2 below.

number

[0060] Generally, this is written as shown in Equations 3 and 4 below.

number

[0061] Instead of using the same clipping range for all clipping operations in different parts of a codec, the coding processes and techniques described herein enable the use of different clipping ranges (x, y) (sometimes referred to as “clipping ranges”) for different parts of a codec. In some embodiments, a codec is associated with different sets of clipping ranges, and one particular clipping range is selected from the set of clipping ranges depending on the part of the codec in which the clipping operation is performed. For example, a first clipping range is used for intra-prediction, a second clipping range (e.g., another clipping range) different from the first clipping range is used for inter-prediction, and a third clipping range (e.g., yet another clipping range) different from the first and second clipping ranges is used for in-loop filtering. In some embodiments, which clipping range is selected depends on the presence and / or use of one or more coding tools, such as one or more signal reshaping coding tools.

[0062] The coding processes and techniques described herein may be applied to perform conversions between visual media files and bitstreams of visual media data, and the conversions may be performed according to the codec. For example, bit-depth based clipping can be applied at multiple locations to avoid data overflow. The range of bit-depth based clipping may be defined as [0, 2^bit depth-1]. Clipping operations can be applied at filtering, sampling, interpolation, weighted prediction, weighted joining, reconstruction, and / or other stages to ensure that the generated prediction and reconstruction samples remain within a defined dynamic range.

[0063] For example, in the case of video captured with 10-bit precision, useful information may be collected in the range (0, 1023). For example, data less than 0 or greater than 1023 should be clipped by (0, 1023). As another example, if the data is encoded within the range (16, 235) for video captured in 8 bits (e.g., under BT.2020), then as the signal increases from 8 bits to 10 bits, the signal range can quadruple and reach (64, 940). Therefore, in such a 10-bit signal (e.g., a signal ranging from 0 to 1023), useful information in the video may be captured (e.g., nearly captured or only captured) in the range of 64 to 940 (e.g., there is no useful data collected from 0 to 63, or from 941 to 1023). An offset may be signaled in the bitstream to set the clipping range as (64, 940), or the clipping range of (64, 940) may be set as the default, with an offset to the range of (64, 940) being signaled in the bitstream. For example, a signaled offset of 7 further shifts the clipping range from (64, 940) to (57, 933). In some embodiments, two offsets are signaled in the bitstream to set the clipping range to shift asymmetrically at the upper and lower limits. The use of adaptively determined clipping ranges can improve the coding process by reducing noise from the clipped portion (e.g., boundary portion) of the decoded video data, and / or regions where useful data is not encoded, which may be introduced by the codec itself during the quantization step (e.g., as quantization noise) (e.g., such codec noise may affect prediction accuracy).

[0064] As used herein, without loss of generality, “clipping range (singular)” or “clipping range (plural)” refers to a clipping range relating to any color component (Y, Cb, or Cr), such as one or more of the signal color components (Y, Cb, Cr, R, G, or B), each or all, or any combination thereof.

[0065] Figure 4A is a flowchart illustrating a method 400 for encoding video according to several embodiments. Method 400 can be performed in a computing system (e.g., a server system 112, a source device 102, or an electronic device 120) having a control circuit and a memory that stores instructions for execution by the control circuit. In some embodiments, method 400 is performed by executing instructions stored in the memory of the computing system (e.g., memory 314). The system receives the current picture (402) and determines two or more clipping ranges from the current picture (e.g., to clip portions of the current picture) (404). The system encodes the information indicating the two or more clipping ranges into a video bitstream containing encoded video data (406).

[0066] Figure 4B is a flowchart illustrating a method 450 for decoding video according to several embodiments. The method 450 may be performed on a computing system (e.g., a server system 112, a source device 102, or an electronic device 120) having a control circuit and a memory for storing instructions for execution by the control circuit. In some embodiments, the method 450 is performed by executing instructions stored in the memory of the computing system (e.g., memory 314). The system receives a video bitstream containing encoded video data (452). The system obtains information indicating two or more different clipping ranges for the encoded video data (454). The system derives two or more different clipping ranges based on the obtained information (456). The system performs a first clipping operation on a first portion of the encoded video data based on a first clipping range among the two or more different clipping ranges (458), and performs a second clipping operation on a second portion of the encoded video data based on a second clipping range different from the first clipping range (460).

[0067] In some embodiments, multiple clipping ranges (e.g., two or more clipping ranges) are specified and used for clipping operations within a processing pipeline. In some embodiments, multiple clipping ranges are defined in the codec and used for clipping operations within a processing pipeline depending on the processing stage. For example, one clipping range (e.g., a first clipping range) is specified and used for intra-prediction, and another clipping range (e.g., a second clipping range) is specified and used for all other processing stages. Alternatively or additionally, one clipping range (e.g., a first clipping range) is specified and used for intra-prediction, another clipping range (e.g., a second clipping range) is specified and used for inter-prediction, and yet another clipping range (e.g., a third clipping range) is specified and used for all other processing stages.

[0068] Alternatively or additionally, one clipping range (e.g., a fourth clipping range distinct from one or more of the first, second, or third clipping ranges) is specified and used for a loop filter, which includes, but is not limited to, deblocking, sample adaptive offset (SAO), adaptive loop fill (ALF), and cross-component adaptive loop filter (CC-ALF). In some embodiments, the loop filter includes (e.g., related to) multiple clipping ranges, including, but not limited to, deblocking, SAO, ALF, and CC-ALF. For example, these clipping ranges are specified to be used for a particular filter (or multiple filters) within the loop filter.

[0069] In some embodiments, the designation of multiple clipping ranges and / or usage for a particular processing stage is signaled in high-level syntax, including, but not limited to, sequence-level flags, picture-level flags, sub-picture-level flags, slice-level flags, and / or tile-level flags. In some embodiments, HLS elements are signaled at a level higher than the block level. For example, HLS elements may correspond to sequence level, frame level, slice level, or tile level. As another example, HLS elements may be signaled in the video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptive parameter set (APS), slice header, picture header, tile header, and / or CTU header.

[0070] In some embodiments, multiple clipping ranges are signaled in high-level syntax, and each clipping range is associated with an index value. For example, for each predefined processing stage (e.g., intra-prediction, inter-prediction, or loop filtering), the index of the applied clipping range is signaled.

[0071] In some embodiments, multiple clipping ranges are defined in the codec and used for clipping within the processing pipeline, depending on other coding tools used / may be used in the coding / decoding process.

[0072] In some embodiments, when certain tools are enabled at a high level, such as the sequence level, picture level, or slice level, one clipping range (e.g., a first clipping range) is specified and used; when these tools are disabled, another clipping range (e.g., a second clipping range different from the first clipping range) is specified and used. For example, if any coding tool modifies the signal dynamic range during processing, one clipping range (e.g., a third clipping range) is used before the coding tool is applied, and another clipping range (e.g., a fourth clipping range) is used after the coding tool is applied. More than three coding ranges may be used in the processing. An example of such a coding tool is Luma Mapping with Chroma Scaling (LMCS).

[0073] LMCS has two main components: 1) a process for mapping input lumacode values ​​to a new set of code values ​​for use within the code loop, and 2) a luma-dependent process for scaling chroma residual values. The first process, luma mapping, aims to improve the coding efficiency of standard and high dynamic range video signals by making more effective use of the range of lumacode values ​​allowed at a given bit depth. The second process, chroma scaling, manages the relative compression efficiency of the luma and chroma components of the video signal. The luma mapping process of LMCS may be applied at the pixel sample level and may be carried out using a piecewise linear model. The chroma scaling process may be applied at the chroma block level and may be implemented using a scaling factor derived from reconfigured neighboring luma samples of the chroma block.

[0074] In some embodiments, when certain tools are enabled at a low level, such as the block unit level, coding unit level, prediction unit level, or transformation unit level, one clipping range (e.g., a first clipping range) is specified and used, and when these tools are disabled, another clipping range (e.g., a second clipping range) is specified and used. For example, if any coding tool modifies the signal dynamic range during processing, one clipping range (e.g., a third clipping range) is used before the application of such tool(s), and another clipping range (e.g., a fourth clipping range) is used afterward. Three or more coding ranges may be used throughout the entire processing.

[0075] In some embodiments, one or more indicators, such as syntax elements (e.g., offsets, or minimum and maximum values ​​of the clipping range), are signaled within the bitstream of visual media data and used to derive each clipping range for one or more clipping operations. The clipping ranges in these embodiments are determined based on the characteristics of the bitstream. For example, a video bitstream may be received that includes encoded video data and one or more syntax elements indicating one or more clipping ranges. In this example, one or more clipping ranges are derived based on one or more syntax elements, and one or more clipping operations are performed on the encoded video data using the derived clipping range(s).

[0076] For example, if the signaled syntax element includes an offset of a first clipping range (e.g., the original clipping range or the modified clipping range), the derived range is different from the first clipping range. In some embodiments, the signaled syntax element includes the minimum and maximum values ​​of the clipping range (e.g., the range includes the minimum and maximum values). In some embodiments, the derived clipping range is smaller than the first clipping range (e.g., one or more regions of the first clipping range are truncated). In some embodiments, the derived clipping range is larger than the first clipping range (e.g., beyond the range of 0 to 1023). For example, coding processes and techniques can operate with 2 bytes (or 16 bits) in filtering or other precision modification processes within the codec that scale up the input signal (e.g., using an expanded range over the entire original, unmodified range) at some intermediate processing stages, even if the codec has a bit depth smaller than 16 bits.

[0077] In some embodiments, a single clipping range (x, y) is signaled in the bitstream at a high level (e.g., a sequence, picture, slice, or tile parameter set, a syntax table including one or more offsets as headers) for the entire sequence, picture, slice, or tile. For example, one or more additional syntax elements are used to represent a range in the bitstream. In some embodiments, multiple clipping ranges (x, y) are signaled in the bitstream and used for their respective clipping operations. In some embodiments, clipping ranges are signaled with high-level syntax elements.

[0078] The bitstream includes encoded video data having one or more grouping types, such as sequences, pictures, slices, or tiles. In some embodiments, a picture includes several blocks and a clipping range (x i , y i ) are signaled separately in the bitstream for each block i. In some embodiments, the signaling is performed at a low level, such as a coding unit, a prediction unit, or a transformation unit, or other smaller units, and one or more additional syntax elements are used to represent one or more clipping ranges in the bitstream.

[0079] In some embodiments, the clipping range (x, y) is determined based on information related to the input (source) signal and / or compression (codec) processing, and the derived clipping range is used for one or more clipping processes. For example, the input (source) signal corresponds to the original picture and / or pre-filtered picture from video source 104. The clipping range can be derived from the pre-filtered picture (e.g., minimum and maximum values ​​related to the clipping range). For example, motion-compensated time filtering (MCTF) is a pre-processing technique that may be employed before video coding to improve compression efficiency. In such a case, the minimum and maximum values ​​related to the clipping range can be derived based on an MCTF pre-filtered picture. For example, an encoder performs pre-filtering or receives a pre-filtered picture, and the encoder signals the minimum and maximum values ​​related to the pre-filtered picture. Alternatively or additionally, the encoder signals a flag indicating that the encoded picture is a pre-filtered picture or was derived from a pre-filtered picture, and the decoder determines minimum and maximum values ​​from the pre-filtered picture to derive the clipping range for the clipping operation.

[0080] In some embodiments, the clipping range (x, y) is determined based on the dynamic range of the input signal. For example, if the dynamic range of the input signal is an unrestricted (full) range from 0 to (1 ≪ BitDepth) - 1, the clipping range is set to equal (0, (1 ≪ BitDepth) - 1). In such scenarios, one or more syntax elements can be signaled within the bitstream to identify the source signal as having an unlimited range.

[0081] In some embodiments, the dynamic range of the input signal is a limited range from A to B, where A and B are predetermined constants, and the clipping range is set to equal (A, B). In such scenarios, one or more syntax elements can be signaled in the bitstream to derive the constants A and B. For example, the dynamic range of the input signal may be a limited range specified by one of the signal processing standards, e.g., (64, 940) for a 10-bit signal, and the clipping range is set to equal (64, 940). In some embodiments, the clipping range includes two endpoints indicating the minimum and maximum values ​​of the clipping range (e.g., 64 and 940). In such scenarios, one or more syntax elements can be signaled in the bitstream to specify the use of predetermined constants.

[0082] In some embodiments, the clipping range (x, y) is determined based on the bit depth of the input signal. For example, with an input signal of n bits and a codec of m bits, if m ≥ n, the clipping range is set to equal (0, ((1≪n)-1)≪(mn)). Thus, if the bit depth n of the input signal is 8 bits and the bit depth m of the codec is 10 bits, the clipping range may be set to equal (0, 1020). In such scenarios, one or more syntax elements may be signaled in the bitstream to derive the bit depth n of the input signal. For example, the bit depth n of the input signal may be directly signaled and / or derived based on the delta from the bit depth m of the codec.

[0083] In some embodiments, the clipping range (x, y) is determined based on an internal (codec) signal dynamic range. For example, when an internal (codec) signal is converted from one range domain to another range domain, the clipping range is set equal to the final range domain. In such cases, the clipping range is determined based on information from a bitstream corresponding to internal signal conversion.

[0084] In some embodiments, the internal (codec) signal comprises a representation of an original dynamic range (A, B) that is changed to a dynamic range of (C, D) during compression processing. In such a scenario, the clipping range may be set equal to (C, D).

[0085] In some embodiments, the internal (codec) signal comprises a representation of a more restricted original dynamic range, and within the compression process, the original dynamic range is extended to a full (e.g., extended or expanded) dynamic range of 0~(1<<BitDepth)-1. In such a scenario, the clipping range may be set equal to (0, (1<<BitDepth)-1).

[0086] In some embodiments, for an input signal having a bit depth of n bits and a codec having a bit depth of m bits, when m < n, the clipping range is set equal to (0, ((1<<m)-1)). For example, when the bit depth of the input signal is 10 bits and the bit depth of the codec is 8 bits, the clipping range may be set equal to (0, 255). In such a scenario, one or more syntax elements may be signaled in the bitstream to derive the bit depth of the input signal.

[0087] In some embodiments, clipping is performed on a Luma-Mapping Chroma-Scaling (LMCS) mapped domain. In some embodiments, the upper and lower minimum and maximum values ​​of the clipping range are derived by applying a forward LMCS lookup table (LUT) to the signaled minimum and maximum values. For example, if LMCS is enabled and clipping is performed on an LMCS-mapped domain, the minimum and maximum values ​​may be derived by applying a forward LMCS LUT to the signaled minimum and maximum values ​​such that the signaled values ​​are mapped to new values ​​(e.g., derived values) in the LMCS-mapped domain using the LMCS LUT.

[0088] As described above, multiple clipping ranges (e.g., two or more clipping ranges) can be used for clipping operations within a processing pipeline. If two or more different clipping ranges are used for each clipping operation, the first clipping operation may be performed on a first portion of the encoded video data based on a first clipping range signaled in the bitstream, and the second clipping operation may be performed on a second portion of the encoded video data based on a second clipping range different from the first clipping range. As described above, each of the clipping ranges may be signaled or derived. For example, the first clipping range may be signaled, and the second clipping range may be derived based on the signaled first clipping range (e.g., by applying a forward LMCS LUT based on signaled maximum and minimum values). In this way, instead of signaling two pairs of minimum and maximum values, only one pair is signaled (for example, the signaled pair of values ​​is associated with the maximum and minimum values ​​of the first clipping range), and the second pair is derived from the first pair. In this example, after LMCS, the signaled pair of values ​​is applied to the first clipping range (e.g., outside the LMCS-mapped domain), and the derived pair of values ​​(e.g., modified values ​​from the signaled values) is applied to the second clipping range (e.g., inside the LMCS-mapped domain).

[0089] In some embodiments, the adaptive clipping range may be determined based on the picture group (GOP) structure, particularly in the case of a low-latency B (LDB) configuration. The picture group, or GOP structure, specifies the order in which frames are arranged within and between frames. In a low-latency configuration, the first frame is an intra-frame, and the others are encoded as generalized P or B-pictures. A low-latency B (LDB) configuration includes a first frame that is an intra-frame, and the rest are encoded as B-pictures. In some embodiments, one or more parameters of the adaptive clipping range are set based on the GOP structure.

[0090] In some embodiments, one or more clipping ranges (or clipping range parameters of an adaptive clipping range) are identified based on whether a low-latency B (LDB) configuration is active. For example, an adaptive clipping range, or a clipping range different from that used in a random-access configuration, is used in an LDB configuration. In this way, a first clipping range can be used when the GOP structure is in a random-access configuration, and a second clipping range (e.g., different from the first clipping range) can be used when the GOP structure is in an LDB configuration. In some embodiments, the clipping ranges and / or clipping range parameters(s) are derived or identified after sorting has been performed to obtain a GOP pyramidal structure (e.g., a hierarchical coding structure).

[0091] In some embodiments, an offset or delta value is signaled in relation to (for example, in association with) the clipping range.

[0092] In some embodiments, one or more clipping ranges (x i , y i ) is the reference clipping range value

number

[0093] In some embodiments, the clipping range (x, y) is designed (configured) to allow overwriting at a specific level (e.g., can be overwritten). For example, a sequence-level range (x 0 , y 0 ) is initially signaled, but the signaled range may be overwritten by signaling a new range (x', y') at a lower level (e.g., picture level, slice level, or tile level), and overwrites (e.g., determines) the range used at that level (e.g., for the picture, if the new range (x', y') is signaled at the picture level). In some embodiments, the overwritten range (x', y') is directly signaled. In some embodiments, the overwritten range (x', y') and the original range (x 0 , y 0The offset between ) is signaled. For example, in a concert video, a sequence-level picture may have certain characteristics for the next n pictures in the sequence (for example, the next 10 pictures may be very dark). As a result, the next n pictures may have a reduced clipping range because the video data is likely to be of lower quality. Thus, the ability to override makes it possible to adaptively adjust the clipping range temporally and / or locally. For example, overrides are performed at a lower level than the original signaled unit, at the picture level when different settings are signaled at the sequence level, and at the block level when different settings are signaled at the picture level.

[0094] In some embodiments, for example as described above, the basic clipping range value (x) is determined based on information related to the input (source) signal and / or compression (codec) processing. 0 , y 0 ) is determined, and multiple signaled offset values ​​(Δx i Δy i ) are signaled separately in the bitstream at lower levels (e.g., per block). For example, each block may have one range having different blocks of different ranges. In some embodiments, a lookup table stores a plurality of predetermined offset values, and a signaled indicator (e.g., a signaled index, or a plurality of signaled indices) provides information on which of the predetermined offset values ​​should be used. Signaling may also occur at lower levels, such as coding units, prediction units, transformation units, or other lower levels, and one or more additional syntax elements may indicate ranges in the bitstream (e.g., basic clipping range (x)). 0 , y 0 It is used to represent )).

[0095] In some embodiments, as described above, multiple base clipping range values ​​are used per block based on information related to the input (source) signal and / or compression (codec) processing.

number

number

[0096] In some embodiments, the offset value (Δx i Δy i The derivation of ) depends on the value of the quantization parameter or the quantization step size.

[0097] In some embodiments, different clipping ranges are used for different components (e.g., color components). For example, a first range (x L , y L ) is used for Luma, and the second range (x c , y c ) is used for chroma. Alternatively, the first range (x L , y L ) is used for Luma, and the second range (xcb , y cb ) is used for the first chromatic Cb component, and the third range (x cr , y cr ) is used for the second chroma Cr component.

[0098] In some embodiments, one or more clipping ranges are derived using coded information, such as the reconstructed sample value ranges of one or more prediction blocks (e.g., prediction blocks to which single prediction is applied, or prediction blocks to which dual prediction is applied).

[0099] Figures 4A and 4B show several logical stages in a specific order, but the stages that are not order-dependent may be rearranged, and other stages may be combined or separated. Several rearrangements or other groupings not specifically mentioned will be obvious to those skilled in the art, and therefore the rearrangements and groupings presented herein are not exhaustive. Furthermore, it should be recognized that the stages may be implemented in hardware, firmware, software, or any combination thereof.

[0100] Here, we refer to some exemplary embodiments.

[0101] (A1) In one embodiment, several embodiments include a video decoding method (e.g., Method 450). This method includes (i) receiving a video bitstream containing encoded video data; (ii) obtaining information indicating two or more different clipping ranges for the encoded video data (e.g., information defined in a codec rather than the video bitstream, or information signaled and obtained from the video bitstream); (iii) deriving two or more different clipping ranges based on the obtained information; (iv) performing a first clipping operation on a first portion of the encoded video data (e.g., at the sequence level, picture level, slice level, tile level, or coding unit level, prediction unit level, or transformation unit level) based on a first clipping range among the two or more different clipping ranges; and (v) performing a second clipping operation on a second portion of the encoded video data based on a second clipping range different from the first clipping range. For example, multiple clipping ranges are specified and used for clipping operations in a processing pipeline. In some embodiments, the value of a first clipping range is signaled, and the value of a second clipping range is derived by the decoder (for example, based on the value of the first clipping range). For example, the value of the second clipping range may be derived by performing a lookup operation using the value of the first clipping range.

[0102] (A2) In some embodiments of A1, a first clipping operation with respect to a first portion of the encoded video data is performed based on a first clipping range in a first processing stage before applying a coding tool to the encoded video data, and a second clipping operation with respect to a second portion of the encoded video data is performed based on a second clipping range in a second processing stage after applying a coding tool to the encoded video data. For example, if some tools are enabled at a high level, such as sequence level, picture level, or slice level, one clipping range is specified and used, and if these tools are disabled, another clipping range is specified and used. In some embodiments, if any of the coding tools change the signal dynamic range during processing, one clipping range is used before applying such tool(s), and then another clipping range is used. In some embodiments, three or more coding ranges are used throughout the entire processing. For example, if some tools are enabled at a low level, such as block / coding units or predictive / transform units, one clipping range is specified and used, and if these tools are disabled, another clipping range is specified and used. In some embodiments, if any of the coding tools alter the signal dynamic range during processing, one clipping range is used before the application of such tool(s), and another clipping range is used afterward. In some embodiments, three or more coding ranges are used throughout the entire processing.

[0103] (A3) In some embodiments of A2, the coding tool includes a Luma Mapping Chromascaling (LMCS) tool. In some embodiments, one of two or more different clipping ranges is derived using an LMCS lookup table. For example, the LMCS lookup table is applied to the values ​​of a first clipping range to derive the values ​​of a second clipping range.

[0104] (A4) In some embodiments of A1 to A3, two or more different clipping ranges of the encoded video data should be used for different processing stages (e.g., intra, inter, filtering, use of enabled / disabled tools). For example, a first clipping range may be used for a portion of the encoded video data outside the LMCS-mapped domain, and a second clipping range may be used for a portion of the encoded video data within the LMCS-mapped domain. In some embodiments, multiple clipping ranges are defined in the codec and used for clipping processes in the processing pipeline depending on the processing stage. For example, one clipping range is specified and used for intra prediction, and another clipping range is specified and used for all other processing stages. For example, one clipping range is specified and used for intra prediction, another clipping range is specified and used for inter prediction, and another clipping range is specified and used for all other processing stages. For example, one clipping range is specified and used for loop filtering, including deblocking, sample adaptive offset (SAO), and / or adaptive loop filtering (ALF). In some embodiments, the clipping range is specified to be used for a particular filter (or multiple filters) within a loop filter, and the loop filter includes, but is not limited to, deblocking, sample-adaptive offset (SAO), and adaptive loop filter (ALF), and includes multiple clipping ranges. In some embodiments, two or more different clipping ranges will be used based on the GOP structure of the encoded video data.

[0105] (A5) In some embodiments of A1 to A4, the acquired information includes information from one or more syntax elements signaled in the video bitstream. In some embodiments, the designation of multiple clipping ranges and / or usage for a particular processing stage is signaled in high-level syntax, including, but not limited to, sequence-level flags, picture-level flags, sub-picture-level flags, slice-level flags, and / or tile-level flags. In some embodiments, one or more syntax elements include high-level syntax elements.

[0106] (A6) In some embodiments of A5, information from one or more syntax elements signaled in the video bitstream indicates two or more different clipping ranges for different processing stages. In one example, each clipping range is associated with an index value. For example, for each predefined processing stage (e.g., intra-prediction, inter-prediction, loop filtering), an index of the applied clipping range may be signaled.

[0107] (A7) In some embodiments of A5 or A6, information from one or more syntax elements signaled in the video bitstream indicates the use of each clipping range of two or more different clipping ranges for a corresponding processing stage. For example, a specification of multiple clipping ranges and / or uses for a particular processing stage is signaled in high-level syntax.

[0108] (A8) In some embodiments of A5 to A7, one or more syntax elements indicate the respective clipping range indices of two or more different clipping ranges in high-level syntax, and each of the two or more different clipping ranges is associated with an index value of its respective clipping range index. For example, multiple clipping ranges are signaled in HLS, and each clipping range is associated with an index value.

[0109] (A9) In some embodiments of A8, the index value of each clipping range index is signaled for a corresponding predetermined processing stage. For example, for each predefined processing stage (e.g., intra prediction, inter prediction, loop filtering), the index of the applied clipping range is signaled.

[0110] (A10) In some embodiments of A1 to A9, encoded video data is encoded according to a codec, and information relating to two or more different clipping ranges is defined in the codec, and the information includes indicators that relate two or more different clipping ranges to corresponding processing stages of different processing stages. In some embodiments, multiple clipping ranges are defined in the codec and used for clipping processing in the processing pipeline depending on the processing stage, for example, the processing stage includes one or more of intra-prediction, inter-prediction, loop filtering such as deblocking, sample adaptive offset (SAO), and adaptive loop filter (ALF).

[0111] (A11) In some embodiments of A10, the indicator specifies the use of a first clipping range for the intra-predictive processing stage, a second clipping range for the inter-predictive processing stage, and the use of one or more clipping ranges for the loop filtering processing stage. In one example, one clipping range is specified and used for intra-predictive processing, and another clipping range is specified and used for all other processing stages. In some embodiments, one clipping range is specified and used for intra-predictive processing, another clipping range is specified and used for inter-predictive processing, and another clipping range is specified and used for all other processing stages. In some embodiments, one clipping range is specified and used for loop filtering, including but not limited to deblocking, SAO, and ALF. In some embodiments, the loop filter includes multiple clipping ranges. These clipping ranges may be specified to be used for specific filters (one or more) within the loop filter.

[0112] (A12) In some embodiments of A10 or A11, an indicator that associates two or more different clipping ranges with corresponding processing stages further describes the status of the coding tool associated with each processing stage, the coding tool modifying the signal dynamic range of each portion of the encoded video data during each processing stage. In some embodiments, one clipping range is used before the application of any such tool(s) that is modifying the signal dynamic range during processing, and then another clipping range is used afterward, according to the determination that any of the coding tools is modifying the signal dynamic range during processing. In some embodiments, three or more coding ranges are used throughout the entire processing.

[0113] (A13) In some embodiments of A12, the indicator associates a first clipping range with a first portion of encoded video data (e.g., at a high level such as sequence level, picture level, slice level, or a low level such as block unit level, coding unit level, prediction unit level, or transformation unit level) based on the status of the coding tool corresponding to the enabled coding tool, and the indicator associates a third clipping range, different from the first clipping range, with a database-listed first portion of encoded video based on the status of the coding tool corresponding to the enabled coding tool. In some embodiments, multiple clipping ranges are defined in the codec and used for clipping in the processing pipeline depending on other coding tools used / may be used in the coding / decoding process. In some embodiments, if several tools are enabled at a high level such as sequence level, picture level, slice level, etc., one clipping range is specified and used, and if these tools are disabled, another clipping range is specified and used. In some embodiments, if any of the coding tools modify the signal dynamic range during processing, one clipping range is used before the application of such tool(s), and then another clipping range is used. In some embodiments, three or more coding ranges can be used for the entire process. In some embodiments, one clipping range is specified and used when some tools are enabled at a low level, and another clipping range is specified and used when these tools are disabled. For example, one of the coding tools may be LMCS.

[0114] (B1) In other embodiments, some embodiments include a method for video coding (e.g., Method 400). This method includes the steps of: receiving a current picture (e.g., an original or pre-filtered photograph); determining two or more clipping ranges from the current picture (e.g., based on the dynamic range of signals in the current image and / or quantization noise associated with the current image); and encoding information associated with indicating the two or more clipping ranges (e.g., signaled high-level syntax or indicators, sequence / picture / subpicture / slice / tile flags, where each clipping range is associated with a signaled index value) into a video bitstream containing encoded video data. In some embodiments, a first clipping range is signaled in the video bitstream, and a second clipping range is derived during video decoding based on the first clipping range.

[0115] (B2) In some embodiments of B1, the two or more clipping ranges include a first clipping range for clipping a first portion of the current image in a first processing step, and a second clipping range, different from the first clipping range, for clipping a second portion of the current image in a second processing step different from the first processing step.

[0116] (B3) In some embodiments of B1 or B2, the encoded information relating to indicating two or more clipping ranges includes information relating to a first portion of the current picture that is clipped by a first clipping range in a first processing step, and a second portion of the current picture that is clipped by a second clipping range in a second processing step (e.g., intra-prediction, inter-prediction, loop filtering).

[0117] (C1) In other embodiments, some embodiments include a method for performing a conversion between a visual media file (e.g., from a video source 104) and a bitstream of visual media data (e.g., an encoded video sequence). The method includes the steps of obtaining a visual media file and performing a conversion between the visual media file and a bitstream of visual media data, wherein the bitstream includes information indicating two or more clipping ranges of the visual media data.

[0118] (C2) In some embodiments of C1, two or more clipping ranges include a first clipping range for clipping a first portion of the visual media data in a first processing step, and a second clipping range, different from the first clipping range, for clipping a second portion of the visual media data in a second processing step different from the first processing step.

[0119] (C3) In some embodiments of C1 or C2, encoded information indicating two or more clipping ranges includes information about a first portion of visual media data that is clipped by a first clipping range in a first processing step, and a second portion of visual media data that is clipped by a second clipping range in a second processing step.

[0120] In other embodiments, some embodiments include a computing system (e.g., server system 112) which includes a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, the memory storing one or more instruction sets configured to be executed by the control circuit, the one or more instruction sets including instructions for executing any of the methods described herein (e.g., A1-A13, B1-B3, and C1-C3 above).

[0121] In yet another embodiment, some embodiments include a non-temporary computer-readable storage medium that stores one or more instruction sets for execution by a control circuit of a computing system, the one or more instruction sets including instructions for executing any of the methods described herein (e.g., A1-A13, B1-B3, and C1-C3 above).

[0122] Terms such as “first,” “second,” etc., may be used herein to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the scope of the claims. As used in the descriptions of embodiments and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural form unless the context clearly indicates otherwise. The terms “and / or” as used herein will be understood to refer to and encompass one or any possible combination of the enumerated items relating to the subject. The terms “comprises” and / or “comprising,” as used herein, identify the presence of a described feature, integer, step, operation, element, and / or component, but it will be further understood that they do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0123] As used herein, the term “if” can be interpreted, depending on the context, to mean “when” or “upon” or “in response to determining” or “in accordance with a determination” or “in response to detecting” the stated premise. Similarly, the phrases “if it is determined that [the stated premise is true]” or “when [the stated premise is true]” or “when [the stated premise is true]” can be interpreted, depending on the context, to mean “upon determining” or “in response to determining” or “in accordance with a determination” or “upon detecting” or “in response to detecting” the stated premise.

[0124] The above description is provided with reference to specific embodiments for illustrative purposes. However, the above exemplary description is not intended to be exhaustive or to limit the claims to the exact form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been selected and described to best illustrate the operating principle and practical applications, and thereby to be available to those skilled in the art. [Explanation of Symbols]

[0125] 100 Communication Systems 102 Source Device 104 Video Sources 106 Encoder Components 108 video bitstreams 110 Network 112 Server Systems 114 Coder Components 116 Video Data 120 Electronic Devices 122 Decoder Components 124 displays 202 Source Coder 204 Controller 206 Predictors 208 Reference Picture Memory 210 Decoder 212 Coding Engine 214 Entropy Coder 216 video sequences 218 communication channels 252 buffer memory 254 Parsa 256 Loop Filter Unit 258 Scaler / Inverse Unit 260 Motion Compensation Prediction Units 262 IntraPicture Prediction Units 264 Current Picture Memory 266 Reference Picture Memory 268 Aggregators 270 symbols 302 Control Circuit 304 Network Interface 306 User Interface 308 Output Devices 310 Input Devices 312 Communications Bus 314 memory 316 Operating Systems 318 Network Communication Module 320 coding modules 322 Decoding Module 324 Parsing Module 326 Conversion Module 328 Prediction Modules 330 Filter Module 340 Encoding Modules 342 Code Modules 344 Prediction Modules 352 Picture Memory

Claims

1. A video decoding method performed on a computing system having memory and one or more processors, The steps include receiving a video bitstream containing encoded video data, The steps include obtaining information indicating two or more different clipping ranges of the encoded video data, The steps include: deriving two or more different clipping ranges based on the acquired information; The steps include performing a first clipping operation on a first portion of the encoded video data based on a first clipping range among the two or more different clipping ranges, A step of performing a second clipping operation on a second portion of the encoded video data based on a second clipping range different from the first clipping range, Methods that include...

2. The method according to claim 1, wherein the first clipping operation with respect to the first portion of the encoded video data is performed based on the first clipping range in a first processing step before applying a coding tool to the encoded video data, and the second clipping operation with respect to the second portion of the encoded video data is performed based on the second clipping range in a second processing step after applying the coding tool to the encoded video data.

3. The method according to claim 2, wherein the coding tool includes a Luma Mapping Chroma Scaling (LMCS) tool.

4. The method according to claim 1, wherein the two or more different clipping ranges of the encoded video data are used in different processing stages.

5. The method according to claim 1, wherein the acquired information includes information from one or more syntax elements signaled within the video bitstream.

6. The method according to claim 5, wherein the information from the one or more syntax elements signaled within the video bitstream indicates the two or more different clipping ranges for different processing stages.

7. The method according to claim 5, wherein the information from the one or more syntax elements signaled within the video bitstream indicates the use of each of the two or more different clipping ranges for a corresponding processing step.

8. The method according to claim 5, wherein one or more syntax elements represent the respective clipping range indices of the two or more different clipping ranges using high-level syntax, and each of the two or more different clipping ranges is associated with the index value of the respective clipping range index.

9. The method according to claim 8, wherein the index value of each of the clipping range indices is signaled for a corresponding predetermined processing step.

10. The method according to claim 1, wherein the encoded video data is encoded according to a codec, the information relating to the two or more different clipping ranges is defined in the codec, and the information includes an indicator that associates the two or more different clipping ranges with corresponding processing stages of a plurality of different processing stages.

11. The method according to claim 10, wherein the indicator specifies the use of the first clipping range for an intra-predictive processing stage, the use of the second clipping range for an inter-predictive processing stage, and the use of one or more clipping ranges for a loop filtering processing stage.

12. The method according to claim 10, wherein the indicators relating the two or more different clipping ranges to the corresponding processing stages further describe the status of the coding tool associated with each processing stage, and the coding tool modifies the signal dynamic range of each portion of the encoded video data during each processing stage.

13. The method according to claim 12, wherein the indicator associates a first clipping range with a first portion of the encoded video data based on the status of the coding tool corresponding to the enabled coding tool, and the indicator associates a third clipping range different from the first clipping range with a first portion of the video data based on the status of the coding tool corresponding to the enabled coding tool.

14. A computing system, Control circuit and Memory and One or more instruction sets stored in the memory and configured for execution by the control circuit, The set comprises one or more instruction sets, Receive a video bitstream containing encoded video data, Information is obtained indicating two or more different clipping ranges of the encoded video data. Based on the information obtained, two or more different clipping ranges are derived. A first clipping operation is performed on a first portion of the encoded video data based on a first clipping range among the two or more different clipping ranges. A second clipping operation is performed on a second portion of the encoded video data based on a second clipping range different from the first clipping range. Includes instructions for, Computing system.

15. The computing system according to claim 14, wherein the first clipping operation with respect to the first portion of the encoded video data is performed based on the first clipping range in a first processing step before applying a coding tool to the encoded video data, and the second clipping operation with respect to the second portion of the encoded video data is performed based on the second clipping range in a second processing step after applying the coding tool to the encoded video data.

16. The computing system according to claim 15, wherein the coding tool includes a lumina mapping chromascaling (LMCS) tool.

17. The computing system according to claim 14, wherein the acquired information includes information from one or more syntax elements signaled within the video bitstream.

18. A non-temporary computer-readable storage medium storing one or more instruction sets configured to be executed by a computing device having a control circuit and memory, wherein the one or more instruction sets are Receive a video bitstream containing encoded video data, Information is obtained indicating two or more different clipping ranges of the encoded video data. Based on the information obtained, two or more different clipping ranges are derived. A first clipping operation is performed on a first portion of the encoded video data based on a first clipping range among the two or more different clipping ranges. A second clipping operation is performed on a second portion of the encoded video data based on a second clipping range different from the first clipping range. A non-temporary computer-readable storage medium containing instructions for use.

19. The non-temporary computer-readable storage medium according to claim 18, wherein the first clipping operation with respect to the first portion of the encoded video data is performed based on the first clipping range in a first processing stage before applying a coding tool to the encoded video data, and the second clipping operation with respect to the second portion of the encoded video data is performed based on the second clipping range in a second processing stage after applying the coding tool to the encoded video data.

20. The coding tool comprises a Luma Mapping Chromascaling (LMCS) tool, wherein the non-temporary computer-readable storage medium is as described in claim 19.