Adaptive range for clipping processing in video and image compression

By adaptively selecting the crop range in video encoding and decoding technology, the problem of insufficient adaptive range of cropping processing in the prior art is solved, the noise problem in the encoding process is improved, and the video quality and encoding and decoding efficiency are improved.

CN120226359APending Publication Date: 2025-06-27TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380080318.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-04
Filing Date
2023-11-09
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the existing video encoding and decoding technology, the adaptive range of the cropping process is insufficient, resulting in noise during the encoding process, affecting the prediction accuracy and video quality.

Method used

By adaptively selecting the cropping range based on the additional information of the input signal and the internal codec characteristics, the cropping noise in the decoded video data is reduced. The specific method includes receiving syntax elements in the video code stream, derive a cropping range, and performing a cropping operation on the encoded video data.

Benefits of technology

The noise problem during the encoding process is improved, the prediction accuracy and video quality are improved, and the efficiency of video encoding and decoding is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226359A_ABST
    Figure CN120226359A_ABST
Patent Text Reader

Abstract

Various implementations described herein include methods and systems for encoding video. In one aspect, a method includes receiving a video bitstream, the video bitstream including encoded video data and at least one syntax element indicating at least one clipping range. The method includes deriving at least one clipping range of a portion of the encoded video data based on the at least one syntax element. Each of the at least one clipping range is used for modifying a sampling value range of the received video code stream. The method includes performing at least one clipping operation on the encoded video data using the derived at least one clipping range.
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 532,664, filed on Aug. 14, 2023, entitled "Adaptive Range for Clipping Processes for Video and Image Compression", and is a continuation of, and claims priority to, U.S. Patent Application No. 18 / 502,003, filed on Nov. 4, 2023, entitled "Adaptive Range for Clipping Processes for Video and Image Compression", all of which are hereby incorporated by reference in their entirety. Technical Field

[0002] The disclosed embodiments generally relate to video coding and decoding, including but not limited to systems and methods for providing an adaptive range for clipping processes for video and image compression. Background Art

[0003] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices transmit and receive digital video data over a communication network or otherwise convey digital video data, and / or store digital video data on a storage device. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding can be used to compress video data according to at least one video coding standard before transmitting or storing the video data. Video coding and decoding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services.

[0004] Video coding and decoding typically utilize prediction methods that exploit the redundancy inherent in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video coding and decoding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was published by ITU-T and ISO / IEC in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended as a successor to HEVC. The VVC / H.266 standard was published by ITU-T and ISO / IEC in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verified version 1.0.0 of the specification with errata 1 was released. Summary of the Invention

[0005] The present disclosure describes adaptively selecting a cropping range based on additional information from an input signal and / or internal codec characteristics. Using the adaptively determined cropping range can improve the encoding process by reducing noise from the cropped portion of the decoded video data, which may be introduced by the codec itself (e.g., such codec noise may affect prediction accuracy), such as during the quantization step (e.g., as quantization noise) and / or in regions where no useful data is encoded.

[0006] According to some embodiments, a video decoding method is provided. The method includes: receiving a video bitstream (e.g., an encoded video sequence), the video bitstream including encoded video data and at least one syntax element indicating at least one cropping range. The method includes: deriving, based on the at least one syntax element, the at least one cropping range for a portion of the encoded video data. Each of the at least one cropping ranges is used to modify a sampling value range of the received video bitstream. The method includes: performing at least one cropping operation on the encoded video data using the derived at least one cropping range.

[0007] According to some embodiments, a video encoding method is provided. The method includes: receiving a current picture; determining, based on characteristics of the current picture, a first portion of the current picture to be cropped with a first cropping range, where the first cropping range is different from a default range associated with the current picture; and encoding information indicating the first cropping range as a video bitstream.

[0008] According to some embodiments, a computing system is provided, such as a streaming media system, a server system, a personal computer system, or other electronic devices. The computing system includes a control circuit and a memory storing at least one set of instructions. The at least one set of instructions includes instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component, e.g., a transcoder.

[0009] According to some embodiments, a non - volatile computer - readable storage medium is provided. The non - volatile computer - readable storage medium stores at least one set of instructions for execution by a computing system. The at least one set of instructions includes instructions for performing any of the methods described herein.

[0010] Thus, devices and systems having methods for encoding and decoding video are disclosed. Such methods, devices, and systems may supplement or replace conventional methods, devices, or systems for video encoding / decoding. The features and advantages described in the specification are not necessarily all - encompassing. In particular, given the figures, specification, and claims provided in this disclosure, additional features and advantages will be apparent to those of ordinary skill in the art. Further, it should be noted that the language used in the specification is primarily selected for readability and teaching purposes and is not necessarily for describing or limiting the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] For a more detailed understanding of the present disclosure, a more specific description can be made by referring to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the drawings only illustrate the relevant features of the present disclosure and are therefore not necessarily considered restrictive, as those skilled in the art will understand that there may be other effective features in this specification after reading the present disclosure.

[0012] Figure 1 FIG. 15 is a block diagram of an example communication system according to some embodiments.

[0013] Figure 2A FIG. 19 is a block diagram of example elements of an encoder component according to some embodiments.

[0014] Figure 2B FIG. 23 is a block diagram of example elements of a decoder component according to some embodiments.

[0015] Figure 3 FIG. 27 is a block diagram of an example server system according to some embodiments.

[0016] Figure 4A FIG. 31 is a flowchart of an example method of video encoding according to some embodiments.

[0017] Figure 4BFlowchart of an example method for video decoding according to some embodiments.

[0018] In accordance with conventional practice, the various features illustrated in the drawings need not be drawn to scale, and like reference numerals may be used throughout the specification and drawings to designate like features. Detailed Description

[0019] The present disclosure describes adaptively selecting a cropping range based on additional information from an input signal and / or internal codec characteristics, among other things. The additional information may be signaled in the bitstream or may be derived at the decoder. As used herein, "cropping range" may relate to any and all combinations of signal color components (e.g., Y, Cb, Cr, R, G, B, or others). Using an adaptively determined cropping range can improve the encoding process by reducing noise from cropped portions of the decoded video data that may be introduced by the codec itself (e.g., such codec noise may affect prediction accuracy), such as during a quantization step (e.g., as quantization noise) and / or in regions where no useful data is encoded. Exemplary systems and devices

[0020] Figure 1 A block diagram of a communication system 100 according to some embodiments is shown. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via at least one network. In some embodiments, the communication system 100 is a streaming system, for example, used with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0021] The source device 102 includes a video source 104 (e.g., a camera component or a media memory) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates at least one encoded video bitstream from the video bitstream. The video bitstream from the video source 104 may be of high data volume compared to the encoded video bitstream generated by the encoder component 106. Since the encoded video bitstream 108 has a lower data volume (less data) compared to the video bitstream from the video source, the encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared to the video bitstream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video to at least one network 110).

[0022] At least one network 110 represents any number of networks for transmitting information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wired (wired) and / or wireless communication networks. The at least one network 110 can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0023] The at least one network 110 includes a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as an encoded video stream from the source device 102). The server system 112 includes codec components 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec components 114 include an encoder component and / or a decoder component. In various embodiments, the codec components 114 are instantiated as hardware, software, or a combination thereof. In some embodiments, the codec components 114 are configured to decode the encoded video stream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video stream 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 can be configured to trim the encoded video stream 108 to customize potentially different streams for at least one of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0024] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be presented on a display or other type of rendering device. In some embodiments, at least one of the electronic devices 120 does not include a display component (e.g., is communicatively coupled to an external display device and / or includes a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0025] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or at least one of the electronic devices 120 are examples of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0026] In an example operation of the communication system 100, the source device 102 transmits the encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may decode and / or encode the encoded video bitstream 108 using the codec component 114. For example, the server system 112 may apply an encoding that is more suitable for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., at least one encoded video bitstream) to at least one of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.

[0027] Figure 2A A block diagram showing example elements of an encoder component 106 in accordance with some embodiments is shown. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that, when viewed in sequence, produce motion. The pictures themselves may be organized as a spatial array of pixels, where each pixel may include at least one sample depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can readily understand the relationship between pixels and samples. The following description focuses on examples.

[0028] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skipping, quantizer, and / or λ (lambda) values of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as they may relate to the encoder component 106 optimized for a particular system design.

[0029] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and at least one reference picture), and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since decoding the symbol stream results in a bit-exact result regardless of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the same sample values as the sample values that the decoder will interpret during decoding when using prediction as reference picture samples. This principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, e.g., due to channel errors) is known to those of ordinary skill in the art.

[0030] The operation of the decoder 210 can be the same as the operation of a remote decoder (such as the decoder component 122), which will be described in detail below in conjunction with Figure 2B However, briefly referring to Figure 2B , since the symbols are available and the entropy encoder 214 and the parser 254 encoding / decoding the symbols into the encoded video sequence can be lossless, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.

[0031] Except for parsing / entropy decoding, the decoder techniques described herein may exist in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operation. The description of encoder techniques may be brief as they may be the opposite of decoder techniques.

[0032] As part of its operation, the source encoder 202 may perform motion-compensated predictive coding, where the source encoder predictively encodes an input frame by referring to at least one previously encoded frame designated as a reference frame from a video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of at least one reference frame that may be selected as at least one predictive reference for the input frame. The controller 204 may manage the encoding operation of the source encoder 202, including, for example, setting parameters and sub-group parameters for encoding video data.

[0033] The decoder 210 decodes the encoded video data of a frame that may be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 may advantageously be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 2A (not shown)), the reconstructed video sequence may be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that may be performed by a remote video decoder on a reference frame and may cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder component 106 stores locally copies of the reconstructed reference frames that have the same content as the reconstructed reference frames that would be obtained by a remote video decoder (in the absence of transmission errors).

[0034] The predictor 206 may perform a prediction search on the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. The predictor 206 may operate on a per-pixel-block basis of sample blocks to find an appropriate prediction reference. As determined by the search results obtained by the predictor 206, the input picture may have a prediction reference extracted from a plurality of reference pictures stored in the reference picture memory 208.

[0035] The outputs of all the above functional units may be entropy encoded in the entropy encoder 214. The entropy encoder 214 converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).

[0036] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer at least one encoded video sequence created by the entropy encoder 214 to make them ready for transmission via a communication channel 218, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to merge the encoded video data from the source encoder 202 with other data to be transmitted (e.g., encoded audio data and / or auxiliary data streams (not shown in the source)). In some embodiments, the transmitter may transmit additional data along with the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data (such as redundant pictures and slices), supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0037] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a specific encoded picture type to each encoded picture, which may affect the encoding techniques applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi - predictive picture (B picture). Intra pictures can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example, independent decoder refresh (IDR) pictures. Those of ordinary skill in the art are aware of those variants of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. Predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and reference index to predict the sample values of each block. Bi - predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures can be used to reconstruct a single block using more than two reference pictures and associated metadata.

[0038] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and encoded on a block-by-block basis. A block can be encoded predictively by referring to other (already encoded) blocks determined by the coding assignments applied to the corresponding picture of the block. For example, blocks of an I picture can be encoded non-predictively, or they can be encoded predictively (spatial prediction or intra prediction) by referring to already encoded blocks of the same picture. Pixel blocks of a P picture can be encoded non-predictively via spatial prediction or via temporal prediction by referring to a previously encoded reference picture. Blocks of a B picture can be encoded non-predictively via spatial prediction or via temporal prediction by referring to one or two previously encoded reference pictures.

[0039] Video can be captured as multiple source pictures (video pictures) in a time series. Intra picture prediction (commonly abbreviated as intra prediction) exploits spatial correlations within a given picture, and inter picture prediction exploits (temporal or other) correlations between pictures. In an example, a particular picture (which is referred to as the current picture) in encoding / decoding is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously encoded and still buffered reference picture in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture and can have a third dimension identifying the reference picture in the case of using multiple reference pictures.

[0040] The encoder component 106 can perform encoding operations according to a predetermined video coding technique or standard (such as any of the techniques or standards described herein). In its operation, the encoder component 106 can perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0041] Figure 2B A block diagram showing example elements of a decoder component 122 according to some embodiments is shown. Figure 2B The decoder component 122 in is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to a loop filter 256 and is configured to send data to the display 124 (e.g., via a wired or wireless connection).

[0042] In some embodiments, decoder component 122 includes a receiver that is coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive at least one encoded video sequence to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, and these other data may be forwarded to their respective consuming entities (not shown). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data with the encoded video. The additional data may be included as part of at least one encoded video sequence. Decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant strips, redundant pictures, forward error correction codes, etc.

[0043] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaler / inverse transform unit 258, intra prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, decoder component 122 may be implemented at least partially in software.

[0044] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to counter network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 inside the decoder component 122 (e.g., the buffer memory 252 is configured to handle playback timing), a separate buffer memory is also provided outside the decoder component 122 (e.g., to counter network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory 252 may not be needed, or the buffer memory 252 can be small. For use on a best-effort packet network such as the Internet, the buffer memory 252 may be required, the buffer memory 252 can be relatively large and can advantageously have an adaptive size, and can be implemented at least partially in an operating system or similar element (not shown) outside the decoder component 122.

[0045] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols can include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device can be in the form of, for example, Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not shown). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technology or standard, and can follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the set. Subgroups can include Group of Pictures (GOP), pictures, tiles, strips, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser 254 can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0046] Depending on the type of the encoded video picture or a part thereof (such as: inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols 270 may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information, which is parsed by the parser 254 from the encoded video sequence. For clarity, this subgroup control information flow between the parser 254 and the multiple units is not depicted below.

[0047] The decoder component 122 can be conceptually subdivided into multiple functional units. In some embodiments, many of these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, the conceptual subdivision of the functional units is maintained here.

[0048] The scaler / inverse transform unit 258 receives the quantized transform coefficients and control information (such as which transform to use, block size, quantization factor, and / or quantization scaling matrix) as at least one symbol 270 from the parser 254. The scaler / inverse transform unit 258 can output a block including sample values, which can be input into the aggregator 268.

[0049] In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is: blocks that do not use predictive information from a previously reconstructed picture, but can use predictive information from a previously reconstructed part of the current picture. Such predictive information can be provided by the intra prediction unit 262. The intra prediction unit 262 can use the surrounding reconstructed information obtained from the current (partially reconstructed) picture in the current picture memory 264 to generate a block having the same size and shape as the block being reconstructed. The aggregator 268 can add, on a per-sample basis, the predictive information already generated by the intra prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258.

[0050] In other cases, the output samples of the scaler / inverse transform unit 258 belong to inter-coded and possibly motion-compensated blocks. In such cases, the motion compensation prediction unit 260 can access the reference picture memory 266 to obtain samples for prediction. After motion-compensating the obtained samples according to the symbol 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The address in the reference picture memory 266 from which the motion compensation prediction unit 260 obtains the prediction samples can be controlled by a motion vector. The motion vector can be provided to the motion compensation prediction unit 260 in the form of a symbol 270, which can have, for example, an X component, a Y component, and a reference picture component. Motion compensation can also include interpolation of the sample values obtained from the reference picture memory 266 when using sub-sampled accurate motion vectors, motion vector prediction mechanisms, etc.

[0051] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and provided to loop filter unit 256 as symbols 270 from parser 254, but can also respond to meta-information obtained during the decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to sample values of previously reconstructed and loop filtered samples. The output of loop filter unit 256 can be a sample stream that can be output to a rendering device such as display 124, and stored in reference picture memory 266 for use in future inter-frame prediction.

[0052] Once reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is reconstructed and the encoded picture has been identified (by, e.g., parser 254) as a reference picture, the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.

[0053] Decoder component 122 can perform decoding operations according to predetermined video compression techniques that can be documented in standards such as any of the standards described herein. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that it follows the syntax of the video compression technique or standard as specified in the video compression technique document or standard, particularly in the profile document thereof. Additionally, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, and so on. In certain cases, the limits set by the level can be further restricted by the hypothetical reference decoder (HRD) specifications and metadata signaled in the encoded video sequence for HRD buffer management.

[0054] Figure 3 A block diagram of server system 112 according to some embodiments is shown. Server system 112 includes control circuit 302, at least one network interface 304, memory 314, user interface 306, and at least one communication bus 312 for interconnecting these components. In some embodiments, control circuit 302 includes at least one processor (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes at least one field programmable gate array (FPGA), hardware accelerator, and / or at least one integrated circuit (e.g., application specific integrated circuit).

[0055] At least one network interface 304 may be configured to interface with at least one communication network (e.g., wireless, wired, and / or optical network). The communication network may be a local area network, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, and so on. Examples of communication networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANBus, and so on. Such communication may be one-way receive-only (e.g., broadcast TV), one-way transmit-only (e.g., CANbus to certain CANbus devices), or two-way (e.g., to other computer systems using local area digital networks or wide area digital networks). Such communication may include communication with at least one cloud computing network.

[0056] The user interface 306 includes at least one output device 308 and / or at least one input device 310. The at least one input device 310 may include one or more of the following: keyboard, mouse, touchpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. The at least one output device 308 may include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or monitors), etc.

[0057] The memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as at least one disk storage device, optical disk storage device, flash memory device, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes at least one storage device located remotely from the control circuit 302. The memory 314 or optionally at least one non-volatile solid-state memory device within the memory 314 includes non-volatile computer-readable storage media. In some embodiments, the memory 314 or the non-volatile computer-readable storage media of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: ● An operating system 316 that includes processes for handling various basic system services and for performing hardware-related tasks; ● A network communication module 318 that is used to connect the server system 112 to other computing devices via at least one network interface 304 (e.g., via wired and / or wireless connections); ● A codec module 320, which is used to perform various functions related to encoding and / or decoding data (such as video data). In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes at least one of the following: ○ A decoding module 322, which is used to perform various functions related to decoding encoded data, such as those functions described previously for the decoder component 122; and ○ An encoding module 340, which is used to perform various functions related to encoding data, such as those functions described previously for the encoder component 106; and ● A picture memory 352, which is used to store pictures and picture data, for example, for use by the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0058] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions described previously for the parser 254), a transform module 326 (e.g., configured to perform various functions described previously for the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions described previously for the motion compensation prediction unit 260 and / or the intra prediction unit 262), and a filter module 330 (e.g., configured to perform various functions described previously for the loop filter 256).

[0059] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions described previously for the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions described previously for the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 a subset of the modules shown. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.

[0060] Each of the above modules stored in the memory 314 corresponds to a set of instructions for performing the functions described herein. The above modules (e.g., instruction sets) need not be implemented as separate software programs, routines, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the above modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0061] Although Figure 3 FIG. illustrates a server system 112 according to some embodiments, but Figure 3 is more of a functional description of the various features that may be present in at least one server system than a structural diagram of the embodiments described herein. In practice, as will be appreciated by those of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in may be implemented on a single server, and a single item may be implemented by at least one server. The actual number of servers used to implement the server system 112, and how the features are distributed among them will vary depending on the implementation, and optionally, will depend in part on the amount of data traffic processed by the server system during peak usage as well as during average usage periods. Exemplary encoding processes and techniques

[0062] The encoding processes and techniques described below may be performed at the devices and systems described above (e.g., the source device 102, the server system 112, and / or the electronic device 120).

[0063] A hybrid video codec may include two or more of the following: intra prediction, inter prediction, transform coding, quantization, entropy coding, and post-loop in-loop filtering. The encoding processes and techniques described herein may be used in the tailoring process for different parts of the codec. In some embodiments, using an adaptive cropping range, rather than relying only on a fixed cropping range based on the internal (codec) bit depth, improves the compression efficiency of the video codec. Cropping based on the bit depth (e.g., for a cropping range of (0, 2^bitdepth - 1)) may be applied to multiple operations in the codec (such as in filtering, sampling, interpolation, weighted prediction, weighted combination, reconstruction, and other stages) to avoid data overflow, thus ensuring that the generated prediction samples and reconstructed samples remain within the defined dynamic range.

[0064] The clipping range (x, y) of the clipping process determines the minimum and maximum values of the samples after the clipping process. In some embodiments, the clipping range (x, y) includes the upper and lower bounds of the clipping range or the minimum value x and the maximum value y. For example, if the internal bit depth is set to 8 bits, each sample of the object signal is in the range of (0, 255), and the clipping process for each sample z can be set as shown in Equation 1 below.

[0065] Similarly, if the internal bit depth is set to 10 bits, the clipping process for each sample z during the entire encoding / decoding process can be as shown in Equation 2 below.

[0066] In general form, this can be written as shown in Equation 3 and Equation 4 below. ClipG(x) = Clip(0, (1 << BitDepth) - 1, x) Equation 4 - Example BitDepth Clipping Range where BitDepth represents the internal (codec) bit depth.

[0067] The encoding processes and techniques described herein allow the use of an adaptively determined clipping range (x, y) (sometimes also referred to as the "clipping process range"), which is based on additional information such as a) the dynamic range of the input signal; b) the bit depth of the input signal; and c) the dynamic range of the internal (encoded) signal. This additional information can be signaled in the bitstream and / or derived at the decoder side. In some embodiments, different color components (e.g., Y, Cb, Cr; or R, G, B) use different clipping ranges (x, y) that can be derived in the same or different ways. For example, the clipping range of the luminance component can be different from that of the two chrominance components. In some embodiments, different image regions (e.g., blocks, partitions, or sub - partitions) can use n different clipping ranges (x n , y n ) that can be derived in the same or different ways. As used herein, "clipping range / clipping ranges" generally refers to the clipping range associated with any color component (Y, Cb, or Cr) (such as at least one, each, all, or any combination of the signal color components (Y, Cb, Cr; or R, G, B)).

[0068] The encoding processes and techniques described herein can be applied to perform conversions between visual media files and bitstreams of visual media data, and the conversions can be performed according to codecs. For example, bit-depth-based cropping can be applied at multiple locations to avoid data overflow. The range of bit-depth-based cropping can be defined as [0, 2^bitdepth - 1]. Cropping operations can be applied during filtering, sampling, interpolation, weighted prediction, weighted combination, reconstruction, and / or other stages to ensure that the generated predicted and reconstructed samples remain within the defined dynamic range.

[0069] For example, for video captured with 10-bit precision, useful information can be collected within the range (0, 1023). For example, any data less than 0 or greater than 1023 should be cropped to (0, 1023). As another example, for video captured with 8 bits (e.g., following BT.2020), the data is encoded within the range (16, 235), and when the signal is increased from 8 bits to 10 bits, the signal range can be multiplied by 4 to reach (64, 940). Thus, in such a 10-bit signal (e.g., which spans 0 to 1023), useful information of the video (e.g., mainly captured or only captured) can be captured from 64 to 940 (e.g., no useful data is collected from 0 to 63 or 941 to 1023). An offset can be signaled in the bitstream to set the cropping range to (64, 940), or the cropping range (64, 940) can be set as the default value, and an offset relative to the range (64, 940) can be signaled in the bitstream. For example, an offset of 7 signaled will further shift the cropping range from (64, 940) to (57, 933). In some embodiments, two offsets are signaled in the bitstream to asymmetrically set the shifted cropping range at the upper and lower bounds. Using an adaptively determined cropping range can improve the encoding process by reducing noise from cropped portions (e.g., boundary portions) of decoded video data, which may be introduced by the codec itself (e.g., such codec noise may affect prediction accuracy), e.g., during the quantization step (e.g., as quantization noise) and / or in regions where no useful data is encoded.

[0070] Figure 4AFIG. is a flowchart illustrating a method 400 for encoding a video according to some embodiments. Method 400 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 400 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system. The system receives (402) a current picture, and the system determines (404) a first portion of the current picture to be cropped with a first cropping range based on characteristics of the current picture, wherein the first cropping range is different from a default range associated with the current picture. The system encodes (406) information indicating the first cropping range into the video bitstream.

[0071] Figure 4B FIG. is a flowchart illustrating a method 450 for decoding a video according to some embodiments. Method 450 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 450 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system. The system receives (452) a video bitstream that includes encoded video data and at least one syntax element indicating at least one cropping range (e.g., an offset increment of a base cropping range). The system derives (454) at least one cropping range for a portion of the encoded video data based on the at least one syntax element, wherein each of the at least one cropping ranges modifies a range of sample values of the received video bitstream. The system performs (456) at least one cropping operation on the encoded video data using the derived at least one cropping range.

[0072] In some embodiments, at least one indicator, such as a syntax element (e.g., an offset, or minimum and maximum values of a cropping range), is signaled in a bitstream of visual media data and used to derive a corresponding cropping range of at least one cropping process. The cropping ranges in these embodiments are determined based on characteristics of the bitstream. For example, a video bitstream may be received that includes encoded video data and at least one syntax element indicating at least one cropping range. In this example, at least one cropping range is derived based on the at least one syntax element, and at least one cropping operation is performed on the encoded video data using the derived at least one cropping range.

[0073] As an example, when the signaling syntax element includes an offset of a first cropping range (e.g., the original cropping range or a modified cropping range), the derived range is different from the first cropping range. In some embodiments, the signaling syntax element includes the minimum and maximum values of the cropping range (e.g., the range includes the minimum and maximum values). In some embodiments, the derived cropping range is smaller than the first cropping range (e.g., at least one region of the first cropping range is truncated). In some embodiments, the derived cropping range is larger than the first cropping range (e.g., outside the range of 0 to 1023). For example, the encoding process and techniques may use 2-byte (or 16-bit) operations in filtering or other precision-changing processes within the codec, and even if the codec has a bit depth smaller than 16 bits, the encoding process and techniques will amplify the input signal at some intermediate processing stages (e.g., using an expanded range relative to the original, unaltered full range).

[0074] In some embodiments, for an entire sequence, picture, slice, or tile, a cropping range (x, y) is signaled at a high level (e.g., a sequence parameter set, picture parameter set, slice parameter set, or tile parameter set, as a header, containing a syntax table with at least one offset). For example, at least one additional syntax element is used to represent the range in the bitstream. In some embodiments, more than one cropping range (x, y) is signaled in the bitstream and used for corresponding cropping processes. In some embodiments, the cropping range is signaled in a high-level syntax (HLS) element. In some embodiments, the HLS element is signaled at a level higher than the block level. For example, the HLS element may correspond to the sequence level, frame level, slice level, or tile level. As another example, the HLS element may be signaled in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptive parameter set (APS), slice header, picture header, tile header, and / or CTU header.

[0075] The bitstream includes encoded video data having at least one type of packet, including: for example, as a sequence, picture, slice, or tile. In some embodiments, a picture includes many blocks, and for each block i, a cropping range (x i , y i ) is signaled separately in the bitstream. In some embodiments, the signaling is done at a low level (such as at an encoding unit, prediction unit, or transform unit or other smaller unit), and at least one additional syntax element is used to represent at least one cropping range in the bitstream.

[0076] In some embodiments, the cropping range (x, y) is determined based on information related to the input (source) signal and / or the compression (codec) process, and the derived cropping range is used for at least one cropping process. For example, the input (source) signal corresponds to the original picture and / or the pre-filtered picture from video source 104. The cropping range (e.g., the minimum and maximum values associated with the cropping range) can be derived from the pre-filtered picture. For example, motion compensated temporal filtering (MCTF) is a preprocessing method that can be employed before video coding to improve compression efficiency. In such cases, the minimum and maximum values associated with the cropping range can be derived based on the picture pre-filtered by MCTF. For example, the encoder performs pre-filtering or receives the pre-filtered picture, and the encoder signals the minimum and maximum values associated with the pre-filtered picture. Alternatively or additionally, the encoder signals a flag indicating that the encoded picture is a pre-filtered picture or derived from a pre-filtered picture, and the decoder determines the minimum and maximum values based on the pre-filtered picture to derive the cropping range for the cropping operation.

[0077] In some embodiments, the cropping range (x, y) is determined based on the dynamic range of the input signal. For example, if the dynamic range of the input signal is the unrestricted (full) range from 0 to (1 << BitDepth) - 1, the cropping range is set to be equal to (0, (1 << BitDepth) - 1). In such a scenario, at least one syntax element can be signaled in the bitstream to identify the source signal as having an unrestricted range.

[0078] In some embodiments, the dynamic range of the input signal is a finite range from A to B, where A and B are predefined constants, and the cropping range is set to be equal to (A, B). In such a scenario, at least one syntax element can be signaled in the bitstream to derive the constants A and B. For example, the dynamic range of the input signal can be a finite range specified in one of the signal processing standards, e.g., (64, 940) for a 10-bit signal, and the cropping range is set to be equal to (64, 940). In some embodiments, the cropping range includes two endpoints representing the minimum and maximum values of the cropping range (e.g., 64 and 940). In such a scenario, at least one syntax element can be signaled in the bitstream to specify the use of the predefined constants.

[0079] In some embodiments, the clipping range (x, y) is determined based on the bit depth of the input signal. For example, for an input signal with a bit depth of n bits and a codec with a bit depth of m bits, where m ≥ n, the clipping range is set to be equal to (0, ((1 << n) - 1) << (m - n)). Thus, if the bit depth n of the input signal is 8 bits and the bit depth m of the codec is 10 bits, the clipping range can be set to be equal to (0, 1020). In such a scenario, at least one syntax element can be signaled in the bitstream to derive the bit depth n of the input signal. For example, the bit depth n of the input signal can be directly signaled and / or the bit depth n of the input signal can be derived based on the increment of the bit depth m of the codec.

[0080] In some embodiments, the clipping range (x, y) is determined based on the dynamic range of the internal (codec) signal. For example, if the internal (codec) signal is converted from one range region to another range region, the clipping range is set to be equal to the final range region. In such a case, the clipping range is determined based on the information in the bitstream corresponding to the conversion of the internal signal.

[0081] In some embodiments, the internal (codec) signal includes an indication that the original dynamic range (A, B) changes to the dynamic range (C, D) during the compression process. In such a scenario, the clipping range can be set to be equal to (C, D).

[0082] In some embodiments, the internal (codec) signal includes an indication that the original dynamic range is more restricted and, within the compression process, the original dynamic range is expanded to the full (e.g., extended or enlarged) dynamic range from 0 to (1 << BitDepth) – 1. In such a scenario, the clipping range can be set to be equal to (0, (1 << BitDepth) - 1)).

[0083] In some embodiments, for an input signal with a bit depth of n bits and a codec with a bit depth of m bits, where m < n, the clipping range is set to be equal to (0, ((1 << ) - 1)). For example, if the bit depth of the input signal is 10 bits and the bit depth of the codec is 8 bits, the clipping range can be set to be equal to (0, 255). In such a scenario, at least one syntax element can be signaled in the bitstream to derive the bit depth of the input signal.

[0084] In some embodiments, clipping is performed in the luminance mapped chrominance scaling (LMCS) mapping domain. LMCS has two main components: 1) a process for mapping input luminance code values to a new set of code values used inside the encoding loop; and 2) a luminance-dependent process for scaling chrominance residual values. The first process, luminance mapping, aims to improve the encoding efficiency of standard and high dynamic range video signals by better utilizing the range of luminance code values allowed at a specified bit depth. The second process, chrominance scaling, manages the relative compression efficiency of the luminance and chrominance components of the video signal. The luminance mapping process of LMCS is applied at the pixel sample level and is implemented using a piecewise linear model. The chrominance scaling process is applied at the chrominance block level and is implemented using a scaling factor derived from adjacent luminance samples of the reconstructed chrominance block. In some embodiments, the minimum and maximum values of the upper and lower bounds of the clipping range are derived by applying a forward LMCS lookup table (LUT) to the signaled minimum and maximum values. For example, if LMCS is enabled and clipping is performed in the LMCS mapping domain, the minimum and maximum values can be derived by applying the forward LMCS LUT to the signaled minimum and maximum values such that the signaled values are mapped to new values (e.g., the derived values) in the LMCS mapping domain.

[0085] As described above, more than one clipping range (e.g., two or more clipping ranges) can be used for the clipping process within the processing pipeline. When two or more different clipping ranges are used for corresponding clipping operations, a first clipping operation can be performed on a first portion of the encoded video data based on a first clipping range signaled in the bitstream, and a second clipping operation can be performed on a second portion of the encoded video data based on a second clipping range different from the first clipping range. As described above, each of the clipping ranges can be signaled or derived. For example, the first clipping range can be signaled, and the second clipping range can be derived based on the signaled first clipping range (e.g., by applying a forward LMCS LUT based on the signaled maximum and minimum values). In this way, instead of signaling two pairs of minimum and maximum values, only one pair (e.g., the pair of signaled values associated with the maximum and minimum values of the first clipping range) is signaled, and the second pair is derived based on the first pair. In this example, after LMCS, the signaled pair of values is applied to the first clipping range (e.g., outside the LMCS mapping domain), and the derived pair of values (e.g., the values modified from the signaled values) is applied to the second clipping range (e.g., inside the LMCS mapping domain).

[0086] In some embodiments, the adaptive cropping range may be determined based on the Group of Pictures (GOP) structure, particularly for Low-Delay B (LDB) configurations. The Group of Pictures or GOP structure specifies the order of intra-frames and inter-frames. In low-delay configurations, the first frame is an intra-frame and the other frames are encoded as generalized P-pictures or B-pictures. The Low-Delay B (LDB) configuration includes a first frame that is an intra-frame and the remaining frames are encoded as B-pictures. In some embodiments, at least one parameter of the adaptive cropping range is set based on the GOP structure.

[0087] In some embodiments, at least one of the cropping ranges (or the cropping range parameters of the adaptive cropping range) is identified based on whether the Low-Delay B (LDB) configuration is valid. For example, an adaptive cropping range or a different cropping range is used in the LDB configuration compared to a random access configuration. Thus, when the GOP structure is in a random access configuration, a first cropping range may be used, and when the GOP structure is in the LDB configuration, a second cropping range (e.g., different from the first cropping range) may be used. In some embodiments, after reordering is completed to obtain a GOP pyramid structure (e.g., a hierarchical coding structure), the cropping range and / or at least one cropping range parameter is derived or otherwise identified.

[0088] In some embodiments, an offset or delta value compared to the cropping range (e.g., relative to the cropping range) is signaled.

[0089] In some embodiments, at least one cropping range (x i , y i ) is determined based on the sum of a base cropping range value and a signaled offset value (Δx i , Δy i ) and is used for the cropping process. In some embodiments, as described above, the base cropping range value (x 0 , y 0 ) is determined based on information related to the input (source) signal and / or the compression (codec) process; and the offset values (Δx, Δy) are signaled in the bitstream for the entire sequence, picture, slice, tile, or other unit. In some embodiments, the offset values (Δx, Δy) are a single number (e.g., when the offset is symmetric at the upper and lower bounds of the cropping range). In some embodiments, the offset values (Δx, Δy) are a pair of numbers for a first offset to the lower bound of the cropping range and a second offset to the upper bound of the cropping range. Signaling may be performed at a higher level elsewhere (such as in the sequence, picture, slice, or tile (or other unit) parameter set).

[0090] In some embodiments, the cropping range (x, y) is designed (configured) to allow coverage at a specific level (e.g., can be covered). For example, first, the sequence level range (x 0 ,y 0 ) is signaled, but the signaled range can be overridden by signaling a new range (x', y') at a lower level (e.g., picture level, stripe level, or tile level), and the signaled range overrides (e.g., dominates) the range used at that level (e.g., for that picture, if a new range (x', y') is signaled at the picture level). In some embodiments, the override range (x', y') is signaled directly. In some embodiments, the offset between the signaled override range (x', y') and the original range (x 0 ,y 0 ) is signaled. For example, in a video of a concert, a picture at the sequence level can have certain characteristics of the next n pictures in the sequence (e.g., the next 10 pictures may become very dark). Thus, due to the potentially low quality of the video data, the cropping range of the next n pictures may be reduced. Therefore, the override capability allows for adaptive temporal and / or local adjustment of the cropping range. For example, an override is performed at a level lower than the originally signaled unit, an override is performed at the picture level while different settings are signaled at the sequence level, or an override is performed at the block level while different settings are signaled at the picture level.

[0091] In some embodiments, the base cropping range values (x 0 ,y 0 ) are determined based on information related to the input (source) signal and / or the compression (codec) process, e.g., as described above, and multiple signaled offset values (Δx i ,Δy i ) are signaled separately in the bitstream for lower levels (e.g., for each block). For example, each block can have a range, where different blocks have different ranges. In some embodiments, a lookup table stores multiple predetermined offset values, and the signaled indicator (e.g., the signaled index or at least one signaled index) provides information on which predetermined offset value to use. Signaling can be done at a low level (such as in an encoding unit, prediction unit, transform unit, or other lower level), and at least one additional syntax element is used to represent the range in the bitstream (e.g., the base cropping range (X 0 ,Y 0 ))).

[0092] In some embodiments, as described above, multiple base cropping range values are determined separately for each block based on information related to the input (source) signal and / or the compression (codec) process and a plurality of signaling offset values (ΔX i , ΔY i ) are signaled separately in the bitstream for a lower level (e.g., for each block). Signaling can be done at a lower level such as in a coding unit, a prediction unit, a transform unit, or other lower levels, and at least one additional syntax element is used to represent a range in the bitstream (e.g., a plurality of base clipping range values ). For example, two lookup tables can store a plurality of predetermined offset values and a plurality of (e.g., corresponding) base clipping ranges, and two signaling indicators (e.g., a signaling index or at least one signaling index) provide information respectively on which predetermined offset value will be used with which base clipping range.

[0093] In some embodiments, the derived offset values (Δx i , Δy i ) depend on the value of a quantization parameter or a quantization step size.

[0094] In some embodiments, different clipping ranges are used for different components (e.g., color components). For example, a first range (x L , y L ) is used for luminance, and a second range (x c , y c ) is used for chrominance. Alternatively, a first range (x L , y L ) is used for luminance, a second range (x cb , y cb ) is used for a first chrominance Cb component, and a third range (x cr , y cr ) is used for a second chrominance Cr component.

[0095] In some embodiments, at least one clipping range is derived using encoded information such as a range of reconstructed sample values of at least one prediction block (e.g., a prediction block applying uni - directional prediction, or a prediction block applying bi - directional prediction).

[0096] In some embodiments, more than one clipping range (e.g., two or more clipping ranges) is specified and used for the clipping process within a processing pipeline. In some embodiments, the value of a first clipping range is signaled, and the value of a second clipping range is derived at the decoder (e.g., based on the value of the first clipping range). For example, the value of the second clipping range can be derived by performing a lookup operation using the value of the first clipping range. In some embodiments, more than one clipping range is defined in an encoder - decoder and is used for the clipping process within a processing pipeline according to processing stages.

[0097] For example, specify a cropping range (e.g., a first cropping range) and use it for intra prediction, and specify another cropping range (e.g., a second cropping range) and use it for all other processing stages.

[0098] Alternatively or additionally, specify a cropping range (e.g., a first cropping range) and use it for intra prediction, specify another cropping range (e.g., a second cropping range) and use it for inter prediction, and specify yet another cropping range (e.g., a third cropping range) and use it for all other processing stages.

[0099] Alternatively or additionally, specify a cropping range (e.g., a fourth cropping range, different from at least one of the first cropping range, the second cropping range, or the third cropping range) and use it for loop filters, including but not limited to deblocking, sample adaptive offset (SAO), adaptive loop filter (ALF), and cross-component adaptive loop filter (CC-ALF). In some embodiments, different cropping ranges among two or more different cropping ranges will be used based on the GOP structure of the encoded video data.

[0100] In some embodiments, loop filters (including but not limited to deblocking, SAO, ALF, and CC-ALF) involve more than one cropping range (e.g., are associated with more than one cropping range). For example, these cropping ranges are specified for a particular filter (or filters) within the loop filter.

[0101] In some embodiments, the specification of more than one cropping range and / or the use of a particular processing stage are signaled in the high-level syntax (including but not limited to sequence-level flags, picture-level flags, sub-picture-level flags, slice-level flags, and / or tile-level flags).

[0102] In some embodiments, multiple cropping ranges are signaled in the high-level syntax (HLS), and each cropping range is associated with an index value. For each predefined processing stage (e.g., intra prediction, inter prediction, or loop filtering), the index of the cropping range applied is signaled.

[0103] In some embodiments, more than one cropping range is defined in the codec and used for the cropping process within the processing pipeline, depending on other coding tools used / can be used during the encoding / decoding process.

[0104] In some embodiments, if a certain tool or certain tools are enabled at a high level (such as at the sequence level, picture level, or slice level), a cropping range (e.g., a first cropping range) is specified and used, and if this or these tools are disabled, another cropping range (e.g., a second cropping range different from the first cropping range) is specified and used. For example, if any of the coding tools changes the signal dynamic range during processing, one cropping range (e.g., a third cropping range) is used before applying the coding tool, and another cropping range (e.g., a fourth cropping range) is used after applying the coding tool. More than two coding ranges may be used in this process. An example of such a coding tool is luminance mapping and chrominance scaling (LMCS). In some embodiments, an LMCS look-up table is used to derive one of two or more different cropping ranges. For example, the LMCS look-up table is applied to the values of the first cropping range to derive the values of the second cropping range. For example, the first cropping range may be used for a portion of the encoded video data outside the LMCS mapping domain, and the second cropping range may be used for a portion of the encoded video data within the LMCS mapping domain.

[0105] Although Figure 4A and Figure 4B a number of logical stages are shown in a particular order, stages that are not order-dependent may be reordered and other stages may be combined or split. A certain reordering or other grouping not specifically recited will be apparent to those of ordinary skill in the art, so the order and grouping presented herein are not exhaustive. Additionally, it should be recognized that these stages may be implemented in hardware, firmware, software, or any combination thereof.

[0106] Now turning to some example embodiments.

[0107] (A1) In one aspect, some embodiments include a method of video decoding (e.g., method 450). The method includes receiving a video bitstream (e.g., an encoded video sequence) that includes encoded video data and at least one syntax element that indicates (e.g., an offset increment from a base cropping range) at least one cropping range. The method includes: deriving, based on the at least one syntax element, at least one cropping range for a portion of the encoded video data (e.g., a sequence, picture, slice, tile, coding unit, prediction unit, or transform unit), wherein each of the at least one cropping ranges is for modifying a sample value range of the received video bitstream. The method includes: performing at least one cropping operation (e.g., not just on a single color component) on the encoded video data using the derived at least one cropping range. In some embodiments, these portions are high-level portions such as a sequence parameter set, picture parameter set, slice parameter set, or tile parameter set (e.g., a header, a syntax table with an offset), and a cropping range (x, y) is signaled in the bitstream for an entire sequence, picture, slice, or tile. In some embodiments, at least one cropping range (x, y) is signaled in the bitstream and used in the cropping process. In some embodiments, the one cropping range (x, y) is signaled in the bitstream for an entire sequence, picture, slice, or tile. The signaling can be performed in a high-level portion such as a sequence parameter set, picture parameter set, slice parameter set, or tile parameter set, and at least one additional syntax element can be used to represent the range in the bitstream. In some embodiments, a high-level syntax (HLS) element is used to signal the at least one cropping range.

[0108] (A2) In some embodiments of A1, the at least one cropping range includes an inclusive cropping range from 64 to 940. In some embodiments, the at least one syntax element indicates at least one shift offset, and the operation of deriving at least one cropping range can include: shifting a base cropping range based on the at least one shift offset to obtain a corresponding modified cropping range different from the base cropping range (e.g., determining a cropping range (x i , y i ) based on the sum of a base cropping range value (x i , y i ) and a signaled offset value (Δx i , Δy i ).

[0109] (A3) In some embodiments of A1 or A2, the at least one syntax element indicates at least one shift offset, and wherein deriving the at least one cropping range comprises: shifting a base cropping range based on the at least one shift offset to obtain a corresponding modified cropping range different from the base cropping range.

[0110] (A4) In some embodiments of A3, the at least one shift offset is determined based on a value of a quantization parameter or a quantization step size (e.g., determined by an encoder and signaled in a bitstream). In some embodiments, the derivation of the offset values (Δx i , Δy i ) depends on the value of the quantization parameter or the quantization step size.

[0111] (A5)In some embodiments of any one of A1 - A4, the encoded video data is encoded according to a compression process, and the method further includes: (i) deriving the bit depth n of the received video bitstream based on the at least one syntax element; (ii) when determining that the bit depth n of the received video bitstream is less than or equal to the bit depth m of the compression process, setting the clipping range of the at least one clipping range to (0, ((1 << n) - 1) << (m - n)); and (iii) when determining that the bit depth n of the received video bitstream is greater than the bit depth m of the compression process, setting the clipping range of the at least one clipping range to (0, ((1 << m) - 1)). In some embodiments, the clipping range (x, y) is determined based on the bit depth of the input signal. In some embodiments, if the bit depth of the input signal is n bits and the bit depth of the codec is m bits, where m ≥ n, the clipping range is set to be equal to (0, ((1 << n) - 1) << (m - n)). For example, if the bit depth of the input signal is 8 bits and the bit depth of the codec is 10 bits, the clipping range is set to be equal to (0, 1020). In some embodiments, at least one syntax element is signaled in the bitstream to derive the bit depth of the input signal. In some embodiments, if the bit depth of the input signal is n bits and the bit depth of the codec is m bits, where m < n, the clipping range is set to be equal to (0, ((1 << m) - 1)). For example, if the bit depth of the input signal is 10 bits and the bit depth of the codec is 8 bits, the clipping range can be set to be equal to (0, 255). In some embodiments, at least one syntax element is signaled in the bitstream to derive the bit depth of the input signal. In some embodiments, the clipping range (x, y) is determined based on the dynamic range of the input signal. In some embodiments, if the dynamic range of the input signal is the unrestricted (full) range from 0 to (1 << BitDepth) - 1, the clipping range is set to be equal to (0, (1 << BitDepth) - 1). In some embodiments, at least one syntax element is signaled in the bitstream to identify that the source signal has an unrestricted range. In some embodiments, if the dynamic range of the input signal is a finite range from A to B, where A and B are predefined constants, the clipping range is set to be equal to (A, B). In this case, at least one syntax element is signaled in the bitstream to derive the constants A and B. For example, if the dynamic range of the input signal is the finite range specified in one of the signal processing standards, e.g., (64, 940) for a 10 - bit signal, the clipping range can be set to be equal to (64, 940). In some embodiments, at least one syntax element is signaled in the bitstream to specify this case. In some embodiments, the clipping range (x, y) is determined based on the internal (codec) signal dynamic range.For example, if the internal (codec) signal is converted from one range region to another, the clipping range is set to be equal to the final range region. In this case, the clipping range is determined based on information from the bitstream corresponding to the internal signal conversion. In some embodiments, the internal (codec) signal initially has a dynamic range (A, B), and during the compression process, the dynamic range is changed to (C, D), and then the clipping range is set to be equal to (C, D).

[0112] (A6) In some embodiments of A5, the bit depth of the compression process includes an extended range associated with an intermediate processing step of the compression process (e.g., the compression process uses 2 bytes or 16 bits and may include a filtering / precision process that amplifies information at some intermediate stages), and this extended range is greater than the original range, and the bit depth is greater than the original bit depth of the compression process with a subset of the extended range. In some embodiments, the internal (codec) signal has a limited dynamic range, but during the compression process, it is extended to the full dynamic range from 0 to (1 << BitDepth)-1, and the clipping range is set to be equal to (0, (1 << BitDepth)-1).

[0113] (A7) In some embodiments of any of A1 - A6, the method further includes determining at least one base clipping range (e.g., (x 0 , y 0 )) based on the received video bitstream (e.g., the input (source) signal) and / or parameters of the compression process used to encode the encoded video data, which includes: when determining that the at least one syntax element includes a single offset parameter for an inclusive base clipping range for the at least one base clipping range, where the base clipping range is associated with a portion of the received video bitstream, setting the clipping range in the at least one clipping range based on the single offset parameter and the base clipping range; and (ii) when determining that the at least one syntax element includes multiple offset values for individual blocks: (a) when determining that the individual blocks of the received video bitstream are associated with the base clipping range in the at least one base clipping range, setting a corresponding clipping range for the individual blocks based on the multiple offset values and the base clipping range; and when determining that the individual blocks of the received video bitstream are associated with each corresponding base clipping range in the at least one base clipping range, setting a corresponding clipping range for the individual blocks based on the multiple offset values and each corresponding base clipping range. In some embodiments, at least one clipping range (x and signal offset values (Δx i , Δy i ) can be determined based on the sum of the base clipping range value i , y i), and use it in the cropping process. In some embodiments, a base cropping range value (x 0 , y 0 ) is determined based on information related to the input signal and / or the compression process; and an offset value (Δx, Δy) is signaled in the bitstream. The signaling can be performed at a high level, such as in a sequence parameter set, a picture parameter set, a slice parameter set, or a tile parameter set. In some embodiments, a base cropping range value (x 0 , y 0 ) is determined based on information related to the input signal and / or the compression process; and multiple offset values (Δx i , Δy i ) signaled for each block in the bitstream. The signaling can be performed at a low level (such as in a coding, prediction, or transform unit), and at least one additional syntax element can be used to represent the range in the bitstream. In some embodiments, multiple base cropping range values are determined separately for each block based on information related to the input (source) signal and / or the compression (codec) process and multiple offset values (Δx i , Δy i ) signaled for each block in the bitstream. The signaling can be done at a low level, such as in a coding, prediction, or transform unit, and at least one additional syntax element can be used to represent the range in the bitstream.

[0114] (A8) In some embodiments of any one of A1 - A7, the method further includes: covering a cropping range at a specified lower level in at least one cropping range, where the at least one syntax element includes the covered cropping range or the offset between the covered cropping range and the base cropping range. For example, the system is configured to apply some local changes in time. For example, a sequence - level range (x 0 , y 0 ) is first signaled, but the signaled sequence - level range can be covered by signaling a new range (x', y') at the picture level. In some embodiments, the cropping range (x, y) is designed to allow covering at a specific level. For example, a sequence - level range (x 0 , y 0 ) is signaled, and the signaled sequence - level range can be covered (e.g., at the picture level) by signaling a new range (x', y'). In some embodiments, the covering range (x', y') can be signaled directly, or by signaling the offset between the covering range (x', y') and the original range (x 0 , y 0 ).

[0115] (A9)In some embodiments of any one of A1 - A8, different clipping ranges are used for performing clipping operations on different color components. In some embodiments, different clipping ranges are used for different components. In some embodiments, the first range (x L , y L ) is for the luminance component, and the second range (x c , y c ) is for the chrominance component. In some embodiments, the first range (x L , y L ) is for luminance, the second range (x cb , y cb ) is for the first chrominance (Cb) component, and the third range (x cr , y cr ) is for the second chrominance (Cr) component.

[0116] (A10)In some embodiments of any one of A1 - A9, encoded information is used to derive at least one clipping range. In some embodiments, the clipping range can be derived based on the range of reconstructed sample values of at least one prediction block, e.g., whether single prediction or dual prediction is applied.

[0117] (A11)In some embodiments of any one of A1 - A10, the corresponding clipping range for each block i (e.g., (x i , y i )) is signaled separately in the bitstream. For example, the signaling can be performed at a low level, and at least one additional syntax element can be used to represent the range in the bitstream. In some embodiments, the clipping range values of the respective clipping ranges are determined based on information related to the input signal and / or the compression process and are used for the clipping process.

[0118] (B1)On the other hand, some embodiments include a method of video encoding (e.g., method 400). The method includes: (i) receiving a current picture (e.g., an original or pre - filtered picture); (ii) based on the characteristics of the current picture, determining a first portion of the current picture to be clipped with a first clipping range, where the first clipping range is different from the default range associated with the current picture; and (iii) encoding information indicating the first clipping range into a video bitstream. For example, the characteristics of the current picture can include the dynamic range of the signal in the current picture, the quantization noise associated with the current picture, and / or the input signal bit depth of the current picture. In some embodiments, the first clipping range is signaled in the video bitstream, and a second clipping range is derived based on the first clipping range during video decoding.

[0119] (B2)In some embodiments of B1, the method includes: receiving a video sequence including the current picture (e.g., from video source 104).

[0120] (B3) In some embodiments of B1 or B2, the method further includes encoding a current picture in a video bitstream.

[0121] (B4) In some embodiments of any one of B1 - B3, information indicating a first cropping range is signaled by at least one indicator (e.g., at least one syntax element). In some embodiments, at least one indicator is signaled using a high-level syntax.

[0122] (B5) In some embodiments of any one of B1 - B4, the method further includes: determining a second cropping range and encoding the second cropping range in a video bitstream.

[0123] (B6) In some embodiments of any one of B1 - B5, the method further includes: determining a second portion of the current picture to be cropped using the second cropping range, and signaling the second portion and / or the second cropping range in a video bitstream.

[0124] (C1) In another aspect, some embodiments include a method of performing a conversion between a visual media file and a bitstream of visual media data. The method includes: obtaining a visual media file (e.g., from video source 104); and performing a conversion between the visual media file and a bitstream of visual media data, wherein the bitstream indicates a first cropping range to be applied to a first portion of the visual media data. In some embodiments, the bitstream further indicates a second cropping range to be applied to a second portion of the visual media data. In some embodiments, the bitstream includes at least one offset value indicating a difference between the first cropping range and a default cropping range.

[0125] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including a control circuit (e.g., control circuit 302) and a memory coupled to the control circuit (e.g., memory 314), the memory storing at least one set of instructions configured to be executed by the control circuit, the at least one set of instructions including instructions for performing any one of the methods described herein (e.g., the above A1 - A10, B1 - B6, and C1).

[0126] In another aspect, some embodiments include: a non-transitory computer-readable storage medium storing at least one set of instructions, the instructions being executed by a control circuit of a computing system, the instructions including instructions for performing any one of the methods described herein (e.g., the above A1 - A10, B1 - B6, and C1).

[0127] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of at least one of the associated listed items. It should also be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of at least one other feature, integer, step, operation, element, component, and / or combination thereof.

[0128] As used herein, depending on the context, the term "if" can be interpreted to mean "when the stated precondition is true" or "once the stated precondition is true" or "in response to determining that the stated precondition is true" or "in accordance with determining that the stated precondition is true" or "in response to detecting that the stated precondition is true". Similarly, depending on the context, the phrases "if it is determined that [the stated precondition is true]" or "if [the stated precondition is true]" or "when [the stated precondition is true]" can be interpreted to mean "upon determining that the stated precondition is true" or "in response to determining that the stated precondition is true" or "in accordance with determining that the stated precondition is true" or "upon detecting that the stated precondition is true" or "in response to detecting that the stated precondition is true".

[0129] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Given the above teachings, many modifications and variations are possible. The embodiments were chosen and described in order to best explain the operating principles and practical applications, thereby enabling others skilled in the art to understand.

Claims

1. A method for video decoding performed at a computing system having a memory and at least one processor, characterized in that, The method includes: Receiving a video bitstream, the video bitstream including encoded video data and at least one syntax element indicating at least one cropping range; Deriving, based on the at least one syntax element, the at least one cropping range for a portion of the encoded video data, wherein each of the at least one cropping ranges is for modifying a range of sample values of the received video bitstream; and Performing at least one cropping operation on the encoded video data using the derived at least one cropping range.

2. The method according to claim 1, characterized in that The at least one cropping range includes an inclusive cropping range from 64 to 940.

3. The method according to claim 1, wherein The at least one syntax element indicates at least one shift offset, and wherein deriving the at least one cropping range includes: shifting a base cropping range based on the at least one shift offset to obtain a corresponding modified cropping range different from the base cropping range.

4. The method according to claim 0, characterized in that, Determining the at least one shift offset based on a value of a quantization parameter or a quantization step size.

5. The method according to claim 1, characterized in that, The encoded video data is encoded according to a compression process, and the method further includes: Deriving a bit depth n of the received video bitstream based on the at least one syntax element; When determining that the bit depth n of the received video bitstream is less than or equal to a bit depth m of the compression process, setting the cropping range in the at least one cropping range to (0, ((1 << n) - 1) << (m - n)); and When determining that the bit depth n of the received video bitstream is greater than the bit depth m of the compression process, setting the cropping range in the at least one cropping range to (0, ((1 << m) - 1)).

6. The method according to claim 0, characterized in that The bit depth m of the compression process includes an extended range associated with an intermediate processing step of the compression process, the extended range being greater than an original range, and the bit depth m being greater than an original bit depth of a subset of the compression process having the extended range.

7. The method according to claim 1, wherein Further includes: Determining at least one base cropping range based on the received video bitstream and / or parameters of a compression process used to encode the encoded video data, which includes: When determining that the at least one syntax element includes a single offset parameter for an inclusive base cropping range for the at least one base cropping range, wherein the base cropping range is associated with a portion of the received video bitstream, setting the cropping range in the at least one cropping range based on the single offset parameter and the base cropping range; and When determining that the at least one syntax element includes multiple offset values for respective blocks: When determining that the respective blocks of the received video bitstream are associated with the base cropping range in the at least one base cropping range, setting a corresponding cropping range for the respective blocks based on the multiple offset values and the base cropping range; and When determining that each block of the received video bitstream is associated with a respective one of the at least one base cropping range, a respective cropping range is set for each block based on the plurality of offset values and the respective base cropping ranges.

8. The method according to claim 1, wherein Further comprising: Covering the cropping range at a specified lower level in the at least one cropping range, wherein the at least one syntax element includes the covered cropping range or an offset between the covered cropping range and the base cropping range.

9. The method according to claim 1, wherein Different cropping ranges are used for performing cropping operations on different color components.

10. The method according to claim 1, characterized in that Using the encoded information to derive the at least one cropping range.

11. A computing system, characterized in that, Comprising: A control circuitry; A memory; And At least one set of instructions stored in the memory and configured to be executed by the control circuitry, the at least one set of instructions including instructions for: Receiving a video bitstream, the video bitstream including encoded video data and at least one syntax element indicating at least one cropping range; Deriving the at least one cropping range for a portion of the encoded video data based on the at least one syntax element, wherein each of the at least one cropping ranges modifies the range of sample values of the received video bitstream; And Performing at least one cropping operation on the encoded video data using the derived at least one cropping range.

12. The computing system according to claim 11, wherein The at least one cropping range includes an inclusive cropping range from 64 to 940.

13. The computing system according to claim 11, wherein The at least one syntax element indicates at least one shift offset, and wherein deriving the at least one cropping range includes: shifting the base cropping range based on the at least one shift offset to obtain a respective modified cropping range different from the base cropping range.

14. The computing system according to claim 0, wherein Determining the at least one shift offset based on the value of the quantization parameter or the quantization step size.

15. The computing system according to claim 11, wherein The encoded video data is encoded according to a compression process, and wherein the one or more sets of instructions further include instructions for: Deriving the bit depth n of the received video bitstream based on the at least one syntax element; When determining that the bit depth n of the received video bitstream is less than or equal to the bit depth m of the compression process, setting the cropping range in the at least one cropping range to (0, ((1 << n) - 1) << (m - n)); And When determining that the bit depth n of the received video bitstream is greater than the bit depth m of the compression process, setting the cropping range in the at least one cropping range to (0, ((1 << m) - 1)).

16. A non-volatile computer-readable storage medium, characterized in that, The non - volatile computer - readable storage medium stores at least one set of instructions configured to be executed by a computing device having a control circuitry and a memory, the at least one set of instructions including instructions for: Receiving a video bitstream, the video bitstream including encoded video data and at least one syntax element indicating at least one cropping range; Derive the at least one cropping range for a portion of the encoded video data based on the at least one syntax element, wherein each of the at least one cropping ranges modifies a range of sample values of the received video bitstream; And Perform at least one cropping operation on the encoded video data using the derived at least one cropping range.

17. The non-volatile computer-readable storage medium according to claim 16, wherein The at least one cropping range includes an inclusive cropping range from 64 to 940.

18. The non-volatile computer-readable storage medium according to claim 16, wherein The at least one syntax element indicates at least one shift offset, and wherein deriving the at least one cropping range includes: shifting a base cropping range based on the at least one shift offset to obtain a corresponding modified cropping range different from the base cropping range.

19. The non-volatile computer-readable storage medium according to claim 0, wherein Determine the at least one shift offset based on a value of a quantization parameter or a quantization step size.

20. The non-volatile computer-readable storage medium according to claim 16, wherein The encoded video data is encoded according to a compression process, and wherein the at least one set of instructions further includes instructions for: Derive a bit depth n of the received video bitstream based on the at least one syntax element; When determining that the bit depth n of the received video bitstream is less than or equal to a bit depth m of the compression process, set the cropping range in the at least one cropping range to (0, ((1 << n) - 1) << (m - n)); And When determining that the bit depth n of the received video bitstream is greater than the bit depth m of the compression process, set the cropping range in the at least one cropping range to (0, ((1 << m) - 1)).