Adaptive clipping process for video and image compression
Patent Information
- Application Number
- CN202580016917.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-04-07
- Filing Date
- 2025-04-08
- Publication Date
- 2026-09-22
Smart Images

Figure CN122804401A_ABST
Abstract
Description
[0001] Incorporation This application claims priority to U.S. Patent Application No. 19 / 172,501, filed April 7, 2025, which claims priority to U.S. Provisional Application No. 63 / 631,408, filed April 8, 2024. The entire disclosure of these earlier applications is incorporated herein by reference. Technical Field
[0002] This invention describes various aspects generally related to video encoding and decoding. Background Technology
[0003] The background description provided in this application is intended to generally present the context of the invention. To the extent that the work described in this background section is intended, the work of the currently named inventors and aspects of the description that may not have otherwise qualified as prior art at the time of filing are neither explicitly nor implicitly considered as prior art to the invention.
[0004] Image / video compression facilitates the transfer of image / video data across different devices, storage devices, and networks with minimal quality loss. In some examples, video codecs can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which compresses images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current image being reconstructed to predict samples. In another example, a video codec can use a technique called inter-frame prediction, which compresses images based on temporal redundancy. For example, inter-frame prediction can utilize motion compensation to predict samples in the current image based on previously reconstructed images. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention
[0005] Various aspects of the present invention include bitstreams, methods, and apparatus for video encoding / decoding. In some examples, the apparatus for video encoding / decoding includes processing circuitry.
[0006] Some aspects of the present invention provide a method for video encoding. In one example, at least one clipping range to be applied during encoding of at least one image by a video codec is determined, the clipping range being different from a fixed clipping range based on the bit depth of the video codec. The clipping range is applied to sample values of at least one image to generate clipped sample values. Encoded information in the bitstream is generated based on the clipped sample values. Whether a syntax element indicating the clipping range is included in the bitstream is determined based on the codec configuration of the video codec.
[0007] Some aspects of the present invention provide a method for video decoding. In one example, an encoded video stream is received. The encoded video stream includes encoded information of sample values of at least one image. At least one clipping range is determined based on at least one syntax element in the encoded video stream, which is different from a fixed clipping range based on the bit depth of a video codec used to encode / decode the encoded information. A configuration of the encoding / decoding tools in the video codec is determined based on the clipping range. At least one image is reconstructed based on the video codec, which has encoding / decoding tools configured according to the configuration.
[0008] Some aspects of the present invention provide a method for processing visual media data. In this method, a conversion between a visual media file and a bitstream of visual media data is performed according to format rules. In one example, the bitstream includes encoded information of sample values of at least one image. The format rules specify that at least one limiting range is determined based on at least one syntax element in the encoded video bitstream, which is different from a fixed limiting range based on the bit depth of a video codec used to process the sample values of at least one image. The format rules also specify that the configuration of the encoding / decoding tools in the video codec is determined based on the limiting range, and at least one image is reconstructed based on the video codec having encoding / decoding tools configured according to the configuration.
[0009] The present invention also provides an apparatus for video decoding. The apparatus for video decoding includes processing circuitry configured to implement any of the described methods for video decoding.
[0010] Aspects of the present invention also provide an apparatus for video encoding. The apparatus for video encoding includes processing circuitry configured to implement any of the described methods for video encoding.
[0011] Aspects of the present invention also provide a non-transitory computer-readable storage medium storing instructions that, when executed by a computer, cause the computer to perform any of the described methods for video decoding / encoding. Attached Figure Description
[0012] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which: Figure 1 An example block diagram of a communication system is shown schematically.
[0013] Figure 2 An example block diagram of a decoder is shown schematically.
[0014] Figure 3 An example block diagram of an encoder is shown schematically.
[0015] Figure 4 A flowchart outlining the encoding process according to some aspects of the present invention is shown.
[0016] Figure 5 A flowchart outlining the decoding process according to some aspects of the present invention is shown.
[0017] Figure 6 A computer system is illustrated schematically according to one aspect. Detailed Implementation
[0018] Figure 1 Block diagrams of some example video processing systems (100) are shown. The video processing system (100) is an example application of the disclosed subject matter, namely, video encoders and video decoders in a streaming media environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming media services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0019] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101), such as a digital camera, for creating, for example, an uncompressed video image stream (102). In one example, the video image stream (102) includes samples captured by a digital camera. The video image stream (102) is depicted as a thick line compared to the encoded video data (104) (or encoded video stream) to emphasize the high data volume of the video image stream, which may be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to enable or implement aspects of the disclosed subject matter described in more detail below. The encoded video data (104) (or encoded video stream) is depicted as a thin line compared to the video image stream (102) to emphasize the low data volume of the encoded video data, which may be stored on a streaming media server (105) for future use. At least one streaming media client subsystem, such as Figure 1Client subsystems (106) and (108) can access a streaming media server (105) to retrieve copies (107) and (109) of encoded video data (104). Client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an outgoing video image stream (111) that can be presented on a display (112) (e.g., a screen) or other display device (not depicted). In some streaming media systems, the encoded video data (104), (107), and (109) (e.g., a video stream) may be encoded according to certain video coding / compression standards. For example, those standards include the ITU-T H.265 Recommendation. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed topics are applicable in the context of VVC.
[0020] It should be noted that electronic devices (120) and (130) may include other components (not shown). For example, electronic device (120) may also include a video decoder (not shown), while electronic device (130) may also include a video encoder (not shown).
[0021] Figure 2 An example block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., receiving circuitry). The video decoder (210) may replace... Figure 1 The video decoder (110) in the example is used.
[0022] The receiver (231) can receive, for example, at least one encoded video sequence included in the bitstream, which is decoded by the video decoder (210). In one aspect, one encoded video sequence is received at a time, and the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences can be received from a channel (201), which can be a hardware / software link connected to a storage device storing the encoded video data. The receiver (231) can receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams that can be forwarded to their respective user entities (not depicted). The receiver (231) can separate the encoded video sequences from other data. To combat / mitigate network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be external to the video decoder (210) (not depicted). In other applications, a buffer memory (not depicted) may exist external to the video decoder (210), for example, to combat / mitigate network jitter, and another buffer memory (215) may exist internally to the video decoder (210), for example, to handle playback timing. The buffer memory (215) may not be necessary or may be made smaller when the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isosynchronous network. For use on best-effort packet networks such as the Internet, a buffer memory (215) may be required. The buffer memory (215) may be relatively large and may advantageously have an adaptive size, and may be implemented at least partially in a similar component (not depicted) external to the operating system or the video decoder (210).
[0023] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the encoded video sequence. Categories of those symbols include information for managing the operation of the video decoder (210) and may include information for controlling a display device, such as a component of a non-electronic device (230) but which may be coupled to the electronic device (230) (e.g., a display screen). Figure 2As shown in the diagram. Control information for one or more display devices may be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (220) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set from the encoded video sequence for at least one pixel subgroup / subgroup in the video decoder based on at least one parameter corresponding to a group. Subgroups may include Group of Pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (220) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0024] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) in order to create symbols (221).
[0025] Depending on the type of the encoded video image or its portions (e.g., inter-frame and intra-frame images, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (220). For clarity, the flow of such subgroup control information between the parser (220) and the multiple units described below is not depicted.
[0026] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units as described below. In practical implementations under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.
[0027] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives from the parser (220) quantization transform coefficients as one or more symbols (221) and control information, including the transform to be used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output blocks containing sample values, which can be input into the aggregator (255).
[0028] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks do not use predictive information from previously reconstructed images, but instead can use predictive information from previously reconstructed portions of the current image. Such predictive information can be provided by the intra-image prediction unit (252). In some cases, the intra-image prediction unit (252) uses surrounding reconstructed information obtained from the current image buffer (258) to generate blocks of the same size and shape as the blocks in the reconstruction. For example, the current image buffer (258) is used to buffer partially reconstructed current images and / or fully reconstructed current images. In some cases, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.
[0029] In other cases, the output samples of the scaler / inverse transform unit (251) may involve inter-frame coded and possibly motion-compensated blocks. In this case, the motion compensation prediction unit (253) can access the reference image memory (257) to obtain samples for prediction. After motion compensation of the obtained samples according to the block-related symbols (221), these samples can be added to the output of the scaler / inverse transform unit (251) (referred to in this case as residual samples or residual signals) by the aggregator (255) to generate output sample information. The address of the predicted samples obtained by the motion compensation prediction unit (253) in the reference image memory (257) can be controlled by motion vectors used by the motion compensation prediction unit (253) in the form of symbols (221), which may have, for example, X, Y, and reference image components. Motion compensation may also include interpolation of sample values obtained from the reference image memory (257) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0030] The output samples of the aggregator (255) can be employed by various loop filtering techniques in the loop filter unit (256). Video compression techniques may include loop filtering techniques controlled by parameters included in the encoded video sequence (also known as the encoded video stream) and available as symbols (221) from the parser (220) for use in the loop filter unit (256). Video compression may also respond to metadata obtained during the decoding of a previous (in decoding order) portion of the encoded image or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0031] The output of the loop filter unit (256) can be a sample stream, which can be output to the display device (212) and stored in the reference image memory (257) for future inter-frame image prediction.
[0032] Some encoded images can be used as reference images for future predictions after they have been fully reconstructed. For example, after the encoded image corresponding to the current image has been fully reconstructed and the encoded image has been identified as a reference image (by, for example, a parser (220)), the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.
[0033] The video decoder (210) can perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T Recommendation H.265. The encoded video sequence may conform to the syntax of the video compression technology or standard and the configuration file documented therein if the encoded video sequence follows both the syntax of the video compression technology or standard and the configuration file documented therein. Specifically, the configuration file may select certain tools from all available tools in the video compression technology or standard as the only tools available under that configuration file. For compliance, the complexity of the encoded video sequence must also be within the limits defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer indicated by signaling in the encoded video sequence.
[0034] In one aspect, the receiver (231) may receive supplemental (redundant) data along with the encoded video. This supplemental data may be part of one or more encoded video sequences. The supplemental data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The supplemental data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0035] Figure 3 An example block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) may replace Figure 1 The video encoder (103) in the example is used.
[0036] The video encoder (303) can obtain data from the video source (301) (non-video source). Figure 3 In one example, an electronic device (320) receives video samples, and a video source (301) can capture one or more video images to be encoded by a video encoder (303). In another example, the video source (301) is part of the electronic device (320).
[0037] A video source (301) can provide a sequence of source video samples in the form of a digital video sample stream to be encoded by a video encoder (303). This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual images that, when viewed sequentially, produce an animated effect. The images themselves can be organized as spatial pixel arrays, where each pixel can include at least one sample, depending on the sampling structure, color space, etc., used. Samples are described in detail below.
[0038] According to one aspect, the video encoder (303) can encode and compress images of a source video sequence into an encoded video sequence (343) in real time or as needed, according to any other time constraints. Implementing an appropriate encoding rate is a function of the controller (350). In some aspects, the controller (350) controls and is functionally coupled to other functional units described below. For clarity, the coupling is not depicted. Parameters set by the controller (350) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, GOP layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions related to the video encoder (303) optimized for a specific system design.
[0039] In some respects, the video encoder (303) is configured to operate in an encoding loop. Simply put, in one example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and one or more reference images), and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to that created by the (remote) decoder. The reconstructed sample stream (sample data) is input to a reference image memory (334). Since the decoding of the symbol stream produces a bit-by-bit accurate result unaffected by the decoder's location (local or remote), the contents of the reference image memory (334) are also bit-by-bit accurate between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same as the sample values "seen" by the decoder during decoding when using the prediction. This fundamental principle of reference image synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related fields.
[0040] The operation of the “local” decoder (333) can be the same as that of a “remote” decoder such as a video decoder (210), which has already been described above. Figure 2 A detailed description has been provided. However, a brief additional reference is also available. Figure 2 When symbols are available and the entropy encoder (345) and parser (220) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (210), including the buffer (215) and parser (220), may not be fully implemented in the local decoder (333).
[0041] In one respect, decoder techniques other than parsing / entropy decoding, which exist in the decoder, exist in the corresponding encoder with the same or substantially the same functional form. Therefore, the disclosed subject focuses on decoder operation. The description of encoder techniques can be simplified, as encoder techniques are the reverse operation of the comprehensively described decoder techniques. More detailed descriptions are only required in certain areas, and these descriptions are provided below.
[0042] In some examples, during operation, the source encoder (330) may perform motion-compensated predictive coding, i.e., predictively coding the input image against at least one previously encoded image designated as a “reference image” in the reference video sequence. In this way, the coding engine (332) encodes the differences between pixel blocks of the input image and pixel blocks of one or more reference images, which may be selected as one or more predictive references for the input image.
[0043] The local video decoder (333) can decode encoded video data of an image that can be designated as a reference image, based on symbols created by the source encoder (330). Advantageously, the operation of the encoding engine (332) can be a lossy process. When the encoded video data can be decoded by the video decoder (333), Figure 3 During decoding at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process, which can be performed by the video decoder on the reference image, and can store the reconstructed reference image in the reference image memory (334). In this way, the video encoder (303) can store a copy of the reconstructed reference image locally, which shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.
[0044] The predictor (335) can perform a prediction search against the encoding engine (332). That is, for a new image to be encoded, the predictor (335) can search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can be used as suitable prediction references for the new image. The predictor (335) can operate on a pixel-by-pixel basis based on sample blocks to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (334).
[0045] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for video data encoding.
[0046] The outputs of all the aforementioned functional units can be entropy encoded in the entropy encoder (345). The entropy encoder (345) converts these symbols into an encoded video sequence by losslessly compressing the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, and arithmetic coding.
[0047] The transmitter (340) may buffer one or more encoded video sequences created by the entropy encoder (345) in preparation for transmission via a communication channel (360), which may be a hardware / software link connected to a storage device that will store the encoded video data. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0048] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a specific encoded image type to each encoded image, which may affect the encoding techniques applicable to the corresponding image. For example, one of the following image types can typically be assigned to an image.
[0049] Intraframe images (I-images) are images that are encoded and decoded without using any other images in the sequence as prediction sources. Some video codecs allow different types of intraframe images, including, for example, Independent Decoder Refresh (IDR) images.
[0050] Predictive images (P-images) can be images that are encoded and decoded by predicting sample values for each block via intra-frame prediction or via inter-frame prediction using motion vectors and reference indices.
[0051] A bidirectional predictive image (B-image) can be an image encoded and decoded by predicting sample values for each block via intra-frame prediction or via inter-frame prediction using two motion vectors and a reference index. Similarly, a multi-predictive image can reconstruct a single block using more than two reference images and associated metadata.
[0052] The source image can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and encoded block by block. Blocks can be predicted with reference to other (already encoded) blocks, which are determined by the coding assignments of the corresponding images applied to the block. For example, blocks of an I-image can be unpredictably coded, or these blocks can be predicted with reference to already encoded blocks of the same image (spatial prediction or intra-frame prediction). Pixel blocks of a P-image can be unpredictably coded, encoded via spatial prediction, or encoded via temporal prediction with reference to a previously encoded reference image. Blocks of a B-image can be unpredictably coded, encoded via spatial prediction, or encoded via temporal prediction with reference to one or two previously encoded reference images.
[0053] The video encoder (303) can perform encoding operations according to predetermined video coding techniques or standards, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0054] In one respect, the transmitter (340) may transmit additional data along with the encoded video. The source encoder (330) may transmit such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant images and slices, SEI messages, VUI parameter set fragments, etc.
[0055] Video can be captured as multiple source images (video images) in a time-series manner. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In one example, the specific image being encoded / decoded is called the current image, which is divided into blocks. Where a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image and, in the case of multiple reference images, may have a third dimension identifying the reference images.
[0056] In some aspects, inter-frame image prediction can utilize bidirectional prediction techniques. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly representing past and future images in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. A block can be predicted using a combination of the first and second reference blocks.
[0057] Furthermore, merging mode techniques can be used to improve coding efficiency in inter-frame image prediction.
[0058] According to some aspects of the invention, prediction is performed on a block-by-block basis, such as inter-frame image prediction and intra-frame image prediction. For example, according to the HEVC standard, images in a video image sequence are divided into coding tree units (CTUs) for compression. A single CTU in an image has the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU comprises three coding tree blocks (CTBs): one luma coding tree block and two chroma coding tree blocks. Each CTU can be recursively quadtree-divided into at least one coding unit. For example, a 64×64 pixel CTU can be divided into one 64×64 pixel coding unit, or four 32×32 pixel coding units, or sixteen 16×16 pixel coding units. In one example, each coding unit is analyzed to determine its prediction type, such as inter-frame prediction or intra-frame prediction. The coding unit is divided into at least one prediction unit based on temporal and / or spatial predictability. Typically, each prediction unit comprises one luma prediction block (PB) and two chroma prediction blocks. In one aspect, prediction operations in encoding / decoding are performed on a block-by-block basis. Using the luma prediction block as an example, a prediction block comprises a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0059] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using any suitable technology. In one aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using at least one integrated circuit (IC). In another aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using at least one processor that executes software instructions.
[0060] Some aspects of the present invention provide techniques for improving compression efficiency and encoding / decoding flexibility by using different limiting ranges in different parts of the codec during the clipping process in video and image processing. The techniques of the present invention can be used individually or in combination in any order. Furthermore, each technique can be implemented by processing circuitry (e.g., at least one processor or at least one integrated circuit). In one example, at least one processor executes a program stored in a non-transitory computer-readable storage medium.
[0061] These technologies can be implemented in various video codec standards such as H.264, H.265, H.266 (VVC), AV1, and AVS.
[0062] In some embodiments, employing an adaptive limiting range, rather than a fixed limiting range that depends solely on the internal (codec) bit depth, can improve the compression efficiency of the video codec. Bit-depth-based limiting can be applied to multiple operations in the codec to avoid data overflow, for example in filtering, sampling, interpolation, weighted prediction, weighted combination, reconstruction, and other stages, ensuring that the generated predicted and reconstructed samples remain within a defined dynamic range.
[0063] Limiting operations can be used in various parts of video and image codecs. Typically, limiting is performed based on a parameter called the limiting range. The limiting range of a limiting process (also called a limiting operation) determines the minimum and maximum values of the samples after the limiting process. For example, a limiting range (x, y) includes an upper and lower limit, or the minimum value x and the maximum value y of that limiting range.
[0064] In some examples, a fixed limiting range can be used based on the codec's internal bit depth. Bit-depth-based limiting can be applied to multiple operations in the codec to avoid data overflow in stages such as filtering, sampling, interpolation, weighted prediction, weighted combination, reconstruction, and others, ensuring that the generated predicted and reconstructed samples remain within a defined dynamic range.
[0065] In one example, when the codec's internal bit depth is set to 8 bits, each sample of the target signal can be clipped within a clipping range of (0, 255). In one example, the clipping process for sample z can be represented by formula (1): Formula (1) Similarly, in one example, when the internal bit depth is set to 10 bits, for each sample throughout the encoding / decoding process... The limiting process can be represented by formula (2): Formula (2) In some respects, the general form of a limiting process with a limiting range (x, y) can be expressed by formula (3): Formula (3) In some respects, it has an internal (codec) bit depth The limiting process can be represented by formula (4): Formula (4) In some embodiments, an adaptive limiting range is used instead of a fixed limiting range that depends solely on the internal (codec) bit depth. Adaptive limiting can improve the compression efficiency of the video codec. In some examples, the limiting range can be indicated via signaling within the encoded video stream and can be applied at specific steps within the reconstruction process.
[0066] Some aspects of the present invention provide techniques for further improving the signaling indication and use of one or more clipping ranges. For example, an encoder can determine at least one clipping range to be applied during encoding of at least one image by a video codec, which differs from a fixed clipping range based on the bit depth of the video codec. The clipping range can be applied to sample values of at least one image to generate clipped sample values. The encoder can generate encoded information in the bitstream based on the clipped sample values; and can determine whether a syntax element indicating a clipping range is included in the bitstream based on the codec configuration of the video codec. At the decoding end, the decoder can determine at least one clipping range based on at least one syntax element in the encoded video bitstream, which differs from a fixed clipping range based on the bit depth of the video codec used to encode / decode the encoded information. The decoder can determine the configuration of the encoding / decoding tools in the video codec based on the clipping range, and reconstruct at least one image based on the video codec having encoding / decoding tools configured according to the configuration.
[0067] It should be noted that one or more limiting ranges are used in this invention without loss of generality. The limiting range in this invention can refer to the limiting range of any color component, such as Y, Cb, or Cr. It should be noted that the techniques described in this invention can be applied to each, all, or any combination of signal color components such as Y, Cb, Cr, R, G, and B.
[0068] In some instances, the clipping range is used at the encoding end and is not indicated by signaling in the bitstream. In some examples, the clipping range is determined at the encoding end and applied during the pre-filtering stage by clipping the filter output signal. The clipped filter output signal is then encoded into the bitstream. This clipping range is not indicated by signaling in the bitstream and is not used at the decoding end. In some examples, the pre-filtering stage is the first stage in video encoding and is performed before actual encoding. In some examples, the pre-filtering stage precedes video encoding. In one example, a denoising filter (e.g., a weak denoising filter) is used as the first stage of encoding, the filtered output from the denoising filter is clipped according to the clipping range, and the clipped signal is encoded into the bitstream. In one example, the clipping range is not included in the bitstream.
[0069] In some examples, a limiting range is applied to sample values to generate limited sample values that are the final values from the processing stage of the video codec, and intermediate values within the processing stage are not limited at all. For example, a limiting range is applied to the output from the pre-filtering stage, and no limiting is applied to any intermediate values within the pre-filtering stage.
[0070] In some respects, one or more clipping ranges can be included in the bitstream, and the number of clipping ranges indicated by signaling in the bitstream can vary depending on the codec configuration. For example, the encoder can determine the number of clipping ranges indicated by signaling based on the encoded configuration, and then indicate the number of clipping ranges in the bitstream via signaling.
[0071] In some examples, the number of clipping ranges indicated by signaling in the bitstream is determined by a specific set of codec tools used during compression / decompression. In one example, multiple clipping ranges may be used to assist each specific codec tool.
[0072] In some examples, when using a pre-filtering stage (before actual encoding), and employing a limiting range in the limiting operation after the pre-filtering stage at the encoding end, two limiting ranges can be indicated in the bitstream via signaling to represent the signal dynamic range before and after the pre-filtering stage. The two limiting ranges can then be used separately to assist specific codec tools during the reconstruction process.
[0073] For example, the two limiting ranges include a first limiting range and a second limiting range. The first limiting range represents the “pre-filtering” range of the signal input to the pre-filtering stage (e.g., the dynamic range before the pre-filtering stage), and the second limiting range represents the “pre-filtering” range of the signal output from the pre-filtering stage (e.g., the dynamic range after the pre-filtering stage). In one example, the first limiting range can be used during the loop filtering process. For example, the loop filtering process includes a Wiener loop filter configured to minimize the error signal between the estimated signal (e.g., the reconstructed signal) and the desired signal. In one example, the filter parameters of the Wiener loop filter can be efficiently determined based on the first limiting range representing the “pre-filtering” range. In another example, the second range representing the “pre-filtering” range is used during the motion compensation process. For example, at least one motion compensation tool can be configured based on the second range to achieve efficient processing.
[0074] In some respects, the applicability of one or more limiting ranges is restricted at a particular coding stage.
[0075] In some examples, when the dynamic range of the target (to be compressed) signal is modified by a function (e.g., a linear or nonlinear function) during the compression / decompression process, one or more corresponding limiting ranges are not modified accordingly by the function, but are used directly because they are derived in the limiting operation on the target signal.
[0076] For example, mapping (also known as reshaping) techniques are used in video coding to modify the dynamic range of a target signal. Mapping techniques can be used to better utilize the distribution of sample codeword values in a video image. Mapping techniques can use functions (also known as mapping functions) to change the value range of samples.
[0077] Mapping and inverse mapping can occur outside the decoding loop. For example, mapping can be applied to the encoder's input samples before core encoding; and inverse mapping can be applied to the decoder's output samples at the decoding end.
[0078] Mapping and inverse mapping can also occur within the decoding loop. In some examples, a technique called luma mapping with chroma scaling (LMCS) is used for loop reshaping. The mapping (also known as reshaping) of the luma or chroma signals is implemented inside the encoding loop.
[0079] In some examples of LMCS, at the encoder, after applying a mapping function to the original sample (to be encoded) and the predicted sample respectively, a residual signal before quantization is generated. For example, the mapping function is applied to the original sample to generate a mapped original sample, and also applied to the predicted sample to generate a mapped predicted sample. The residual signal is then calculated as the difference between the mapped original sample and the mapped predicted sample. The residual of a block can be transformed into transform coefficients, and the transform coefficients can be quantized. At the decoder, the residual can be calculated via dequantization and inverse transform. Furthermore, the mapping function is applied to the predicted sample to generate a mapped predicted sample, and then the mapped predicted sample is combined with the residual signal to form a mapped reconstructed sample. Additionally, an inverse mapping function is applied to the mapped reconstructed sample to generate a reconstructed sample.
[0080] In some examples, when using mapping and inverse mapping, the clipping range is not modified based on the mapping. For instance, the clipping range derived from the bitstream is used for clipping operations without being modified by the mapping.
[0081] In some examples, when the dynamic range of the target (to be compressed) signal is modified by a function (e.g., a linear or nonlinear function) during the compression / decompression process, one or more corresponding limiting ranges are not modified accordingly by the function, and the default limiting range is used in the limiting operation applied to the target signal.
[0082] In some respects, signaling indications for one or more limiting ranges of a target signal can utilize target signal information (e.g., content characteristics).
[0083] In some examples, for compression configurations where the input signal dynamic range is smaller than the codec's internal dynamic range, the clipping range is signaled within the original (smaller) dynamic range. In one example, when the content to be compressed / decompressed has an 8-bit dynamic range and the codec is running with a 10-bit internal dynamic range, the clipping range signaled in the bitstream is reduced to the 8-bit dynamic range. In some examples, the internal dynamic range is 10 bits, and the target signal (e.g., the output signal from the encoder, which is the input signal to the decoder) is appropriately reduced to an 8-bit dynamic range so that the target signal can be encoded using 8 bits to reduce signaling costs. At the decoding end, the decoder can appropriately expand the target signal to a 10-bit dynamic range for higher computational accuracy. In some examples, to include the clipping range for clipping 10-bit content in the codec used for compression / decompression, the clipping range is reduced to 8 bits and then included in the bitstream. At the decoding end, the decoder can expand the clipping range to 10 bits. For example, at the encoding end, the limiting range (40, 1000) used for limiting the internal signal of the video codec can be reduced to (10, 250) to be included in the bitstream. At the decoding end, the decoder decodes the limiting range (10, 250) from the bitstream, and since the internal dynamic range is 10 bits, it expands (10, 250) to (40, 1000) and applies it to the decoder.
[0084] In some examples, signaling indication for one or more clipping ranges is performed based on a sign value, which can be determined by the decoder. In one example, when the signaling range is determined by some predefined values representing the minimum and maximum possible values through offsets, the offset values are indicated as unsigned values in the bitstream. At the decoding end, the decoder can determine the appropriate sign for the offset.
[0085] In one example, the content has 8 bits, and the clipping range indicated by signaling is (10, 240). In this example, the predefined range is (0, 255), and two offsets can be applied to the predefined range to indicate the clipping range via signaling. For example, two unsigned values (10, 15) are indicated by signaling in the bitstream. At the decoding end, the decoder can apply a positive sign to the first unsigned value, which indicates the offset relative to the minimum value of the clipping range. For example, in this example, +10 is applied to 0. The decoder can also apply a negative sign to the second unsigned value, which indicates the offset relative to the maximum value of the clipping range. For example, in this example, -15 is applied to 255. Therefore, the decoder can deduce the clipping range as (10, 240).
[0086] Figure 4A flowchart illustrating an overview process (400) according to one aspect of the present invention is shown. The process (400) can be used with a video encoder. In various aspects, the process (400) is executed by processing circuitry, such as processing circuitry that performs the functions of a video encoder (103), processing circuitry that performs the functions of a video encoder (303), etc. In some aspects, the process (400) is implemented in the form of software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (400). The process begins at (S401) and continues to (S410).
[0087] At (S410), at least one limiting range to be applied during the encoding of at least one image by the video codec is determined, which is different from a fixed limiting range based on the bit depth of the video codec.
[0088] At (S420), the clipping range is applied to the sample values of at least one image to generate clipped sample values.
[0089] At (S430), the encoded information in the bitstream is generated based on the clipped sample values.
[0090] At (S440), the codec configuration of the video codec determines whether the bitstream includes a syntax element indicating the clipping range.
[0091] In some respects, a clipping range is applied to the sample values output from the pre-filtering stage of the video codec to generate clipped sample values. These clipped sample values are then provided as input to the video codec for encoding. In one example, the clipping range is excluded from the bitstream when the reconstruction performed by the video codec is independent of the clipping range.
[0092] In some respects, it is determined whether at least one encoding / decoding tool of the video codec uses the limiting range information, and when at least one encoding / decoding tool of the video codec uses the limiting range information, it indicates that the syntax element of the limiting range will be included in the bitstream.
[0093] In some aspects, multiple limiting ranges to be applied during encoding by the video codec are determined. Based on the codec configuration of the video codec, at least one limiting range is selected from the multiple limiting ranges, each of which is used by at least one tool of the video codec. Therefore, at least one syntax element indicating at least one limiting range is included in the bitstream.
[0094] In some examples, a first clipping range and a second clipping range are determined to be applied during encoding by the video codec. The first clipping range is applied to the input values of the pre-filtering stage of the video codec, and the second clipping range is applied to the output values of the pre-filtering stage of the video codec. The output values clipped within the second clipping range are encoded into the bitstream. The first and second clipping ranges are determined for reconstruction performed by the video codec, respectively. At least one syntax element indicating the first and second clipping ranges is included in the bitstream.
[0095] In one example, information about a first limiting range is used to determine the loop filter used for reconstruction in the video codec, for example, to configure the loop filter. In another example, information about a second limiting range is used to determine the motion compensation tool of the video codec, for example, to configure the motion compensation tool.
[0096] In some examples, the target signal to which the clipping operation is applied (e.g., the input of a stage in a video codec, the output of a stage in a video codec, etc.) has a dynamic range that is modified based on a function. In one example, a clipping operation with a clipping range is applied to the target signal without modifying the clipping range according to a function. In another example, a clipping operation with a default clipping range is applied to the target signal. In yet another example, a clipping operation with a modified clipping range is applied to the target signal. In some examples, the function can be a linear function. In some examples, the function can be a non-linear function.
[0097] In some aspects, the video codec uses an internal dynamic range corresponding to a first bit depth, which is higher than the second bit depth used by the input signal of the video codec. The clipping range of the internal signal to be applied to the video codec is reduced to a scaled clipping range with the second bit depth. The scaled clipping range with the second bit depth is included in the bitstream. In one example, the internal dynamic range of the video codec has a 10-bit depth, and the input signal of the video codec used for compression / decompression has an 8-bit dynamic range. The clipping range to be applied to the internal signal can then be reduced from 10 bits to 8 bits, and the scaled clipping range is included in the bitstream.
[0098] In some examples, at least one unsigned value is included in the bitstream, indicating an offset of the clipping range relative to a predefined clipping range. This offset can be applied to one of the minimum and maximum values of the predefined clipping range.
[0099] In some examples, a limiting range is applied to sample values to generate limited sample values that are the final values from the processing stage of the video codec, and intermediate values within the processing stage are not limited at all. For example, a limiting range is applied to the output from the pre-filtering stage, and no limiting is applied to intermediate values within the pre-filtering stage.
[0100] The process then continues to (S499) and terminates.
[0101] The process (400) can be adjusted as appropriate. One or more steps in the process (400) can be modified and / or omitted. One or more additional steps can be added. Any suitable implementation order can be used.
[0102] Figure 5 A flowchart illustrating an overview process (500) according to one aspect of the present invention is shown. The process (500) can be used with a video decoder. In various aspects, the process (500) is executed by processing circuitry, such as processing circuitry that performs the functions of the video decoder (110), processing circuitry that performs the functions of the video decoder (210), etc. In some aspects, the process (500) is implemented in the form of software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (500). The process begins at (S501) and continues to (S510).
[0103] At (S510), an encoded video stream is received. The encoded video stream includes encoded information of sample values of at least one image.
[0104] At (S520), at least one limiting range is determined based on at least one syntax element in the encoded video bitstream, the limiting range being different from a fixed limiting range based on the bit depth of the video codec used to encode / decode the encoded information.
[0105] At (S530), the configuration of the codec tools in the video codec is determined based on the limiting range.
[0106] At (S540), at least one image is reconstructed based on a video codec having encoding and decoding tools configured according to the configuration.
[0107] In some respects, during the encoding of sample values, a limiting range is applied to the input values in the pre-filtering stage. Information about the limiting range is used to determine the configuration of the loop filters used for reconstruction in the video codec.
[0108] In some respects, during the encoding of sample values, a limiting range is applied to the output values of the pre-filtering stage. Information about the limiting range is used to determine the configuration of the motion compensation tools for the video codec.
[0109] In some examples, the target signal to which the clipping operation is applied has a dynamic range that is modified by a function. In one example, a clipping operation with a clipping range is applied to the target signal without modifying the clipping range according to a function. In another example, a clipping operation with a default clipping range (e.g., a default clipping range based on bit depth) is applied to the target signal. In yet another example, a clipping operation with a modified clipping range (e.g., modified according to a function) is applied to the target signal.
[0110] It should be noted that the function can be a linear function or a nonlinear function.
[0111] In some aspects, video codecs use an internal dynamic range corresponding to a first bit depth, which is higher than the second bit depth used by the input signal of the video codec. In some examples, the clipping range obtained from the encoded video bitstream is based on the second bit depth, and the clipping range can be expanded to obtain a scaled clipping range with the first bit depth. Clipping operations with scaled clipping ranges can be applied to the internal signal of the video codec.
[0112] In some aspects, the clipping range can be indicated by signaling as one or more unsigned values, and the symbol can be appropriately determined at the decoding end. In some examples, at least one unsigned value is determined based on at least one syntax element in the encoded video bitstream. The symbol of this unsigned value can be determined. The clipping range can be determined based on the unsigned value and the symbol.
[0113] In one example, a first unsigned value and a second unsigned value are determined based on at least one syntax element in the encoded video bitstream. A first offset is determined as a combination of a positive sign and the first unsigned value. The first offset is applied to the minimum value of a predefined range (e.g., based on bit depth) to calculate the minimum value of the clipping range. A second offset is determined as a combination of a negative sign and the second unsigned value. The second offset is applied to the maximum value of a predefined range to calculate the maximum value of the clipping range.
[0114] In some examples, a clipping range is applied to sample values to generate clipped sample values that are the final values from the processing stage of the video codec, and intermediate values within the processing stage are not clipped at all. For example, a clipping range is applied to the output from the post-filtering stage, and no clipping is applied to intermediate values within the post-filtering stage.
[0115] The process then continues to (S599) and terminates.
[0116] The process (500) can be adjusted as appropriate. One or more steps in the process (500) can be modified and / or omitted. One or more additional steps can be added. Any suitable implementation order can be used.
[0117] According to one aspect of the present invention, a method for processing visual media data is provided. In this method, a conversion between a visual media file and a bitstream of visual media data is performed according to format rules. For example, the bitstream may be a bitstream decoded / encoded using any of the decoding and / or encoding methods described in this application. The format rules may specify at least one constraint on the bitstream and / or at least one process to be performed by the decoder and / or encoder.
[0118] In one example, the bitstream includes encoded information of sample values for at least one image. The format rules specify that at least one limiting range is determined based on at least one syntax element in the encoded video bitstream. This limiting range differs from a fixed limiting range based on the bit depth of the video codec used to process the sample values of at least one image. The format rules also specify that the configuration of the encoding / decoding tools in the video codec is determined based on the limiting range, and that at least one image is reconstructed based on the video codec, which has encoding / decoding tools configured according to this configuration.
[0119] The above-described technologies can be implemented in the form of computer software using computer-readable instructions and physically stored in at least one computer-readable medium. For example, Figure 6 A computer system (600) suitable for implementing certain aspects of the disclosed subject matter is shown.
[0120] Computer software can be written using any suitable machine code or computer language. This machine code or computer language may be processed by mechanisms such as assembly, compilation, and linking to generate code containing instructions that can be executed directly by at least one computer central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, or other methods.
[0121] These instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0122] Figure 6The components of the computer system (600) shown are merely examples and are not intended to imply any limitation on the scope or functionality of computer software used to implement various aspects of the invention. The configuration of the components should also not be construed as having any dependency or requirement on any component or combination of components shown in the example aspects of the computer system (600).
[0123] The computer system (600) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from at least one human user via, for example, tactile input (e.g., key presses, swipes, data glove movements), audio input (e.g., voice, clapping sounds), visual input (e.g., gestures), or olfactory input (not depicted). The human-machine interface devices may also be used to acquire media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0124] Human-machine interface input devices may include at least one of the following (only one of each is depicted): keyboard (601), mouse (602), touchpad (603), touch screen (610), data glove (not shown), joystick (605), microphone (606), scanner (607), and camera (608).
[0125] The computer system (600) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of at least one human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback provided by a touchscreen (610), a data glove (not shown), or a joystick 605, but there may also be tactile feedback devices that do not act as input devices), audio output devices (e.g., speakers (609), headphones (not depicted)), visual output devices (e.g., screens (610), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input functionality, each with or without tactile feedback functionality—some of which may be able to output two-dimensional or more than three-dimensional visual outputs through stereoscopic output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), and printers (not depicted).
[0126] The computer system (600) may also include user-accessible storage devices and related media, such as optical media or similar media (621) including CD / DVD ROM / RW (620) equipped with CD / DVD, flash drives (622), removable hard disk drives or solid-state drives (623), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), and dedicated ROM / ASIC / PLD devices such as security dongles (not depicted).
[0127] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0128] The computer system (600) may also include an interface (654) connected to at least one communication network (655). The network may be, for example, a wireless network, a wired network, or an optical network. The network may also be a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), an in-vehicle and industrial network, a real-time network, a latency-tolerant network, etc. Examples of networks include: local area networks such as Ethernet and wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; in-vehicle and industrial networks including CAN buses, etc. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system (600)) attached to some general-purpose data port or peripheral bus (649); other networks are typically integrated into the core of the computer system (600) via a system bus described below (e.g., integrated into a PC computer system via an Ethernet interface, or integrated into a smartphone computer system via a cellular network interface). Using any of these networks, the computer system (600) can communicate with other entities. This communication can be one-way (e.g., broadcasting TV), one-way (e.g., CAN bus to certain CAN bus devices), or it can be bidirectional, such as connecting to other computer systems using a local area or wide area digital network. Certain protocols and protocol stacks can be used on each of the networks and network interfaces described above.
[0129] The aforementioned human-machine interface device, user-accessible storage device, and network interface can be attached to the core (640) of the computer system (600).
[0130] The core (640) may include at least one CPU (641), GPUs (642), a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) (643), a hardware accelerator (644) for certain tasks, a graphics adapter (650), etc. These devices, along with read-only memory (ROM) (645), random-access memory (RAM) (646), and internal mass storage devices (647) such as internal non-user-accessible hard disk drives, SSDs, etc., can be connected via a system bus (648). In some computer systems, the system bus (648) may be accessed in the form of at least one physical connector, allowing for expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (648) or connected to the core's system bus (648) via a peripheral bus (649). In one example, a screen (610) may be connected to a graphics adapter (650). Peripheral bus architectures include PCI, USB, etc.
[0131] CPUs (641), GPUs (642), FPGAs (643), and accelerators (644) can execute certain instructions that combine to form the aforementioned computer code. This computer code can be stored in ROM (645) or RAM (646). Temporary data can also be stored in RAM (646), while permanent data can be stored, for example, in an internal mass storage device (647). Fast storage and retrieval of any of the storage devices can be achieved by using a cache memory, which can be closely associated with at least one CPU (641), GPU (642), mass storage device (647), ROM (645), RAM (646), etc.
[0132] The computer-readable medium may contain computer code for performing various computer-implemented operations. The medium and computer code may be designed and constructed specifically for the purposes of this invention, or may belong to a class well-known and available to those skilled in the art of computer software.
[0133] For example, but not as a limitation, a computer system having an architecture (600), particularly a core (640), can provide functionality generated when one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in at least one tangible computer-readable medium. Such a computer-readable medium can be a medium associated with certain storage devices (e.g., internal core mass storage (647) or ROM (645)) that are user-accessible mass storage devices as described above and with the non-volatile nature of the core (640). Software implementing various aspects of the invention can be stored in such a device and executed by the core (640). Depending on specific needs, the computer-readable medium may include at least one storage device or chip. The software can cause the core (640), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute a particular process or a particular portion of a particular process described in this application, including defining data structures stored in RAM (646) and modifying such data structures according to the process defined by the software. Alternatively or as an alternative, the computer system may provide functionality generated by hard-wired logic or otherwise embodied in circuitry (e.g., an accelerator (644)) that operates in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may cover logic, and vice versa. Where appropriate, references to computer-readable media may cover circuitry (e.g., integrated circuits) storing software for execution, circuitry embodying logic for execution, or both. This invention covers any suitable combination of hardware and software.
[0134] The terms “at least one of…” or “one of…” as used in this invention are intended to include any one or a combination of the listed elements. For example, mentioning at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A to C are intended to include only A, only B, only C, or any combination thereof. Mentioning one of A or B and one of A and B are intended to include either A or B or (A and B). Where applicable, such as when the elements are not mutually exclusive, the use of “one of…” does not exclude any combination of the listed elements.
[0135] Although several examples of various aspects have been described in this invention, modifications, substitutions, and various alternative equivalents exist that fall within the scope of this invention. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, while not explicitly stated or described herein, embody the principles of the invention and are therefore within its spirit and scope.
[0136] The aforementioned disclosure also covers the following features. These features can be combined in various ways, and are not limited to the following combinations.
[0137] (1). A video coding method comprising: determining at least one clipping range to be applied during encoding of at least one image by a video codec, the clipping range being different from a fixed clipping range based on the bit depth of the video codec; applying the clipping range to sample values of the at least one image to generate clipped sample values; generating encoded information in a bitstream based on the clipped sample values; and determining whether a syntax element indicating the clipping range is included in the bitstream based on a codec configuration of the video codec.
[0138] (2). The method according to feature (1), wherein the application includes: applying a limiting range to sample values output from the pre-filtering stage of the video codec to generate a limiting sample value; and providing the limiting sample value as input for encoding by the video codec, wherein the limiting range is excluded from the bitstream when the reconstruction performed by the video codec is independent of the limiting range.
[0139] (3). The method according to any one of features (1) to (2) further includes: determining whether at least one encoding / decoding tool of the video codec uses the limiting range information; and when at least one encoding / decoding tool of the video codec uses the limiting range information, including the syntax element indicating the limiting range in the bitstream.
[0140] (4). The method according to any one of features (1) to (3) further includes: determining a plurality of limiting ranges to be applied during encoding performed by the video codec; selecting at least one limiting range from the plurality of limiting ranges based on the codec configuration of the video codec, the at least one limiting range being used by at least one tool of the video codec; and including at least one syntax element indicating at least one limiting range in the bitstream.
[0141] (5). The method according to any one of features (1) to (4) further includes: determining a first limiting range and a second limiting range to be applied during encoding performed by the video codec, applying the first limiting range to the input values of the pre-filtering stage of the video codec, and applying the second limiting range to the output values of the pre-filtering stage of the video codec, encoding the output values limited in the second limiting range into the bitstream; determining that the first limiting range and the second limiting range are respectively used for reconstruction performed by the video codec; and including at least one syntax element indicating the first limiting range and the second limiting range in the bitstream.
[0142] (6). A method according to any one of features (1) to (5), wherein determining the first limiting range and the second limiting range for reconstruction respectively includes: determining information on the use of the first limiting range by the loop filter for reconstruction in the video codec, and applying the first limiting range to the input value of the pre-filtering stage of the video codec.
[0143] (7). The method according to any one of features (1) to (6), wherein determining the first limiting range and the second limiting range for reconstructing includes: determining information on the use of the second limiting range by the motion compensation tool of the video codec, and applying the second limiting range to the output value of the pre-filtering stage of the video codec.
[0144] (8). A method according to any one of features (1) to (7), wherein the target signal for which the limiting operation is applied has a dynamic range that is modified based on a function, and the method includes at least one of the following: applying a limiting operation with a limiting range to the target signal without modifying the limiting range according to a function; applying a limiting operation with a default limiting range to the target signal; or applying a limiting operation with a modified limiting range to the target signal.
[0145] (9). The method according to any one of features (1) to (8), wherein the function includes at least one of linear functions and nonlinear functions.
[0146] (10). The method according to any one of features (1) to (9) further includes: determining that the video codec uses an internal dynamic range corresponding to a first bit depth, the first bit depth being higher than a second bit depth used by the input signal of the video codec; narrowing the limiting range of the internal signal to be applied to the video codec to a scaled limiting range having a second bit depth; and including the scaled limiting range having a second bit depth in the bitstream.
[0147] (11). The method according to any one of features (1) to (10) further includes: including at least one unsigned value in the bitstream to indicate a clipping range, the unsigned value indicating an offset relative to one of the minimum and maximum values of a predefined clipping range.
[0148] (12). A method according to any one of features (1) to (11), wherein: the application includes applying a limiting range to sample values to generate limiting sample values, the sample values being the final values from the processing stage of the video codec, and intermediate values within the processing stage not being limited.
[0149] (13). A method for video decoding, comprising: receiving an encoded video stream, the encoded video stream including encoded information of sample values of at least one image; determining at least one limiting range based on at least one syntax element in the encoded video stream, the limiting range being different from a fixed limiting range based on the bit depth of a video codec used to encode / decode the encoded information; determining a configuration of encoding / decoding tools in the video codec based on the limiting range; and reconstructing at least one image based on the video codec having encoding / decoding tools configured according to the configuration.
[0150] (14). According to the method of feature (13), wherein during the encoding of sample values, the amplitude limiting range is applied to the input values of the pre-filtering stage, determining the configuration of the codec tool includes: determining the configuration of the loop filter for reconstruction in the video codec based on the amplitude limiting range information.
[0151] (15). The method according to any one of features (13) to (14), wherein during the encoding of sample values, the limiting range is applied to the output values of the pre-filtering stage, and the configuration of the codec tool is determined by: determining the configuration of the motion compensation tool of the video codec based on the information of the limiting range.
[0152] (16). A method according to any one of features (13) to (15), wherein the target signal to be applied by the limiting operation has a dynamic range that is modified by a function, and the method includes at least one of the following: applying a limiting operation with a limiting range to the target signal without modifying the limiting range according to a function; applying a limiting operation with a default limiting range to the target signal; or applying a limiting operation with a modified limiting range to the target signal.
[0153] (17). According to any one of the features (13) to (16), wherein the function includes at least one of linear functions and nonlinear functions.
[0154] (18). A method according to any one of features (13) to (17), wherein the video codec uses an internal dynamic range corresponding to a first depth, the first depth being higher than a second depth used by the input signal of the video codec, and the method includes: expanding the limiting range to obtain a scaled limiting range having the first depth; and applying the limiting operation having the scaled limiting range to the internal signal of the video codec.
[0155] (19). The method according to any one of features (13) to (18), wherein determining the limiting range includes: determining at least one unsigned value based on at least one syntax element in the encoded video bitstream; determining the sign of the unsigned value; and determining the limiting range based on the unsigned value and the sign.
[0156] (20). A method according to any one of features (13) to (19), wherein determining the limiting range comprises: determining a first unsigned value and a second unsigned value based on at least one syntax element in the encoded video bitstream; determining a minimum value to which a first offset is applied to a predefined range to calculate the minimum value of the limiting range, the first offset being a combination of a positive sign and a first unsigned value; and determining a maximum value to which a second offset is applied to a predefined range to calculate the maximum value of the limiting range, the second offset being a combination of a negative sign and a second unsigned value.
[0157] (21). The method according to any one of features (13) to (20) further includes: applying the limiting range to the final value from the processing stage of the video codec, wherein intermediate values within the processing stage are not limited.
[0158] (22). A method for processing visual media data, the method comprising processing a bitstream of visual media data according to a format rule, wherein: the bitstream includes encoded information of sample values of at least one image; and the format rule specifies: determining at least one limiting range based on at least one syntax element in the encoded video bitstream, the limiting range being different from a fixed limiting range based on the bit depth of a video codec used to process the sample values of at least one image; determining a configuration of encoding / decoding tools in the video codec based on the limiting range; and reconstructing at least one image based on the video codec having encoding / decoding tools configured according to the configuration.
[0159] (23). An apparatus for video encoding, comprising processing circuitry configured to perform a method according to any one of features (1) to (12).
[0160] (24). An apparatus for video decoding, comprising processing circuitry configured to perform a method according to any one of features (13) to (21).
[0161] (25). A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform a method according to any one of features (1) to (22).
Claims
1. A video encoding method, characterized in that, include: Determine at least one limiting range to be applied during the encoding of at least one image by a video codec, the limiting range being different from a fixed limiting range based on the bit depth of the video codec; The limiting range is applied to the sample values of the at least one image to generate limited sample values; The encoded information in the bitstream is generated based on the amplitude-limited sample values; as well as The codec configuration of the video codec determines whether the bitstream includes a syntax element indicating the clipping range.
2. The method according to claim 1, characterized in that: The applications include: The clipping range is applied to the sample values output from the pre-filtering stage of the video codec to generate the clipped sample values; and The clipped sample values are provided as input for the video codec to perform the encoding. When the reconstruction performed by the video codec does not depend on the clipping range, the clipping range is excluded from the bitstream.
3. The method according to any one of claims 1 to 2, characterized in that, Also includes: Determine whether at least one encoding / decoding tool of the video codec uses the information of the limiting range; as well as When the at least one encoding / decoding tool of the video codec uses the information of the limiting range, it will include the syntax element indicating the limiting range in the bitstream.
4. The method according to any one of claims 1 to 3, characterized in that, Also includes: Determine multiple limiting ranges to be applied during the encoding process performed by the video codec; Based on the codec configuration of the video codec, at least one limiting range is selected from the plurality of limiting ranges, and the at least one limiting range is used by at least one tool of the video codec; as well as At least one syntax element indicating the at least one clipping range shall be included in the bitstream.
5. The method according to claim 4, characterized in that, Also includes: Determine a first limiting range and a second limiting range to be applied during the encoding process performed by the video codec, apply the first limiting range to the input values of the pre-filtering stage of the video codec, and apply the second limiting range to the output values of the pre-filtering stage of the video codec, and encode the output values limited in the second limiting range into the bitstream; The first limiting range and the second limiting range are respectively used for the reconstruction performed by the video codec; as well as At least one syntax element indicating the first clipping range and the second clipping range shall be included in the bitstream.
6. The method according to claim 5, characterized in that, The determination of the first limiting range and the second limiting range for the reconstruction includes: The loop filter used for reconstruction in the video codec is determined to use information from the first limiting range, and the first limiting range is applied to the input values of the pre-filtering stage of the video codec.
7. The method according to claim 5, characterized in that, The determination of the first limiting range and the second limiting range for the reconstruction includes: The motion compensation tool of the video codec is determined to use the information of the second limiting range, and the second limiting range is applied to the output value of the pre-filtering stage of the video codec.
8. The method according to any one of claims 1 to 7, characterized in that, The target signal for which the amplitude limiting operation is applied has a dynamic range, which is modified based on a function, and the method includes at least one of the following: The limiting operation with the stated limiting range is applied to the target signal without modifying the limiting range according to the function; The limiting operation with a default limiting range is applied to the target signal; or The limiting operation with the modified limiting range is applied to the target signal.
9. The method according to claim 8, characterized in that, The function includes at least one of linear functions and nonlinear functions.
10. The method according to any one of claims 1 to 9, characterized in that, Also includes: The video codec is determined to use an internal dynamic range corresponding to a first bit depth, which is higher than the second bit depth used by the input signal of the video codec. The limiting range of the internal signal to be applied to the video codec is reduced to a scaled limiting range with the second bit depth; as well as The scaled clipping range having the second bit depth is included in the bitstream.
11. The method according to any one of claims 1 to 10, characterized in that, Also includes: At least one unsigned value is included in the bitstream to indicate the clipping range, the unsigned value indicating an offset relative to one of the minimum and maximum values of the predefined clipping range.
12. The method according to any one of claims 1 to 11, characterized in that: The applications include: The limiting range is applied to the sample values to generate the limited sample values, which are the final values from the processing stage of the video codec, while intermediate values within the processing stage are not limited.
13. A method for video decoding, characterized in that, include: Receive an encoded video stream, the encoded video stream including encoded information of sample values of at least one image; At least one limiting range is determined based on at least one syntax element in the encoded video bitstream, the limiting range being different from a fixed limiting range, the fixed limiting range being based on the bit depth of the video codec used to encode / decode the encoded information; The configuration of the encoding / decoding tools in the video codec is determined based on the stated limiting range; as well as The at least one image is reconstructed based on the video codec, which has the encoding / decoding tools configured according to the configuration.
14. The method according to claim 13, characterized in that, During the encoding of the sample values, the amplitude limiting range is applied to the input values in the pre-filtering stage, and determining the configuration of the encoding / decoding tool includes: The configuration of the loop filter used for reconstruction in the video codec is determined based on the information about the limiting range.
15. A method for processing visual media data, characterized in that, The method includes: According to the format rules, process the bitstream of visual media data, where: The bitstream includes encoded information of sample values from at least one image; and The format rules stipulate that: At least one limiting range is determined based on at least one syntax element in the bitstream, the limiting range being different from a fixed limiting range, the fixed limiting range being based on the bit depth of a video codec used to process the sample values of the at least one image; The configuration of the encoding / decoding tools in the video codec is determined based on the stated limiting range; and The at least one image is reconstructed based on the video codec, which has the encoding / decoding tools configured according to the configuration.