Bit depth shift control method and apparatus
By introducing bit depth signaling information based on machine target characteristics, the lack of bit depth shift control in video encoding and decoding is solved, data processing capabilities are improved, and the effective execution of machine tasks is supported.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2026-01-19
- Publication Date
- 2026-07-24
AI Technical Summary
The existing video encoding and decoding process lacks a bit depth conversion control mechanism based on machine target characteristics, resulting in a lack of explicit signaling and control for bit depth shifting.
Bit depth signaling information based on machine target features is introduced. By receiving video bitstreams and performing bit depth shifting processing, a transformed video sequence is generated, and a video bitstream containing bit depth signaling information is generated during the encoding process.
It enhances the bit depth control capability during video encoding and decoding, and supports the data processing capabilities of downstream machine tasks.
Smart Images

Figure CN122457779A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 748,406, filed January 22, 2025, and U.S. Application No. 19 / 435,911, filed December 30, 2025, the disclosures of which are incorporated herein by reference in their entirety. Technical Field
[0002] This application relates to video encoding and decoding technology, and more particularly to a bit depth shift control method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0003] Video compression can be used not only by humans but also by machines. Recently, an activity was established under ISO / IEC JTC 1SC 29 WG 4 to standardize Video Coding for Machines (VCM). The VCM reference model uses a traditional video codec as its core codec.
[0004] Controlling the bit depth shifting process during video encoding and decoding is a critical issue. Existing methods shift the bit depth of the decoded sequence left or right at the decoder based on predefined values and / or signal notifications, lacking explicit signaling and control mechanisms based on machine target characteristics. Summary of the Invention
[0005] According to one aspect of this disclosure, a bit depth shift control method is provided, the method comprising: receiving a video bit stream including an encoded video sequence and bit depth signaling information; decoding the encoded video sequence to generate a decoded video sequence; and performing bit depth shift processing on the decoded video sequence based on the bit depth signaling information.
[0006] According to one aspect of this disclosure, a bit depth shift control method is provided, the method comprising: receiving a video sequence; performing bit depth shift processing on the video sequence to generate a transformed video sequence; encoding the transformed video sequence to generate an encoded video sequence; and generating a video bit stream including the encoded video sequence and bit depth signaling information corresponding to the bit depth shift processing.
[0007] According to one aspect of this disclosure, a bit depth shift control apparatus is provided. The apparatus includes: a receiving module for receiving a video bitstream including an encoded video sequence and bit depth signaling information; a decoding module for decoding the encoded video sequence to generate a decoded video sequence; and a shifting module for performing bit depth shifting processing on the decoded video sequence based on the bit depth signaling information.
[0008] According to one aspect of this disclosure, a bit depth shift control apparatus is provided. The apparatus includes: a receiving module for receiving a video sequence; a shifting module for performing bit depth shift processing on the video sequence to generate a transformed video sequence; an encoding module for encoding the transformed video sequence to generate an encoded video sequence; and a generating module for generating a video bitstream including the encoded video sequence and bit depth signaling information corresponding to the bit depth shift processing.
[0009] According to one aspect of this disclosure, an electronic device is provided, including a memory and a processor, the memory storing computer-readable instructions which, when executed by the processor, are used to implement any of the aforementioned methods.
[0010] According to one aspect of this disclosure, a computer-readable storage medium is provided storing computer-readable instructions that, when executed by a processor, are used to implement any of the aforementioned methods.
[0011] According to one aspect of this disclosure, a method for storing a video bitstream is provided, the method comprising: generating a video bitstream according to the method described above; and storing the video bitstream.
[0012] In summary, the embodiments of this application provide a bit depth shift control method and apparatus, which introduce bit depth signaling information based on machine target characteristics to realize bit depth shift control in the video encoding and decoding process, thereby improving the data support capability for downstream machine tasks. Attached Figure Description
[0013] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:
[0014] Figure 1 This is a schematic diagram of a communication system according to an embodiment of the present disclosure.
[0015] Figure 2 This is a schematic diagram of a streaming system according to an embodiment of the present disclosure.
[0016] Figure 3 This is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.
[0017] Figure 4 This is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.
[0018] Figure 5 This is a flowchart of an example process for performing bit depth shift according to an implementation method.
[0019] Figure 6 An example computer system according to an embodiment of this disclosure is shown. Detailed Implementation
[0020] The following detailed description of the exemplary embodiments is with reference to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.
[0021] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the implementation to the exact forms disclosed. Modifications and variations can be made based on the foregoing disclosure, or modifications and variations can be derived from practice of the implementation. Furthermore, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, in the flowcharts and descriptions of operations provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed (at least partially) simultaneously, and the order of one or more operations may be switched.
[0022] It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation method. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code—it should be understood that software and hardware can be designed to implement the systems and / or methods based on the descriptions herein.
[0023] Even if specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. While each dependent claim listed below may directly refer to only one claim, the disclosure of possible implementations includes every dependent claim combined with every other claim in the claim set.
[0024] Unless explicitly stated otherwise, no element, action, or instruction used herein should be construed as critical or necessary. Furthermore, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” The term “an” or similar language is used when referring to only one item. Furthermore, as used herein, the terms “have,” “possess,” “contain,” “include,” “comprise,” etc., are intended to be open-ended terms. Furthermore, unless explicitly stated otherwise, the phrase “based on” is intended to mean “at least partially based on.” Furthermore, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.
[0025] Throughout this specification, references to "one embodiment," "implementation," or similar language mean that a particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of this solution. Therefore, throughout this specification, the phrases "in one embodiment," "in an embodiment," and similar language may, but not necessarily all, refer to the same embodiment.
[0026] Furthermore, the features, advantages, and characteristics described in this disclosure may be combined in any suitable manner in one or more embodiments. Based on the description herein, those skilled in the art will recognize that this disclosure may be practiced without one or more specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments of this disclosure.
[0027] Implementations of this disclosure relate to a syntax signaling method for bit-depth truncation of video compression for machine-task video codecs.
[0028] Reference Figures 1 to 2 This describes one or more embodiments of the encoding and decoding structures for implementing the present disclosure.
[0029] Figure 1A simplified block diagram of a communication system (100) according to an embodiment of the present disclosure is shown. The system (100) may include at least two terminals (110) and (120) interconnected via a network (150). For unidirectional data transmission, the first terminal (110) may encode video data, which may include grid data, at a local location for transmission to the other terminal (120) via the network (150). The second terminal (120) may receive the encoded video data from the other terminal from the network (150), decode the encoded data, and display the recovered video data. Unidirectional data transmission can be common in media service applications, etc.
[0030] Figure 1 A second pair of terminals (130), (140) is shown, provided to support bidirectional transmission of encoded video that may occur, for example, during a video conference. For bidirectional data transmission, each terminal (130), (140) can encode video data captured at a local location for transmission to the other terminal via a network (150). Each terminal (130), (140) can also receive encoded video data transmitted by the other terminal, can decode the encoded data, and can display the recovered video data on a local display device.
[0031] exist Figure 1 In this context, terminals (110) to (140) can be, for example, servers, personal computers, and smartphones, and / or any other type of terminal. For example, terminals (110-140) can be laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. A network (150) refers to any number of networks that transmit encoded video data among terminals (110) to (140), including, for example, wired and / or wireless communication networks. Communication networks (150) can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this discussion, unless otherwise stated below, the architecture and topology of the network (150) may be irrelevant to the operation of this disclosure.
[0032] As an example of the application to the disclosed topic. Figure 2 The placement of a video encoder and decoder in a streaming environment is illustrated. The disclosed subject matter can be used with other video-enabled applications, including, for example, video conferencing, digital TV, and storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0033] like Figure 2As shown, the streaming system (200) may include a capture subsystem (213) comprising a video source (201) and an encoder (203). The streaming system (200) may also include at least one streaming server (205) and / or at least one streaming client (206).
[0034] A video source (201) can create a stream (202) including, for example, a 3D mesh and metadata associated with the 3D mesh. The video source (201) may include, for example, a 3D sensor (e.g., a depth sensor) or a 3D imaging technique (e.g., a digital camera device), and a computing device configured to generate the 3D mesh using data received from the 3D sensor or the 3D imaging technique. A sample stream (202) that may have a high data volume compared to an encoded video bitstream can be processed by an encoder (203) coupled to the video source (201). The encoder (203) may include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoder (203) may also generate an encoded video bitstream (204). An encoded video bitstream (204) that may have a low data volume compared to an uncompressed stream (202) can be stored on a streaming server (205) for future use. One or more streaming clients (206) and streaming clients (207) can access the streaming server (205) to retrieve video bitstreams (208) and (209) respectively, which may be copies of the encoded video bitstream (204).
[0035] The video bitstream (209) is an incoming copy of the encoded video bitstream (204), and an outgoing video sample stream (211) is created, which can be rendered on a display (212) or another rendering device (not depicted). In some streaming systems, the video bitstream (204), video bitstream (208), and video bitstream (209) can be encoded according to a specific video encoding / compression standard.
[0036] Figure 3 An example functional block diagram of a video decoder (210) attached to a display (212) according to an embodiment of the present disclosure is shown.
[0037] The video decoder (210) may include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra-frame prediction unit (352), a motion compensation prediction unit (353), an aggregator (355), a loop filter unit (356), a reference image memory (357), and a current image memory (358). In at least one embodiment, the video decoder (210) may include an integrated circuit, a series of integrated circuits, and / or other electronic circuit systems. The video decoder (210) may also be implemented, in part or in whole, as software running on one or more CPUs with associated memory.
[0038] In this and other embodiments, the receiver (310) may receive one or more encoded video sequences to be decoded by the decoder (210), receiving one encoded video sequence at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences may be received from a channel (312), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive encoded video data as well as other data (e.g., encoded audio data and / or auxiliary data streams), which may be forwarded to their respective user entities (not depicted). The receiver (310) may separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter referred to as the "parser"). The buffer (315) may be unused or small when the receiver (310) receives data from a storage / forwarding device with sufficient bandwidth and controllability or from an isochronous synchronization network. For use on best-effort packet networks such as the Internet, a buffer (315) may be required. The buffer (315) can be relatively large and can have an adaptive size.
[0039] The video decoder (210) may include a parser (320) to reconstruct symbols (321) from the entropy-coded video sequence. Categories of these symbols include, for example, information for managing the operation of the decoder (210); and information potentially for controlling a rendering device (e.g., a display (212)). Figure 2The diagram shows a device that can be coupled to a decoder. Control information for the rendering device can be, for example, in the form of Supplemental Enhancement Information (SEI) messages or Video Availability Information (VUI) parameter set fragments (not depicted). The parser (320) can perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow principles known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (320) can extract a subgroup parameter set from the encoded video sequence for at least one subgroup of pixels in the subgroups used in the video decoder, based on at least one parameter corresponding to a group. Subgroups can include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The parser (320) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc., from the encoded video sequence.
[0040] The parser (320) can perform entropy decoding / parsing operations on the video sequence received from the buffer (315) to create symbols (321).
[0041] Depending on the type of encoded video picture or part of encoded video picture (e.g., inter-frame picture and intra-frame picture, inter-frame block and intra-frame block) and other factors, the reconstruction of the symbol (321) may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (320). For clarity, the flow of such subgroup control information between the parser (320) and the following multiple units is not described.
[0042] In addition to the functional blocks already mentioned, the decoder (210) can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the following functional units.
[0043] A unit may be a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive quantization transform coefficients as symbols (321) and control information (including which transform to use, block size, quantization factor, quantization scaling matrix, etc.) from the parser (320). The scaler / inverse transform unit (351) may output a block containing sample values, which may be input to the aggregator (355).
[0044] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed images, but can use predictive information from previously reconstructed portions of the current image. Such predictive information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses surrounding reconstructed information obtained from the current (partially reconstructed) image from the current image memory (358) to generate blocks of the same size and shape as the blocks in the reconstruction. In some cases, the aggregator (355) adds the predictive information already generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.
[0045] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to inter-frame coded blocks and may be motion-compensated. In this case, the motion compensation prediction unit (353) can access the reference image memory (357) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (321) belonging to the block, these samples can be added to the output of the scaler / inverse transform unit (351) by the aggregator (355) (referred to in this case as residual samples or residual signals) to generate output sample information. The address in the reference image memory (357) can be controlled by motion vectors, from which the motion compensation prediction unit (353) obtains the predicted samples. The motion vectors can be provided to the motion compensation prediction unit (353) in the form of symbols (321), which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values obtained from the reference image memory (357) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0046] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356). The video compression technique may include in-loop filtering techniques controlled by parameters included in the encoded video bitstream as symbols (321) from the parser (320) that can be used in the loop filter unit (356), but may also be responsive to metadata obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0047] The output of the loop filter unit (356) can be a sample stream, which can be output to a rendering device such as a display (212) and stored in a reference image memory (357) for future inter-frame image prediction.
[0048] Once a certain coded image is fully reconstructed, it can be used as a reference image for future predictions. Once a coded image is fully reconstructed and has been identified as a reference image (by, for example, a parser (320)), the current reference image can become part of a reference image memory (357), and a new current image memory can be reallocated before the reconstruction of subsequent coded images begins.
[0049] The video decoder (210) can perform decoding operations according to predefined video compression techniques that can be documented in standards such as ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard (as specified in the video compression technique document or standard, and particularly in the brief document therein). Furthermore, the complexity of the encoded video sequence can be within limits defined by the hierarchy of the video compression technique or standard in order to conform to some video compression techniques or standards. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata signaled in the encoded video sequence for HRD buffer management.
[0050] In this implementation, the receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (210) to correctly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0051] Figure 4 An example functional block diagram of a video encoder (203) associated with a video source (201) according to an embodiment of the present disclosure is shown.
[0052] The video encoder (203) may include, for example, an encoder as a source encoder (430), an encoding engine (432), a (local) decoder (433), a reference image memory (434), a predictor (435), a transmitter (440), an entropy encoder (445), a controller (450), and a channel (460).
[0053] The encoder (203) can receive video samples from a video source (201) (which is not part of the encoder), which can capture video images to be encoded by the encoder (203).
[0054] The video source (201) can be provided as a digital video sample stream of a sequence of source videos to be encoded by the encoder (203). This digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (201) can be a storage device storing previously prepared video. In a video conferencing system, the video source (203) can be a camera device capturing local image information as a video sequence. The video data can be provided as multiple individual pictures that are given motion when viewed in sequence. The pictures themselves can be organized as a spatial array of pixels, wherein, depending on the sampling structure, color space, etc., used, each pixel can include one or more samples. Those skilled in the art can readily understand the relationship between pixels and samples. The following description focuses on samples.
[0055] According to the implementation, the encoder (203) can encode and compress images of the source video sequence into an encoded video sequence (443) in real time or according to any other time constraints required by the application. Implementing an appropriate encoding / decoding speed is a function of the controller (450). The controller (450) can also control other functional units as described below and can be functionally coupled to these units. For simplicity, the coupling is not depicted. The parameters set by the controller (450) may include rate control related parameters (image skipping, quantizer, λ value of rate distortion optimization techniques, etc.), image size, group of images (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller (450) that may belong to a video encoder (203) optimized for a specific system design.
[0056] Some video encoders operate in a manner readily recognized by those skilled in the art as an "encoder-decoder loop." As an oversimplification, the encoder-decoder loop may consist of the encoding portion of a source encoder (430) responsible for creating symbols based on the input image to be encoded and a reference image, and a (local) decoder (433) embedded in the encoder (203) that reconstructs the symbols to create sample data, which is also created by a (remote) decoder when compression between the symbols and the encoded video bitstream is lossless in some video compression techniques. This reconstructed sample stream can be fed into a reference image memory (434). Since decoding the symbol stream results in bit-accurate results independent of the decoder's location (local or remote), the contents of the reference image memory are also bit-accurate between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction portion are exactly the same sample values that the decoder would "see" when using prediction during decoding. The basic principles of this reference image synchronization (and the drift that results in, for example, failure to maintain synchronization due to channel errors) are known to those skilled in the art.
[0057] The operation of the "local" decoder (433) can be combined with the one already mentioned above. Figure 3 The operation of the “remote” decoder (210) described in detail is the same. However, since the symbols are available and can be losslessly encoded / decoded into a encoded video sequence by the entropy encoder (445) and the parser (320), the entropy decoding part of the decoder (210) (including the channel (312), receiver (310), buffer (315) and parser (320)) does not need to be fully implemented in the local decoder (433).
[0058] It can be observed that any decoder technique other than parsing / entropy decoding, which exists in the decoder, may need to exist in the corresponding encoder in essentially the same functional form. For this reason, the subject matter presented focuses on decoder operation. The description of encoder techniques can be simplified, as they may be the opposite of the fully described decoder techniques. More detailed descriptions are provided below only where necessary in certain places.
[0059] As part of the operation of the source encoder (430), the source encoder may perform motion-compensated predictive coding, which references one or more previously encoded frames in the video sequence designated as reference frames to predictively code the input frame. In this way, the coding engine (432) encodes the differences between pixel blocks of the input frame and pixel blocks of the reference frame, which may be selected as the predictive reference for the input frame.
[0060] The local video decoder (433) can decode encoded video data of frames that can be designated as reference frames based on symbols created by the source encoder (430). The operation of the encoding engine (432) can advantageously be a lossy process. When encoded video data can be decoded by the video decoder (433), Figure 4 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (433) replicates the decoding process performed on the reference frame by the video decoder and can store the reconstructed reference frame in the reference image memory (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame that has the same content as the reconstructed reference frame that will be obtained by the remote video decoder (without transmission errors).
[0061] The predictor (435) can perform a prediction search for the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) can search in the reference image memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new image. The predictor (435) can operate pixel-by-pixel based on the sample blocks to find appropriate prediction references. In some cases, such as as determined by the search results obtained by the predictor (435), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (434).
[0062] The controller (450) can manage the encoding operations of the video encoder (430), which include, for example, setting parameters and subgroup parameters for encoding video data.
[0063] The outputs of all the aforementioned functional units can be entropy encoded in the entropy encoder (445). The entropy encoder converts the symbols generated by the various functional units into a encoded video sequence by performing lossless compression on the symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable-length coding, arithmetic coding, etc.).
[0064] The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) in preparation for transmission via a communication channel (460), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (440) can combine the encoded video data from the video encoder (430) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).
[0065] The controller (450) can manage the operation of the encoder (203). During encoding, the controller (450) can assign a specific encoding picture type to each encoded picture, which may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures can typically be assigned as intra-frame pictures (I-pictures), predictive pictures (P-pictures), or bidirectional predictive pictures (B-pictures).
[0066] An intra-frame picture (I-picture) can be a picture that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will understand those variations of I-pictures and their corresponding applications and characteristics.
[0067] Predictive images (P-images) can be images that can be encoded and decoded using intra-frame or inter-frame prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0068] Bidirectional predictive images (B-images) can be encoded and decoded using intra-frame or inter-frame prediction that uses at most two motion vectors and reference indices to predict sample values for each block. Similarly, multi-predictive images can use more than two reference images and associated metadata to reconstruct a single block.
[0069] Source images are typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples) and encoded on a block-by-block basis. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the coding assignment of the corresponding images applied to the blocks. For example, blocks of an I-image can be unpredictably encoded, or blocks of an I-image can be predictively encoded (spatial prediction or intra-frame prediction) with reference to already encoded blocks of the same image. Pixel blocks of a P-image can be encoded unpredictably, via spatial prediction, or via temporal prediction (with reference to a previously encoded reference image). Blocks of a B-image can be encoded unpredictably, via spatial prediction, or via temporal prediction (with reference to one or two previously encoded reference images).
[0070] The video encoder (203) can perform encoding operations according to a predetermined video coding technique or standard, such as that specified in ITU-T H.265 Recommendation. In the operation of the video encoder (203), various compression operations can be performed, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard in use.
[0071] In this implementation, the transmitter (440) can transmit additional data as well as encoded video. The video encoder (430) can include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant images and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.
[0072] The proposed methods can be used individually or in any combination in any order. Furthermore, each of the method (or implementation), encoder, and decoder can be implemented using a processing circuitry system (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.
[0073] In existing methods, the bit depth of the decoded sequence is shifted left or right at the decoder based on predefined values and / or a syntax signaled by a signal. In one or more examples, the bit depth shift can be performed during the quantization or transform process at the encoder or during the dequantization or inverse transform process at the decoder.
[0074] The following table shows an example bit depth truncation syntax.
[0075]
[0076] Table 1
[0077] Some tools / suggestions require additional syntax to signal the number of bits shifted at the luminance channel when bit_depth_shift_flag=0.
[0078] In one or more examples, the syntax for bit depth can be implemented as follows.
[0079]
[0080] Table 2
[0081] Since bit_depth_shift_luma and bit_depth_shift_chroma are always available, other tools can reuse them instead of creating new syntax.
[0082] According to one or more implementations, the number of bits shifted at the decoder is signaled in the bitstream. In one or more examples, a syntax value of 0 can indicate that no shift is performed. In one example, the number of bits shifted at the decoder is signaled as a syntax for all color components.
[0083]
[0084] Table 3
[0085] In one or more examples, the number of bits to be shifted at the decoder is signaled for the luminance and chrominance components, respectively.
[0086]
[0087] Table 4
[0088] In one or more examples, the number of bits shifted at the decoder is signaled for the Y, U, and V components, respectively.
[0089]
[0090] Table 5
[0091] In one or more examples, a bit shift enable flag is signaled in the bitstream. A flag equal to 1 indicates that all color components are shifted by a predefined number of bits at the decoder; otherwise, no bit shifting is performed at the decoder.
[0092]
[0093] Table 6
[0094] In one or more examples, bit shift enable flags are signaled in the bitstream for both luminance and chrominance components. A flag of 1 indicates that a predefined number of bits should be shifted at the decoder for each of the luminance and chrominance components.
[0095]
[0096] Table 7
[0097] In one or more examples, bit shift enable flags are signaled for the Y, U, and V components in the bitstream, respectively. A flag of 1 indicates that a predefined number of bits should be shifted for each of the Y, U, and V components at the decoder.
[0098]
[0099] Table 8
[0100] According to one or more implementations, a bit shift enable flag and the number of bits are signaled in the bit stream. When the bit shift enable flag is equal to 1, the decoder performs the shift based on a predefined value and / or a value inferred from the number of bits signaled.
[0101]
[0102] Table 9
[0103] In one or more examples, when the bit shift enable flag is equal to 1, the decoder can perform shifts according to predefined values. For example, a 1-bit shift for luminance and a 0-bit shift for chroma. In another example, all color components are shifted by 1 bit.
[0104] In another example, when the bit shift enable flag is equal to 1, the decoder can perform a shift according to the value obtained from the number of bits indicated by the signal. For example, when the bit shift enable flag is equal to 1, the decoder shifts the luminance according to the number of bits indicated by the signal for luminance, and shifts the chrominance according to the number of bits indicated by the signal for chrominance.
[0105] In one or more examples, shifts based on predefined values and shifts based on values obtained from the number of bits signaled can be combined. For example, when the bit shift enable flag is equal to 1, the decoder shifts the luminance based on the number of bits signaled for luminance and shifts the chrominance based on predefined values.
[0106] According to one or more implementations, specific tools may be executed differently depending on the syntax signaled in bit_depth_shift(). For example, a specific tool (e.g., histogram equalization for enhancing the luminance component) may be executed only when the bit shift enable flag is equal to 1. In one example, bit_depth_luma_enhance() may be executed only when the bit shift enable flag is equal to 1.
[0107]
[0108] Table 10
[0109] In one or more examples, a specific tool is executed only when the bit shift enable flag is equal to 0.
[0110]
[0111] Table 11
[0112] In one or more examples, a specific tool is executed only if the number of bits signaled by the signal meets a predefined constraint. In one example, the specific tool is executed only if the number of bits signaled for brightness is 1.
[0113]
[0114] Table 12
[0115] In one or more examples, a particular tool is executed only if the number of bits used for chroma signaling is 1.
[0116]
[0117] Table 13
[0118] In one or more examples, a specific tool may reuse the number of bits signaled in bit_depth_shift(). In one or more examples, a specific tool may signal syntax based on the syntax signaled in bit_depth_shift(). In one example, additional syntax may be signaled in colorization() only if bit_depth_shift_flag is equal to 1.
[0119]
[0120] Table 14
[0121] In one or more examples, the additional syntax is signaled in colorization() only when bit_depth_shift_flag is equal to 1.
[0122]
[0123] Table 15
[0124] In one or more examples, the additional syntax is signaled in `colorization()` only if the number of bits signaled meets a predefined condition. In one example, the additional syntax is signaled in `colorization()` only if `bit_depth_shift_luma` is not equal to 1.
[0125]
[0126] Table 16
[0127] According to one or more implementations, byte_alignment() can also be performed at the end of the bit_depth_shift() function.
[0128] Figure 5 A flowchart of an example process 500 for performing bit depth shifting is shown. Process 500 can be executed by video decoder 210.
[0129] The process can begin at operation S502, where a video bitstream including an encoded video sequence and bit depth signaling information is received. The process proceeds to operation S504, where the encoded video sequence is decoded to generate a decoded video sequence. The process proceeds to operation S506, where a bit depth shift is performed on the decoded video sequence using the bit depth signaling information. The bit depth shift can be performed according to any of the embodiments discussed above. For example, as part of the bit depth shift, bit depth signaling information is extracted from the video bitstream, where it is determined whether to perform bit depth shift based on the bit depth signaling information. For example, as discussed above, bit depth shift can be performed on or not on the decoded video sequence based on the value of a bit depth enable flag. As discussed above, the bit depth signaling information can indicate the number of bits shifted in the decoded video sequence and which color component is bit depth shifted.
[0130] The techniques described above can be implemented as computer software that uses computer-readable instructions and is physically stored on one or more computer-readable media. For example, Figure 6 A computer system 600 suitable for implementing certain embodiments of the present disclosure is shown.
[0131] Computer software can be encoded using any suitable machine code or computer language, which can be subjected to mechanisms such as assembly, compilation, and linking to create code containing instructions that can be executed directly by a computer's central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, etc.
[0132] The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0133] Figure 6 The components shown for computer system 600 are examples and are not intended to impose any limitation on the scope of use or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any one or a combination of components shown in a non-limiting embodiment of computer system 600.
[0134] Computer system 600 may include specific human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users via, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., speech, tapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface device may also be used to capture specific media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image capturing devices), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0135] The input human-machine interface device may include one or more of the following (only one of each is depicted): keyboard 601, mouse 602, touchpad 603, touch screen 610, data glove, joystick 605, microphone 606, scanner 607, and camera device 608.
[0136] Computer system 600 may also include specific human-machine interface (HMI) output devices. Such HMI output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. These HMI output devices may include tactile output devices (e.g., tactile feedback via touchscreen 610, data gloves, or joystick 605, but tactile feedback devices that are not used as input devices may also exist). For example, such devices may be audio output devices (e.g., speakers 609, headphones (not depicted)); visual output devices (e.g., screens 610, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability—some of which may be capable of outputting two-dimensional or more than three-dimensional visual output in a manner such as stereoscopic output; virtual reality glasses (not depicted); holographic displays and smoke generators (not depicted)); and printers (not depicted).
[0137] Computer system 600 may also include human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW 620 and CD / DVD or similar media 621, thumb drives 622, removable hard disk drives or solid-state drives 623, conventional magnetic media (such as magnetic tape and floppy disks) (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0138] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0139] Computer system 600 may also include interfaces to one or more communication networks. These networks can be wireless, wired, or optical. They can also be local area, wide area, metropolitan area, vehicular, industrial, real-time, latency-tolerant, etc. Examples of networks include: local area networks (LANs) such as Ethernet and wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; wired or wireless wide area digital TV networks, including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks including CAN buses, etc. Some networks typically require external network interface adapters attached to specific general-purpose data ports or peripheral buses 649 (e.g., USB ports of computer system 600); other networks are typically integrated into the core of computer system 600 by attaching to the system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, computer system 600 can communicate with other entities. Such communication can be one-way (receive-only, e.g., broadcasting TV), one-way (transmit-only, e.g., a CAN bus to a specific CAN bus device), or bidirectional (e.g., using a local or wide area digital network to other computer systems). Such communication can include communication to a cloud computing environment 655. Specific protocols and protocol stacks can be used on each of the networks and network interfaces described above.
[0140] The aforementioned human-machine interface device, human-accessible storage device, and network interface 654 can be attached to the core 640 of the computer system 600.
[0141] Core 640 may include one or more central processing units (CPUs) 641, graphics processing units (GPUs) 642, dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) 643, task-specific hardware accelerators 644, etc. These devices, along with read-only memory (ROM) 645, random access memory 646, and internal mass storage devices (such as internal non-user-accessible hard disk drives, SSDs, etc.) 647, can be connected via system bus 648. In some computer systems, system bus 648 may be accessed via one or more physical connectors to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus 648 or via peripheral bus 649. Peripheral bus architectures include PCI, USB, etc. A graphics adapter 650 may be included in core 640.
[0142] The CPU 641, GPU 642, FPGA 643, and accelerator 644 can execute specific instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM 645 or RAM 646. Transient data can also be stored in RAM 646, while permanent data can be stored, for example, in an internal mass storage device 647. Fast storage and retrieval of any memory device within the memory device can be achieved using a cache memory, which can be closely associated with one or more CPUs 641, GPUs 642, mass storage devices 647, ROM 645, RAM 646, etc.
[0143] Computer-readable media may have computer code thereon for performing operations of various computer implementations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.
[0144] By way of example and not limitation, a computer system 600 having an architecture, and particularly a core 640, can provide functionality produced by a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with a user-accessible mass storage device as described above, as well as certain storage devices of the core 640 having non-transitory characteristics (e.g., internal mass storage device 647 or ROM 645). Software implementing various embodiments of this disclosure can be stored in such a device and executed by the core 640. Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the core 640, and particularly its processors (including CPU, GPU, FPGA, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 646 and modifying such data structures according to the software-defined processes. Alternatively or as an alternative, the computer system may provide functionality generated by logic hardwired or otherwise embodied in circuitry (e.g., accelerator 644), which may operate in place of or with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, reference to software may encompass logic, and conversely, reference to logic may encompass software. Where appropriate, reference to computer-readable medium may encompass circuitry storing software for execution (e.g., integrated circuit (IC)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.
[0145] The process described above can be implemented in an image and / or video decoding process or an image and / or video encoding process. The decoding / encoding process can be used in a video decoder device. Alternatively, the decoding / encoding process can also be used in a video encoder device. In some embodiments, the process is executed by a processing circuitry system, such as a processing circuitry system that performs the functions of a video decoder, etc. In other embodiments, the process is executed by a processing circuitry system, such as a processing circuitry system that performs the functions of a video encoder, etc. In some embodiments, the process is implemented as software instructions; therefore, the processing circuitry system executes the process when it executes the software instructions. In other embodiments, the process can be implemented as a hardware process on a chip; therefore, the processing circuitry system executes the process when it executes the hardware instructions. The process can be appropriately adjusted. Steps in the process described above can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.
[0146] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, a computer system can be adapted to implement certain embodiments of the disclosed subject matter. The computer software can be encoded using any suitable machine code or computer language, which can be subjected to mechanisms such as assembly, compilation, and linking to create code comprising instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc. The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc. The components used in the computer system are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing embodiments of this disclosure. The configuration of the components should not be construed as having any dependency or requirement relating to any component or combination of components shown in the exemplary embodiments of the computer system. Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.
[0147] The use of “at least one of…” or “one of…” in this disclosure is intended to include any one or a combination of the elements described. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A to C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include either A or B or (A and B). The use of “one of…” does not exclude any combination of the elements where applicable, such as when the elements are not mutually exclusive.
[0148] While this disclosure has described several non-limiting embodiments, there are variations, substitutions, and various alternative equivalents that fall within the scope of this disclosure. Therefore, it will be appreciated that those skilled in the art will be able to conceive of many systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are therefore within its spirit and scope.
[0149] The above disclosure also covers the following implementation methods:
[0150] (1) A method executed by at least one processor of a video decoder, the method comprising: receiving a video bitstream including an encoded video sequence and bit depth signaling information; decoding the encoded video sequence to generate a decoded video sequence; and performing bit depth shifting processing on the decoded video sequence based on the bit depth signaling information.
[0151] (2) The method according to feature (1), wherein the bit depth signaling information includes a parameter indicating the number of bits shifted in the decoded video sequence.
[0152] (3) The method according to feature (2), wherein the parameter indicating the number of bits shifted in the decoded video sequence is applied to each color component.
[0153] (4) The method according to feature (2), wherein the parameter is a first parameter indicating the number of bits shifted in the luminance component of the decoded video sequence, and wherein the bit depth signaling information further includes a second parameter indicating the number of bits shifted in the chrominance component of the decoded video sequence.
[0154] (5) The method according to feature (2), wherein the parameter is a first parameter indicating the number of bits shifted by the Y color component of the decoded video sequence, wherein the bit depth signaling information further includes a second parameter indicating the number of bits shifted by the U color component of the decoded video sequence, and wherein the bit depth signaling information further includes a third parameter indicating the number of bits shifted by the V color component of the decoded video sequence.
[0155] (6) The method according to any one of features (1) to (5), wherein the bit depth signaling information further includes a bit depth shift enable flag, the bit depth shift enable flag having a first value indicating the number of bits shifted in the decoded video sequence and a second value indicating that no bit shift is performed at the decoder.
[0156] (7) The method according to any one of features (1) to (6), wherein the bit depth signaling information further includes a first bit depth shift enable flag, the first bit depth shift enable flag having a first value indicating the number of bits shifted for the luminance component of the decoded video sequence and a second value indicating that the luminance component is not bit shifted at the decoder, and wherein the bit depth signaling information further includes a second bit depth shift enable flag, the second bit depth shift enable flag having a first value indicating the number of bits shifted for the chroma component of the decoded video sequence and a second value indicating that the chroma component is not bit shifted at the decoder.
[0157] (8) The method according to any one of features (1) to (7), wherein the bit depth signaling information further includes a first bit depth shift enable flag, the first bit depth shift enable flag having a first value indicating the number of bits shifted for the Y color component of the decoded video sequence and a second value indicating that the Y color component is not bit shifted at the decoder, wherein the bit depth signaling information further includes a second bit depth shift enable flag, the second bit depth shift enable flag having a first value indicating the number of bits shifted for the U color component of the decoded video sequence and a second value indicating that the U color component is not bit shifted at the decoder, and wherein the bit depth signaling information further includes a third bit depth shift enable flag, the third bit depth shift enable flag having a first value indicating the number of bits shifted for the V color component of the decoded video sequence and a second value indicating that the V color component is not bit shifted at the decoder.
[0158] (9) The method according to feature (6), wherein the bit depth signaling information further includes a parameter indicating the number of bits shifted in the decoded video sequence when the bit depth shift enable flag is the first value.
[0159] (10) The method according to feature (6), wherein when the bit depth shift enable flag is the first value, the luminance component is shifted by 1 bit and the chrominance component is shifted by 0 bits.
[0160] (11) The method according to feature (6), wherein when the bit depth enable flag is the first value, the luminance component is shifted according to the number of bits indicated by the first parameter in the bit depth signaling information, and the chrominance component is shifted according to the number of bits indicated by the second parameter in the bit depth signaling information.
[0161] (12) The method according to feature (6), wherein when the bit depth enable flag is the first value, the luminance component is shifted according to the number of bits indicated by the parameter in the bit depth signaling information, and the chrominance component is shifted according to a predefined number of bits.
[0162] (13) The method according to feature (6), wherein the decoder runs a bit depth tool according to the value of the bit depth shift enable flag.
[0163] (14) The method according to feature (13), wherein the bit depth tool is a brightness enhancement tool.
[0164] (15) A method executed by at least one processor in an encoder, the method comprising: receiving a video sequence; performing a bit depth shifting process on the video sequence to generate a transformed video sequence; encoding the transformed video sequence to generate an encoded video sequence; and generating a video bitstream including the encoded video sequence and bit depth signaling information corresponding to the bit depth shifting process.
[0165] (16) The method according to feature (15), wherein the bit depth signaling information includes a parameter indicating the number of bits shifted in the video sequence.
[0166] (17) The method according to feature (16), wherein the parameter indicating the number of bits shifted in the video sequence is applied to each color component.
[0167] (18) The method according to feature (16), wherein the parameter is a first parameter indicating the number of bits shifted for the luminance component of the video sequence, and wherein the bit depth signaling information further includes a second parameter indicating the number of bits shifted for the chrominance component of the video sequence.
[0168] (19) The method according to feature (16), wherein the parameter is a first parameter indicating the number of bits shifted by the Y color component of the video sequence, wherein the bit depth signaling information further includes a second parameter indicating the number of bits shifted by the U color component of the video sequence, and wherein the bit depth signaling information further includes a third parameter indicating the number of bits shifted by the V color component of the video sequence.
[0169] (20) A non-transitory computer-readable medium for storing a video bitstream, the video bitstream being decoded by means of: receiving the video bitstream including an encoded video sequence and bit depth signaling information; decoding the encoded video sequence to generate a decoded video sequence; and performing bit depth shifting processing on the decoded video sequence based on the bit depth signaling information.
Claims
1. A bit depth shift control method, characterized in that, include: Receive a video bitstream that includes encoded video sequences and bit depth signaling information; The encoded video sequence is decoded to generate a decoded video sequence; as well as The decoded video sequence is subjected to bit depth shifting processing based on the bit depth signaling information.
2. The method according to claim 1, characterized in that, The bit depth signaling information includes parameters that indicate the number of bits shifted in the decoded video sequence.
3. The method according to claim 2, characterized in that, The parameters are applied to each color component.
4. The method according to claim 2, characterized in that, The parameters include a first parameter indicating the number of bits shifted for the luminance component of the decoded video sequence and a second parameter indicating the number of bits shifted for the chrominance component of the decoded video sequence.
5. The method according to claim 1, characterized in that, The bit depth signaling information also includes a bit depth shift enable flag, which has a first value indicating the number of bits shifted in the decoded video sequence and a second value indicating that bit shifting is not performed at the decoder.
6. The method according to claim 1, characterized in that, The bit depth signaling information further includes a first bit depth shift enable flag, which has a first value indicating the number of bits shifted in the luminance component of the decoded video sequence and a second value indicating that the luminance component is not bit shifted at the decoder. The bit depth signaling information also includes a second bit depth shift enable flag, which has a first value indicating the number of bits shifted in the chroma component of the decoded video sequence and a second value indicating that the chroma component is not bit shifted at the decoder.
7. The method according to claim 5, characterized in that, The bit depth signaling information also includes a parameter indicating the number of bits shifted in the decoded video sequence when the bit depth shift enable flag is the first value.
8. The method according to claim 5, characterized in that, When the bit depth enable flag is the first value, the luminance component is shifted according to the number of bits indicated by the first parameter in the bit depth signaling information, and the chrominance component is shifted according to the number of bits indicated by the second parameter in the bit depth signaling information.
9. The method according to claim 5, characterized in that, When the bit depth enable flag is the first value, the luminance component is shifted according to the number of bits indicated by the parameters in the bit depth signaling information, and the chrominance component is shifted according to the predefined number of bits.
10. The method according to claim 5, characterized in that, The decoder runs the bit depth tool based on the value of the bit depth shift enable flag.
11. The method according to claim 10, characterized in that, The bit depth tool is a brightness enhancement tool.
12. A bit depth shift control method, characterized in that, include: Receive video sequences; Perform bit depth shifting on the video sequence to generate a transformed video sequence; The transformed video sequence is encoded to generate an encoded video sequence; as well as Generate a video bitstream that includes the encoded video sequence and bit depth signaling information corresponding to the bit depth shifting process.
13. The method according to claim 12, characterized in that, The bit depth signaling information includes parameters that indicate the number of bits shifted in the video sequence.
14. The method according to claim 13, characterized in that, The parameters are applied to each color component.
15. The method according to claim 13, characterized in that, The parameters include a first parameter indicating the number of bits shifted for the luminance component of the video sequence, and a second parameter indicating the number of bits shifted for the chrominance component of the video sequence.
16. A bit depth shift control device, characterized in that, include: The receiving module is used to receive video bitstreams including encoded video sequences and bit depth signaling information; A decoding module is used to decode the encoded video sequence to generate a decoded video sequence; as well as The shift module is used to perform bit depth shifting processing on the decoded video sequence based on the bit depth signaling information.
17. A bit depth shift control device, characterized in that, include: The receiving module is used to receive video sequences; The shift module is used to perform bit-depth shift processing on the video sequence to generate a transformed video sequence; An encoding module is used to encode the transformed video sequence to generate an encoded video sequence; as well as The generation module is used to generate a video bitstream that includes the encoded video sequence and bit depth signaling information corresponding to the bit depth shifting process.
18. An electronic device comprising a memory and a processor, characterized in that, The memory stores computer-readable instructions. The computer-readable instructions, when executed by the processor, are used to implement the method according to any one of claims 1 to 11 or to implement the method according to any one of claims 12 to 15.
19. A computer-readable storage medium storing computer-readable instructions, characterized in that, The computer-readable instructions, when executed by a processor, are used to implement the method according to any one of claims 1 to 11 or the method according to any one of claims 12 to 15.
20. A method for storing video bitstreams, characterized in that, include: Generate a video bitstream by performing the method according to any one of claims 12 to 15; And to store the video bitstream.