Method, apparatus, and readable medium for decoding an encoded video bitstream
By extracting the parameters in the adaptive parameter set and combining the picture header and the video encoding layer NAL unit, adaptive decoding of the encoded pictures is achieved, which solves the problem of difficulty in managing image size and resolution changes in the prior art, and improves the efficiency and flexibility of video decoding.
Patent Information
- Application Number
- CN202080037834.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-30
- Filing Date
- 2020-10-05
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-10-05
AI Technical Summary
When existing video encoding and decoding technologies process encoded video code streams, it is difficult to effectively manage changes in image size and resolution, resulting in waste of resources and inefficient decoding.
By extracting parameters in the adaptive parameter set (APS), combining picture header (PH) and video encoding layer (VCL) NAL units, adaptive decoding of encoded pictures is realized, supporting dynamic changes in picture size and resolution.
It improves the efficiency and flexibility of video decoding, reduces resource consumption, and can more effectively process video data in different scenarios.
Smart Images

Figure CN113874875B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 954,096 filed on December 27, 2019 and U.S. Patent Application No. 17 / 038,541 filed on September 30, 2020, the entire contents of which are incorporated herein. Technical Field
[0003] The disclosed subject matter relates to video encoding and decoding, and more particularly to a method, apparatus, and storage medium for decoding an encoded video stream. Background Art
[0004] Video encoding and decoding using inter-picture prediction with motion compensation is known. Uncompressed digital video may consist of a series of pictures, each picture having a certain spatial dimension, for example, 1920×1080 luma samples and associated chroma samples. The series of pictures may have a fixed or variable picture rate (also informally referred to as frame rate), for example, 60 pictures per second or 60 Hertz (Hz). Uncompressed video has significant bit rate requirements. For example, 1080p60 4:2:0 video (1920×1080 luma sample resolution at 60 Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires more than 600 GB of storage space.
[0005] One purpose of video encoding and decoding can be to reduce redundancy in the input video signal through compression. Compression can help reduce the bandwidth or storage space requirements mentioned above, in some cases, by two or more orders of magnitude. Both lossless compression and lossy compression and combinations thereof can be used for video encoding and decoding. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal. When lossy compression is used, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal can be used for the intended application. Lossy compression is widely used in video. The amount of distortion allowed by lossy compression depends on the application; for example, users of certain consumer streaming applications can tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that the higher the allowable / tolerable distortion, the higher the compression ratio that can be produced.
[0006] There are several broad categories of techniques that video encoders and decoders can use, including, for example, motion compensation, transforms, quantization, and entropy coding. Some of the techniques in these categories are described below.
[0007] Historically, video encoders and decoders have tended to operate on a given picture size, which in most cases is defined and remains constant for a Coded Video Sequence (CVS), Group of Pictures (GOP), or similar multi-picture temporal frame. For example, in MPEG-2, the system design changes the horizontal resolution (and thus the picture size) depending on factors such as scene activity, but only within I-frame pictures, so it usually applies to GOPs. Resampling reference pictures to allow different resolutions to be used in CVS is known, for example from ITU-T Recommendation H.263 Annex P. However, since the picture size does not change here, only the reference pictures are resampled, which may result in only a portion of the picture canvas being used (in the case of downsampling) or only a portion of the scene being captured (in the case of upsampling). In addition, H.263 Annex Q allows individual macroblocks to be resampled up or down by a factor of two (in each dimension). Again, the picture size remains constant. The size of the macroblock is fixed in H.263 and therefore does not need to be signaled. Summary of the invention
[0008] In an embodiment, a method for decoding a coded video stream includes: obtaining a coded video sequence CVS from the coded video stream, the CVS including picture units corresponding to coded pictures; obtaining a picture header PH network abstraction layer NAL unit included in the picture unit; obtaining at least one video coding layer VCL NAL unit included in the picture unit; decoding the coded picture based on the PH NAL unit, the at least one VCL NAL unit and an adaptive parameter set APS to obtain a decoded picture, wherein the APS is included in an APS NAL unit obtained from the CVS, and the APS NAL unit is used for the decoding earlier than the at least one VCL NAL unit; and outputting the decoded picture.
[0009] In an embodiment, a device for decoding a coded video stream includes: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code to execute a method for decoding a coded video stream.
[0010] In an embodiment, a device for decoding a coded video code stream comprises: a first obtaining module, obtaining a coded video sequence CVS from the coded video code stream, wherein the CVS comprises a picture unit corresponding to a coded picture; a second obtaining module, obtaining a picture header PH network abstraction layer NAL unit included in the picture unit; a third obtaining module, obtaining at least one video coding layer VCL NAL unit included in the picture unit; a decoding module, decoding the coded picture based on the PH NAL unit, the at least one VCL NAL unit and an adaptive parameter set APS to obtain a decoded picture, wherein the APS is included in an APS NAL unit obtained from the CVS, and the APSNAL unit is used for the decoding earlier than the at least one VCL NAL unit; and an output module, outputting the decoded picture.
[0011] In an embodiment, a non-transitory computer-readable medium storing instructions includes: one or more instructions, which, when executed by one or more processors of an apparatus for decoding a coded video stream, cause the one or more processors to: execute a method for decoding a coded video stream. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Other features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:
[0013] Figure 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.
[0014] Figure 2 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.
[0015] Figure 3 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.
[0016] Figure 4 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.
[0017] FIG. 5A to FIG. 5E is a diagram of options for signaling ARC parameters according to an embodiment.
[0018] FIG. 6A to FIG. 6B is a diagram of an example of a syntax table according to an embodiment.
[0019] Figure 7 is an example of a prediction structure with scalability for adaptive resolution changes according to an embodiment.
[0020] Figure 8is an example of a syntax table according to an embodiment.
[0021] Fig. 9 is a schematic diagram of a simplified block diagram for parsing and decoding a POC period and an access unit count value for each access unit according to an embodiment.
[0022] Fig.10 is a schematic diagram of a video bitstream structure including multiple layers of sub-pictures according to an embodiment.
[0023] Fig.11 is a schematic diagram of display of a selected sub-picture with enhanced resolution according to an embodiment.
[0024] Fig.12 is a block diagram of decoding and display processing of a video stream including multiple layers of sub-pictures according to an embodiment.
[0025] Fig.13 is a schematic diagram of a 360 video display with an enhancement layer of sub-pictures according to an embodiment.
[0026] Fig.14 is an example of layout information of a sub-picture, a corresponding layer of the sub-picture, and a picture prediction structure according to an embodiment.
[0027] Fig.15 is an example of layout information of a sub-picture having a spatial scalability form of a local area and its corresponding layer and picture prediction structure according to an embodiment.
[0028] Figures 16A to 16B is an example of a syntax table of sub-picture layout information according to an embodiment.
[0029] Fig.17 is an example of a syntax table of a SEI message for sub-picture layout information according to an embodiment.
[0030] Fig.18 is an example of a syntax table for indicating output layers and profile / tier / level information for each output layer set, according to an embodiment.
[0031] Fig.19 is an example of a syntax table for indicating an output layer mode for each output layer set according to an embodiment.
[0032] Fig. 20 is an example of a syntax table for indicating a current sub-picture of each layer of each output layer set according to an embodiment.
[0033] Fig.21 is a schematic diagram of code stream consistency requirements according to an embodiment.
[0034] Fig. 22is a flowchart of an example method for decoding an encoded video stream according to an embodiment.
[0035] Fig.23 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION
[0036] Figure 1 The present invention is a simplified block diagram of a communication system (100) according to an embodiment disclosed in the present application. The communication system (100) includes at least two terminal devices (110-120), and the at least two terminal devices can be interconnected with each other through a network (150). For unidirectional data transmission, a first terminal device (110) can encode video data at a local location for transmission to another terminal device (120) through the network (150). The second terminal device (120) can receive the encoded video data of another terminal device from the network (150), decode the encoded video data and display the recovered video data. Unidirectional data transmission is more common in applications such as media services.
[0037] Figure 1 The present invention illustrates a third terminal device (130) and a fourth terminal device (140) for supporting bidirectional transmission of encoded video, which may occur, for example, during a video conference. For bidirectional data transmission, each of the third terminal device (130) and the fourth terminal device (140) may encode video data collected at a local location for transmission to the other of the third terminal device (130) and the fourth terminal device (140) via a network (150). Each of the third terminal device (130) and the fourth terminal device (140) may also receive the encoded video data transmitted by the other of the third terminal device (130) and the fourth terminal device (140), decode the encoded video data, and display the restored video data on a local display device.
[0038] exist Figure 1In an embodiment of the present invention, the first terminal device (110), the second terminal device (120), the third terminal device (130) and the fourth terminal device (140) may be servers, personal computers and smart phones, but the principles disclosed in the present application may not be limited thereto. The embodiments disclosed in the present application are applicable to laptop computers, tablet computers, media players and / or dedicated video conferencing equipment. The network (150) represents any number of networks that transmit encoded video data between the first terminal device (110), the second terminal device (120), the third terminal device (130) and the fourth terminal device (140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data in circuit switching and / or packet switching channels. Representative networks may include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purpose of discussion in the present application, unless explained below, the architecture and topology of the network (150) may be irrelevant to the operation disclosed in the present application.
[0039] As an example of the application of the subject matter disclosed in this application, Figure 2 The placement of the video encoder and video decoder in a streaming environment is shown. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0040] The streaming system may include an acquisition subsystem (213), which may include a video source (201) such as a digital camera, which creates, for example, an uncompressed video sample stream (202). The video sample stream (202) is depicted as a thick line to emphasize that it has a higher data volume than an encoded video stream. The video sample stream (202) is depicted as a thick line to emphasize that it has a higher data volume than the encoded video stream. The video sample stream (202) may be processed by a video encoder (203) coupled to the digital camera. The video encoder (203) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video stream (204) is depicted as a thin line to emphasize that it has a lower data volume than the video sample stream (202), which may be stored on a streaming server (205) for future use. At least one streaming client (206, 208) may access a streaming server (205) to retrieve a first copy (207) and a second copy (209) of an encoded video stream (204). The client (206) may include a video decoder (210) that decodes the incoming first copy (207) of the encoded video stream and produces an output video sample stream (211) that can be presented on a display (212) or other presentation device (not depicted). In some streaming systems, the video stream (204), the first copy (207), or the second copy (209) may be encoded according to certain video encoding / compression standards. Examples of such standards include ITU-T Recommendation H.265. The video coding standard under development is informally referred to as Next Generation Video Coding. The video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the present application may be used in the context of the VVC standard.
[0041] Figure 3 It may be a block diagram of a video decoder (210) according to an embodiment disclosed in this application.
[0042] The receiver (310) may receive at least one encoded video sequence to be decoded by the video decoder (210). In the same embodiment or another embodiment, the receiver (301) may receive one encoded video sequence at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence may be received from a channel (312), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not shown). The receiver (310) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (315) may be coupled between the receiver (310) and an entropy decoder / parser (320) (hereinafter referred to as "parser"). When the receiver (310) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous synchronous network, it may not be necessary to configure the buffer memory (315), or the buffer memory may be made smaller. For use over a best-effort packet network such as the Internet, a buffer memory (315) may also be required, which may be relatively large and may have an adaptive size.
[0043] The video decoder (210) may include a parser (320) for reconstructing symbols (321) from an entropy-encoded video sequence. The types of symbols include information for managing the operation of the video decoder (210) and potential information for controlling a display device such as a display (212) that is not part of the video decoder but may be coupled to the video decoder, such as Figure 3As shown in . The control information for the display device may be in the form of a Supplemental Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not described). The parser (320) may parse and / or entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser (320) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a sub-picture, a tile, a slice, a brick, a macroblock, a Coding Tree Unit (CTU), a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and the like. A tile may refer to a rectangular area of a CU / CTU within a special tile row in a picture. A brick may refer to a rectangular area of a CU / CTU row within a special tile. A slice may refer to at least one brick of a picture contained in a NAL unit. A sub-picture may refer to a rectangular area of at least one slice in a picture. The entropy decoder / parser may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0044] The parser (320) may perform entropy decoding and / or parsing operations on the video sequence received from the buffer memory (315) to create symbols (321).
[0045] Depending on the type of the coded video picture or a portion of the coded video picture (e.g., inter-frame and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of the symbol (321) may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed by the parser (320) from the coded video sequence. For the sake of brevity, such subgroup control information flow between the parser (320) and the multiple units below is not described.
[0046] In addition to the functional blocks already mentioned, decoder 210 may be conceptually subdivided into several functional units as described below. In a practical embodiment operating under commercial constraints, many of these units closely interact with each other and may be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.
[0047] The first unit is a scaler and / or inverse transform unit (351). The scaler and / or inverse transform unit (351) receives quantized transform coefficients as symbols (321) from the parser (320) and control information including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output a block including sample values, which may be input into an aggregator (355).
[0048] In some cases, the output samples of the scaler and / or inverse transform unit (351) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses surrounding reconstructed information extracted from the current (partially reconstructed) picture (358) to generate a block of the same size and shape as the block being reconstructed. In some cases, the aggregator (355) adds the prediction information generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.
[0049] In other cases, the output samples of the scaler and / or inverse transform unit (351) may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit (353) may access the reference picture buffer (357) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (321), these samples may be added by the aggregator (355) to the output of the scaler and / or inverse transform unit (351) (in this case referred to as residual samples or residual signals) to generate output sample information. The motion compensation unit's retrieval of predicted samples from an address in the reference picture memory may be controlled by a motion vector, and the motion vector is provided to the motion compensation unit in the form of the symbols (321), for example, including X, Y and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.
[0050] The output samples of the aggregator (355) may be used by various loop filtering techniques in a loop filter unit (356). The video compression techniques may include in-loop filter techniques controlled by parameters included in the encoded video bitstream and available to the loop filter unit (356) as symbols (321) from the parser (320), and may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, and to previously reconstructed and loop filtered sample values.
[0051] The output of the loop filter unit (356) may be a sample stream that may be output to a display (212) and stored in a reference picture memory for subsequent inter-picture prediction.
[0052] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once a coded picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (320)), the current picture (358) may become part of the reference picture buffer (357), and new current picture memory may be reallocated before starting reconstruction of a subsequent coded picture.
[0053] The video decoder 210 may perform decoding operations according to a predetermined video compression technology, such as in ITU-T Recommendation H.265. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further defined by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.
[0054] In an embodiment, the receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0055] Figure 4 It may be a block diagram of a video encoder (203) according to an embodiment disclosed in this application.
[0056] The video encoder (203) may receive video samples from a video source (201) (not part of the encoder), which may capture video images to be encoded by the video encoder (203).
[0057] The video source (201) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (203), wherein the digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb4:4:4). In a media service system, the video source (201) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (201) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, wherein each pixel may include at least one sample depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by a person skilled in the art. The following description focuses on the samples.
[0058] According to an embodiment, the video encoder (203) may encode and compress pictures of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller (450). The controller (450) controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, coupling is not indicated in the figure. The parameters set by the controller (450) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, GOP layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller (450), which may involve a video encoder (203) optimized for a certain system design.
[0059] Some video decoders operate in a coding loop that is readily recognizable to those skilled in the art. As a simplified description, the coding loop may include an encoder (430) (hereinafter referred to as the "source encoder"), which is responsible for the coding part of creating symbols based on the input picture to be encoded and the reference picture, and a (local) decoder (433) embedded in the video encoder (203), which reconstructs the symbols to create sample data in a similar way to how the (remote) decoder creates sample data (because in the video compression technology considered in this application, any compression between the symbols and the encoded video code stream is lossless). The reconstructed sample stream is input to a reference picture memory (434). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the reference picture buffer contents are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift caused when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.
[0060] The operation of the "local" decoder (433) can be combined with the above Figure 3 The "remote" decoder (210) is identical to the one described in detail. However, additional brief reference is made to Figure 4 , when symbols are available and the entropy encoder (445) and the parser (320) are capable of losslessly encoding and / or decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210), including the channel (312), the receiver (310), the buffer memory (315), and the parser (320), may not be fully implemented in the local decoder (433).
[0061] At this point, it can be observed that any decoder technology, except for the parsing and / or entropy decoding present in the decoder, must also be present in the corresponding encoder in substantially the same functional form. For this reason, the present application focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. A more detailed description is only needed in certain areas and is provided below.
[0062] During operation, in some embodiments, the source encoder (430) may perform motion compensated predictive coding. The motion compensated predictive coding predictively encodes an input frame with reference to at least one previously encoded frame from a video sequence designated as a "reference frame". In this manner, the encoding engine (432) encodes the difference between a pixel block of the input frame and a pixel block of a reference frame that may be selected as a prediction reference for the input frame.
[0063] The local video decoder (433) may decode the encoded video data of the frame that may be designated as the reference frame based on the symbol created by the source encoder (430). The operation of the encoding engine (432) may be a lossy process. When the encoded video data is available at the video decoder ( Figure 4 When the video sequence is decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (433) replicates the decoding process that may be performed by the video decoder on the reference frame and may cause the reconstructed reference frame to be stored in the reference picture cache (434). In this way, the video encoder (203) may store a copy of the reconstructed reference frame locally that has common content (absent transmission errors) with the reconstructed reference frame to be obtained by the remote video decoder.
[0064] The predictor (435) may perform a prediction search for the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) may search the reference picture memory (434) for sample data (as a candidate reference pixel block) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as a suitable prediction reference for the new frame. The predictor (435) may operate pixel-by-pixel based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (435), it may be determined that the input picture may have a prediction reference obtained from a plurality of reference pictures stored in the reference picture memory (434).
[0065] The controller (450) may manage encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.
[0066] The outputs of all the above functional units may be entropy encoded in an entropy encoder (445). The entropy encoder (445) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.
[0067] The transmitter (440) may buffer the encoded video sequence created by the entropy encoder (445) in preparation for transmission over a communication channel (460), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (440) may combine the encoded video data from the video encoder (430) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0068] The controller (450) may manage the operation of the video encoder (203). During encoding, the controller (450) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following frame types:
[0069] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and features.
[0070] A predictive picture (P picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values for each block.
[0071] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict sample values for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0072] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or the blocks may be predictively coded (spatial prediction or intra-frame prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.
[0073] The video encoder (203) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T Recommendation H.265. In operation, the video encoder (203) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0074] In an embodiment, the transmitter (440) may transmit additional data when transmitting the encoded video. The video encoder (430) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0075] Recently, there has been some interest in partially aggregating or extracting multiple semantically independent pictures into a single video picture in the compressed domain. In particular, in the context of, for example, 360-degree codecs or certain surveillance applications, multiple semantically independent source pictures (e.g., six cube faces for a cube-projected 360-degree scene, or a single camera input in the case of a multi-camera surveillance setup) may require separate adaptive resolution settings to handle different scene activities at a given point in time. In other words, at a given point in time, the encoder may choose to use different resampling factors for different semantically independent pictures that make up the entire 360-degree or surveillance scene. This in turn requires that reference picture resampling be performed on parts of the coded picture when combined into a single picture, and that adaptive resolution codec signaling be available.
[0076] The following introduces some terms that will be referred to in the rest of this specification.
[0077] "Sub-picture" may in some cases refer to a semantically grouped rectangular arrangement of samples, blocks, macroblocks, coded units, or similar entities that may be independently coded with varying resolutions. At least one sub-picture may form a picture. At least one coded sub-picture may form a coded picture. At least one sub-picture may be assembled into a picture, and at least one sub-picture may be extracted from a picture. In some circumstances, at least one coded sub-picture may be assembled into a coded picture in a compressed domain without transcoding it to a sample level, and in the same circumstances or in some other circumstances, at least one coded sub-picture may be extracted from a coded picture in a compressed domain.
[0078] Adaptive Resolution Change (ARC) may refer to a mechanism that allows the resolution of a picture or sub-picture within a coded video sequence to be changed, for example, by resampling a reference picture. ARC parameters hereinafter refer to the control information required to perform adaptive resolution change, which may include, for example, filter parameters, scaling factors, resolution of output pictures and / or reference pictures, various control flags, and the like.
[0079] In an embodiment, encoding and decoding may be performed on a single, semantically independent coded video picture.Before describing the implications of encoding / decoding of multiple sub-pictures with independent ARC parameters and the additional complexity it implies, options for signaling ARC parameters will be described.
[0080] refer to Figures 5A-5E , several embodiments for signaling ARC parameters are shown. As indicated in each embodiment, they have certain advantages and disadvantages from the perspective of coding efficiency, complexity and architecture. Video coding standards or techniques may select at least one of these embodiments, or select options known in the relevant art, for signaling ARC parameters. These embodiments may not be mutually exclusive, and may be interchangeable according to application needs, the standard techniques involved, or the selection of encoders.
[0081] Categories of ARC parameters can include:
[0082] - Upsampling factors and / or downsampling factors in the X and Y dimensions, either separately or combined
[0083] - Add upsampling factor and / or downsampling factor of the time dimension to indicate constant speed upscaling and / or downscaling of a given number of images
[0084] - Either of the above may involve encoding one or more possibly shorter syntax elements, which may point to a table containing the factor(s).
[0085] - Resolution: The resolution of the input picture, output picture, reference picture, coded picture, in samples, blocks, macroblocks, CUs, or any other suitable granularity, combined or separate, in either the X or Y dimensions. If more than one resolution is present (e.g., one for the input picture and one for the reference picture), then in some cases one set of values may be inferred from another. For example, the resolution may be gated using a flag. See below for a more detailed example.
[0086] - "Warping" coordinates: similar to the coordinates used in Annex P of the H.263 standard, which can have an appropriate granularity as described above. Annex P of the H.263 standard defines an efficient way to encode such warping coordinates, but it is conceivable that other, potentially more efficient methods can also be envisioned. For example, the variable-length reversible "Huffman" encoding of the warping coordinates of Annex P can be replaced by a binary encoding of an appropriate length, where the length of the binary codeword can be derived, for example, from the maximum picture size (possibly multiplied by a factor and offset by a value) to allow "warping" outside the boundaries of the maximum picture size.
[0087] - Upsampling filter parameters or downsampling filter parameters: In embodiments, there may be only a single filter used for upsampling and / or downsampling. However, in embodiments, it may be desirable to allow greater flexibility in the design of the filter, which may be achieved by signaling filter parameters. Such parameters may be selected by an index into a list of possible filter designs, the filter may be fully specified (e.g., by a list of filter coefficients using an appropriate entropy coding technique), the filter may be selected implicitly by an upsampling ratio and / or a downsampling ratio, which in turn may be signaled according to any of the mechanisms mentioned above, etc.
[0088] In the following, this specification assumes that a finite set of upsampling factors and / or downsampling factors (using the same factors in the X and Y dimensions) indicated by codewords is encoded. The codewords can be variable-length encoded, for example, using Exp-Golomb coding common to certain syntax elements in video coding specifications (such as H.264 and H.265). A suitable mapping of values to upsampling factors / downsampling factors can be seen, for example, in Table 1:
[0089]
[0090] Many similar mappings can be designed, depending on the needs of the application and the capabilities of the upscaling and downscaling mechanisms available in the video compression technology or standard. This Table 1 can be extended to more values. The values can also be represented using entropy coding mechanisms other than Exp-Golomb codes, such as using binary encoding. Using binary encoding may have certain advantages when the resampling factors are of interest outside the video processing engine (most importantly the encoder and decoder) itself, such as the MANE (Media Aware Network Element). It should be noted that for cases where no resolution change is required, a shorter Exp-Golomb code can be chosen; in the above Table 1, only a single bit. For this most common case, using Exp-Golomb codes can have a coding efficiency advantage over using binary codes.
[0091] The number of entries in Table 1 and their semantics may be fully or partially configurable. For example, a basic form of Table 1 may be conveyed in a "higher layer" parameter set such as a sequence parameter set or a decoder parameter set. In an embodiment, one or more such tables may be defined in a video coding technology or standard and may be selected by, for example, a decoder or sequence parameter set.
[0092] The following describes how to include the up-sampling factor / down-sampling factor (ARC information) encoded as described above in a video coding technique or standard syntax. Similar considerations can be applied to one or more codewords that control the up-sampling filter / down-sampling filter. Regarding when a filter or other data structure requires a relatively large amount of data, see the discussion below.
[0093] like Figure 5A As shown, Annex P of the H.263 standard includes ARC information (502) in the form of four deformation coordinates in the picture header (501), more specifically, in the H.263PLUSPTYPE (503) header extension. This may be a wise design choice when a) there is a picture header available, and b) the ARC information is expected to change frequently. However, the overhead of using H.263-style signaling may be quite high, and the scaling factor of the picture boundary may not be relevant because the picture header may be of transient nature.
[0094] like Figure 5B As shown, JVCET-M135-v1 includes ARC reference information (505) (index) located in a picture parameter set (504), which ARC reference information indexes a table (506) that includes target resolutions that are in turn located inside a sequence parameter set (507). According to a verbal statement made by the author, the location of possible resolutions in the table (506) in the sequence parameter set (SPS) (507) can be justified by using the SPS as an interoperability negotiation point during capability exchange. Within the limits set by the values in the table (506), the resolution can be changed from picture to picture by referencing the appropriate picture parameter set (504).
[0095] refer to FIG. 5C to FIG. 5E , there may be the following embodiments to convey ARC information in a video bitstream. Each of these options has certain advantages over the above embodiments. The embodiments may exist simultaneously in the same video coding technology or standard.
[0096] In an embodiment, for example Figure 5C In the embodiment shown in , ARC information (509) such as a resampling (scaling) factor may be present in a slice header, a GOP header, a tile header, or a tile group header. Figure 5CAn embodiment is illustrated in which a tile group header (508) is used. This may be sufficient if the ARC information is small, such as, for example, a single variable length ue(v) as shown above, or a fixed length codeword of a few bits. The ARC information may apply, for example, to a sub-picture represented by a tile group, rather than to the entire picture, and including the ARC information directly in the tile group header has the additional advantage. See also below. In addition, even if the video compression technique or standard envisions adaptive resolution changes only for the entire picture (as opposed to, for example, adaptive resolution changes based on tile groups), placing the ARC information in the tile group header may have certain advantages from an error resilience perspective compared to placing it in an H.263-type picture header.
[0097] In an embodiment, in e.g. Figure 5D In the illustrated embodiment, the ARC information (512) itself may be present in an appropriate parameter set, such as, for example, a picture parameter set, a header parameter set, a tile parameter set, an adaptation parameter set, or the like. Figure 5D An embodiment is illustrated in which an adaptive parameter set (511) is used. The scope of the parameter set may advantageously be no larger than a picture, such as a group of tiles. The use of ARC information is implicit by activating the relevant parameter set. For example, when a video coding technique or standard considers ARC based only on pictures, then a picture parameter set or an equivalent parameter set may be appropriate.
[0098] In an embodiment, in e.g. Figure 5E In the illustrated embodiment, ARC reference information (513) may be present in a tile group header (514) or similar data structure. The reference information (513) may refer to a subset of ARC information (515) available in a parameter set (516) that has a scope beyond a single picture, such as a sequence parameter set or a decoder parameter set.
[0099] like Fig. 6A As shown, a tile group header (601), as an exemplary syntax structure of a header applicable to a (possibly rectangular) portion of a picture, may conditionally contain a variable-length Exp-Golomb coded syntax element dec_pic_size_idx (602) (shown in bold). The presence of this syntax element in the tile group header may be gated by using adaptive resolution (603). Here, the value of the flag is not shown in bold, which means that the point at which the flag appears in the codestream is when it appears in the syntax table. Whether adaptive resolution is used for the picture or a portion thereof may be signaled in any high-level syntax structure inside or outside the codestream. In the example shown, adaptive resolution is signaled in a sequence parameter set, as described below.
[0100] refer to Figure 6B, an excerpt of a sequence parameter set (610) is also shown. The first syntax element shown is adaptive_pic_resolution_change_flag (611). When true, this flag may indicate that adaptive resolution is used, which in turn may require specific control information. In this example, such control information is conditionally present based on the value of the flag, which is based on an if() statement in the sequence parameter set (612) and the tile group header (601).
[0101] When adaptive resolution is used, in this example, what is encoded is the output resolution in samples (613). Reference numeral 613 refers to both output_pic_width_in_luma_samples and output_pic_height_in_luma_samples, which together can define the resolution of the output picture. Elsewhere in the video coding technique or standard, certain restrictions on either value may be defined. For example, a level definition may limit the number of total output samples, which may be the product of the values of the two syntax elements described above. In addition, certain video coding techniques or standards, or external techniques or standards (e.g., system standards) may limit the range of values (e.g., one or both dimensions must be divisible by a power of 2) or aspect ratio (e.g., width and height must have a relationship of, for example, 4:3 or 16:9). Such restrictions may be introduced to facilitate hardware implementation or for other reasons and are well known in the art.
[0102] In some applications, it is recommended that the encoder instruct the decoder to use a certain reference picture size, rather than implicitly assuming its size to be the output picture size. In this example, the syntax element reference_pic_size_present_flag (614) gates the conditional presence of the reference picture size (615) (again, this number refers to both width and height).
[0103] Finally, one possible table of decoded picture width and height is shown. Such a table may be indicated, for example, by a table indication (num_dec_pic_size_in_luma_samples_minus1) (616). "Minus1" (minus 1) may refer to the interpretation of the value of this syntax element. For example, if the coded value is zero, there is one table entry. If the value is five, there are six table entries. For each "row" in the table, the decoded picture width and height are then contained in the syntax (617).
[0104] The presented table entries may be indexed using the syntax element dec_pic_size_idx (602) in the tile group header, allowing each tile group to have a different decode size (effectively a scaling factor).
[0105] Some video coding techniques or standards (e.g., VP9) support spatial scalability by implementing some form of reference picture resampling (which is signaled in a very different way than disclosed in this application) in conjunction with temporal scalability to achieve spatial scalability. More specifically, certain reference pictures can be upsampled to a higher resolution using ARC-type techniques to form the basis of a spatial enhancement layer. These upsampled pictures can be refined using standard prediction mechanisms at a higher resolution to increase detail.
[0106] The embodiments discussed in this application can be used in such environments. In some cases, in the same or another embodiment, the value in the Network Abstract Layer (NAL) unit header (e.g., the Temporal ID field) can be used to indicate not only the temporal layer, but also the spatial layer. Doing so may have certain advantages for certain system designs; for example, an existing Selected Forwarding Unit (SFU) created and optimized for forwarding selected for the temporal layer based on the Temporal ID value in the NAL unit header can be used in a scalable environment without modification. To achieve this, it may be necessary to map between the coded picture size and the temporal layer indicated by the Temporal ID field in the NAL unit header.
[0107] In some video coding techniques, an access unit (AU) may refer to one or more coded pictures, one or more slices, one or more tiles, one or more NAL units, etc., which are captured and combined into corresponding pictures, slices, tiles and / or NAL units at a given time instance. The time instance may be, for example, the composition time.
[0108] In HEVC and certain other video codecs, a picture order count (POC) value may be used to indicate a reference picture selected from a plurality of reference pictures stored in a decoded picture buffer (DPB). When an access unit (AU) includes one or more pictures, slices, or tiles, each picture, slice, or tile belonging to the same AU may carry the same POC value, from which it may be derived that they are created from content of the same composition time. In other words, in a scenario where two pictures / slices / tiles carry the same given POC value, the same given POC value may indicate that the two pictures / slices / tiles belong to the same AU and have the same composition time. Conversely, two pictures / slices / tiles with different POC values may indicate that those pictures / slices / tiles belong to different AUs and have different composition times.
[0109] In an embodiment, this strict relationship can be relaxed because the access unit can include pictures, slices, or tiles having different POC values. By allowing different POC values within an AU, the POC values can be used to identify potentially independently decodable pictures / slices / tiles having the same presentation time. This in turn enables support for multiple scalable layers without changing the reference picture selection signaling (e.g., reference picture set signaling or reference picture list signaling), as described in more detail below.
[0110] However, it is still desirable to be able to identify the AU to which a picture / slice / tile belongs from the POC value alone, relative to other pictures / slices / tiles having different POC values. This can be achieved as described below.
[0111] In an embodiment, in a high-level syntax structure (such as a NAL unit header, slice header, tile group header, SEI message, parameter set, or AU delimiter), the access unit count (AUC) is signaled. The value of the AUC can be used to identify which NAL units, pictures, slices, or tiles belong to a given AU. The value of the AUC can correspond to different composition time instances. The AUC value can be equal to a multiple of the POC value. The AUC value can be calculated by dividing the POC value by an integer value. In some cases, the division operation may impose a certain burden on the decoder implementation. In such cases, a small restriction in the numbering space of the AUC value can allow the division operation to be replaced by a shift operation. For example, the AUC value can be equal to the most significant bit (MSB) value of the POC value range.
[0112] In an embodiment, in a high-level syntax structure (such as a NAL unit header, slice header, tile group header, SEI message, parameter set, or AU delimiter), the value of the POC cycle for each AU (poc_cycle_au) is signaled. The poc_cycle_au can indicate how many different and consecutive POC values can be associated with the same AU. For example, if the value of the poc_cycle_au is equal to 4, then pictures, slices, or tiles with POC values equal to 0 to 3 (inclusive) are associated with an AU having an AUC value equal to 0, and pictures, slices, or tiles with POC values equal to 4 to 7 (inclusive) are associated with an AU having an AUC value equal to 1. Thus, the value of the AUC can be inferred by dividing the POC value by the value of the poc_cycle_au.
[0113] In an embodiment, the value of poc_cycle_au may be derived from information, such as that located in a video parameter set (VPS), identifying the number of spatial or SNR layers in a coded video sequence. An example of such a possible relationship is briefly described below. Although the derivation as described above may save some bits in the VPS and may therefore improve codec efficiency, in some embodiments, poc_cycle_au may be explicitly encoded in an appropriate high-level syntax structure at a level lower than the video parameter set so that poc_cycle_au may be minimized for a given small portion of a codestream such as a picture. Because POC values and / or values of syntax elements that indirectly reference the POC may be encoded in a low-level syntax structure, this optimization may save more bits than may be saved by the derivation process described above.
[0114] In an embodiment, Figure 8 An example of a syntax table is shown, which signals a syntax element of vps_poc_cycle_au in a VPS (or SPS), which indicates the poc_cycle_au used for all pictures / slices in a coded video sequence; and a syntax element of slice_poc_cycle_au, which indicates the poc_cycle_au of the current slice in the slice header. If the POC value of each AU increases uniformly, vps_contant_poc_cycle_per_au in the VPS is set to 1, and vps_poc_cycle_au is signaled in the VPS. In this case, slice_poc_cycle_au is not explicitly signaled, and the AUC value of each AU is calculated by dividing the value of the POC by vps_poc_cycle_au. If the POC value of each AU increases unevenly, vps_contant_poc_cycle_per_au in the VPS is set to 0. In this case, vps_access_unit_cnt is not signaled, but slice_access_unit_cnt is signaled in the slice header of each slice or picture. Each slice or picture may have a different slice_access_unit_cnt value. The AUC value for each AU is calculated by dividing the value of POC by slice_poc_cycle_au.
[0115] Fig. 9A block diagram illustrating an example of the above process is shown. For example, in operation S910, the VPS (or SPS) may be parsed, and it may be determined at operation S920 whether the POC period of each AU is constant within the encoded video sequence. If the POC period of each AU is constant (yes at operation S920), then at operation S930, the value of the access unit count of the specific access unit may be calculated based on the poc_cycle_au signaled for the encoded video sequence and the POC value of the specific access unit. If the POC period of each AU is not constant (no at operation S920), then at operation S940, the value of the access unit count of the specific access unit may be calculated based on the poc_cycle_au signaled at the picture level and the POC value of the specific access unit. At operation S950, a new VPS (or SPS) may be parsed.
[0116] In an embodiment, pictures, slices or tiles corresponding to an AU having the same AUC value may be associated with the same decoding or output time instance, even though the POC values of the pictures, slices or tiles may be different. Therefore, in the absence of any internal parsing / decoding dependencies on the pictures, slices or tiles in the same AU, all or a subset of the pictures, slices or tiles associated with the same AU may be decoded in parallel and may be output at the same time instance.
[0117] In an embodiment, pictures, slices or tiles corresponding to AUs having the same AUC value may be associated with the same composition time instance / display time instance even though the POC values for the pictures, slices or tiles may be different. When the composition time is included in the container format, pictures may be displayed at the same time instance even if they correspond to different AUs if they have the same composition time.
[0118] In an embodiment, each picture, slice or tile may have the same temporal identifier (temporal_id) in the same AU. All or a subset of pictures, slices or tiles corresponding to one time instance may be associated with the same temporal sublayer. In an embodiment, each picture, slice or tile may have the same or different spatial layer id (layer_id) in the same AU. All or a subset of pictures, slices or tiles corresponding to one time instance may be associated with the same or different spatial layers.
[0119] Figure 7An example of a video sequence structure with a combination of temporal_id, layer_id, POC and AUC values with adaptive resolution change is shown. In this example, a picture, slice or tile in the first AU with AUC=0 may have temporal_id=0 and layer_id=0 or 1, while a picture, slice or tile in the second AU with AUC=1 may have temporal_id=1 and layer_id=0 or 1, respectively. Regardless of the values of temporal_id and layer_id, the POC value of each picture increases by 1. In this example, the value of poc_cycle_au may be equal to 2. In an embodiment, the value of poc_cycle_au may be set equal to the number of (spatial scalability) layers. Therefore, in this embodiment, the value of POC increases by 2, while the value of AUC increases by 1.
[0120] In the above embodiments, all or a subset of inter-frame or inter-layer prediction structures and reference picture indications can be supported by using the existing reference picture set (RPS) signaling or reference picture list (RPL) signaling in HEVC. In the RPS or RPL, the selected reference picture is indicated by signaling the value of the POC between the current picture and the selected reference picture or the incremental value of the POC. In an embodiment, the RPS and RPL can be used to indicate the inter-frame or inter-layer prediction structure without changing the signaling, but with the following restrictions. If the value of the temporal_id of the reference picture is greater than the value of the temporal_id of the current picture, the current picture may not use the reference picture for motion compensation or other predictions. If the value of the layer_id of the reference picture is greater than the value of the layer_id current picture, the current picture may not use the reference picture for motion compensation or other predictions.
[0121] In an embodiment, POC difference-based motion vector scaling for temporal motion vector prediction may be disabled across multiple pictures within an access unit. Thus, although each picture may have a different POC value within an access unit, the motion vectors are not scaled and used for temporal motion vector prediction within an access unit. This is because reference pictures with different POCs in the same AU are considered to be reference pictures with the same temporal instance. Therefore, in this embodiment, the motion vector scaling function may return 1 when the reference picture belongs to the AU associated with the current picture.
[0122] In an embodiment, when the spatial resolution of the reference picture is different from the spatial resolution of the current picture, motion vector scaling based on POC difference for temporal motion vector prediction may be optionally disabled on multiple pictures. When motion vector scaling is allowed, the motion vector is scaled based on both the POC difference and the spatial resolution ratio between the current picture and the reference picture.
[0123] In an embodiment, for temporal motion vector prediction, especially when poc_cycle_au has a non-uniform value (e.g., when vps_contant_poc_cycle_per_au==0), the motion vector may be scaled based on the AUC difference instead of the POC difference. Otherwise (e.g., when vps_contant_poc_cycle_per_au==1), the motion vector scaling based on the AUC difference may be the same as the motion vector scaling based on the POC difference.
[0124] In an embodiment, when the motion vector is scaled based on the AUC difference, the reference motion vector in the same AU (with the same AUC value) as the current picture is not scaled based on the AUC difference, and the reference motion vector is used for non-scaled motion vector prediction or for scaling the motion vector prediction based on the spatial resolution ratio between the current picture and the reference picture.
[0125] In an embodiment, the AUC value may be used to identify the boundary of the AU and for the operation of the hypothetical reference decoder (HRD), the operation of which requires input timing and output timing with AU granularity. In an embodiment, the decoded picture with the highest layer in the AU may be output for display. The AUC value and the layer_id value may be used to identify the output picture.
[0126] In an embodiment, a picture may be composed of one or more sub-pictures. Each sub-picture may cover a partial area or the entire area of the picture. The area supported by a sub-picture may overlap or not overlap with the area supported by another sub-picture. The area composed of one or more sub-pictures may cover or not cover the entire area of the picture. If a picture includes sub-pictures, the area supported by the sub-picture may be the same as the area supported by the picture.
[0127] In an embodiment, a sub-picture may be encoded by an encoding method similar to that used for an encoded picture. A sub-picture may be encoded independently or may be encoded based on another sub-picture or an encoded picture. A sub-picture may or may not have any parsing dependency on another sub-picture or an encoded picture.
[0128] In an embodiment, the encoded sub-picture may be included in one or more layers. The encoded sub-pictures in a layer may have different spatial resolutions. The original sub-picture may be spatially resampled (e.g., upsampled or downsampled), encoded with different spatial resolution parameters, and included in a codestream corresponding to the layer.
[0129] In an embodiment, a sub-picture with (W, H) may be encoded and included in the encoded bitstream corresponding to layer 0, where W indicates the width of the sub-picture and H indicates the height of the sub-picture. w,k , H* S h,k ), the upsampled (or downsampled) sub-picture can be encoded and included in the encoded bitstream corresponding to layer k, where S w,k, S h,k Indicates the horizontal resampling rate and the vertical resampling rate. If S w,k , S h,k If the value of is greater than 1, the resampling can be upsampling. However, if S w,k , S h,k If the value of is less than 1, the resampling can be downsampling.
[0130] In an embodiment, an encoded sub-picture in a layer may have a different visual quality than an encoded sub-picture in another layer in the same sub-picture or in a different sub-picture. i,n is encoded, and the sub-image j in layer m is quantized with parameter Q j,m to encode.
[0131] In an embodiment, a coded sub-picture in a layer may be independently decodable without any parsing dependency or decoding dependency on a coded sub-picture in another layer of the same local region. A sub-picture layer that is independently decodable without reference to another sub-picture layer of the same local region is an independent sub-picture layer. A coded sub-picture in an independent sub-picture layer may or may not have a decoding dependency or parsing dependency on a previously coded sub-picture in the same sub-picture layer, but the coded sub-picture may not have any dependency on a coded picture from another sub-picture layer.
[0132] In an embodiment, an encoded sub-picture in a layer may be dependently decodable, with any parsing dependency or decoding dependency on an encoded sub-picture in another layer from the same local region. A sub-picture layer that is dependently decoded by referencing another sub-picture layer of the same local region is a dependent sub-picture layer. An encoded sub-picture in a dependent sub-picture may reference an encoded sub-picture belonging to the same sub-picture, a previously encoded sub-picture in the same sub-picture layer, or both reference sub-pictures.
[0133] In an embodiment, a coded sub-picture consists of one or more independent sub-picture layers and one or more dependent sub-picture layers. However, for a coded sub-picture, there may be at least one independent sub-picture layer. The value of the layer identifier (layer_id) of the independent sub-picture layer may be equal to 0, which may be present in the NAL unit header or another high-level syntax structure. The sub-picture layer with layer_id equal to 0 may be a base sub-picture layer.
[0134] In an embodiment, a picture may consist of one or more foreground sub-pictures and one background sub-picture. The area supported by the background sub-picture may be equal to the area of the picture. The area supported by the foreground sub-picture may overlap with the area supported by the background sub-picture. The background sub-picture may be a base sub-picture layer and the foreground sub-picture may be a non-base (enhanced) sub-picture layer. One or more non-base sub-picture layers may be decoded with reference to the same base layer. Each non-base sub-picture layer with layer_id equal to a may reference a non-base sub-picture layer with layer_id equal to b, where a is greater than b.
[0135] In an embodiment, a picture may consist of one or more foreground sub-pictures with or without background sub-pictures. Each sub-picture may have its own base sub-picture layer and one or more non-base (enhancement) layers. Each base sub-picture layer may be referenced by one or more non-base sub-picture layers. Each non-base sub-picture layer with layer_id equal to a may reference a non-base sub-picture layer with layer_id equal to b, where a is greater than b.
[0136] In an embodiment, a picture may consist of one or more foreground sub-pictures with or without background sub-pictures. Each coded sub-picture in a (base or non-base) sub-picture layer may be referenced by one or more non-base layer sub-pictures belonging to the same sub-picture, and one or more non-base layer sub-pictures not belonging to the same sub-picture.
[0137] In an embodiment, a picture may consist of one or more foreground sub-pictures with or without background sub-pictures. A sub-picture in layer a may be further partitioned into multiple sub-pictures in the same layer. One or more coded sub-pictures in layer b may reference a partitioned sub-picture in layer a.
[0138] In an embodiment, a coded video sequence (CVS) may be a set of coded pictures. A CVS may consist of one or more coded sub-picture sequences (CSPS), where a CSPS may be a set of coded sub-pictures covering the same local area of a picture. A CSPS may have the same or different temporal resolution as the coded video sequence.
[0139] In an embodiment, the CSPS may be encoded and included in one or more layers. The CSPS may be composed of one or more CSPS layers. By decoding one or more CSPS layers corresponding to the CSPS, a sub-picture sequence corresponding to the same local area may be reconstructed.
[0140] In the same or another embodiment, the number of CSPS layers corresponding to one CSPS may be the same as or different from the number of CSPS layers corresponding to another CSPS.
[0141] In an embodiment, one CSPS layer may have a different temporal resolution (e.g., frame rate) than another CSPS layer. The original (uncompressed) sub-picture sequence may be temporally resampled (e.g., upsampled or downsampled), encoded with different temporal resolution parameters, and included in the codestream corresponding to the layer.
[0142] In an embodiment, a sub-picture sequence with a frame rate F may be encoded and included in the encoded codestream corresponding to layer 0, while a sub-picture sequence with a frame rate F* S may be encoded and included in the encoded codestream corresponding to layer 0. t,k The sub-picture sequence of the temporal up-sampled (or down-sampled) original sub-picture sequence can be encoded and included in the coded bitstream corresponding to layer k, where S t,k indicates the temporal sampling rate of layer k. If S t,k If the value of S is greater than 1, the temporal resampling process may be a frame rate upconversion. t,k If the value of is less than 1, the temporal resampling process may be a frame rate down-conversion.
[0143] In an embodiment, when a sub-picture with CSPS layer a is referenced by a sub-picture with CSPS layer b for motion compensation or any inter-layer prediction, if the spatial resolution of CSPS layer a is different from the spatial resolution of CSPS layer b, the decoded pixels in CSPS layer a are resampled and used as reference. The resampling process may use upsampling filtering or downsampling filtering.
[0144] Fig.10 An example video bitstream including a background video CSPS layer with layer_id equal to 0 and multiple foreground CSPS layers is shown. Although the encoded sub-picture can be composed of one or more CSPS layers, the background area that does not belong to any foreground CSPS layer can include a base layer. The base layer can contain background areas and foreground areas, while the enhanced CSPS layer contains foreground areas. At the same area, the enhanced CSPS layer can have better visual quality than the base layer. The enhanced CSPS layer can refer to the reconstructed pixels and motion vectors of the base layer corresponding to the same area.
[0145] In an embodiment, in a video file, the video bitstream corresponding to the base layer is included in a track, while the CSPS layer corresponding to each sub-picture is included in a separate track.
[0146] In an embodiment, the video bitstream corresponding to the base layer is included in a track, while the CSPS layer having the same layer_id is included in a separate track. In this example, the track corresponding to layer k includes only the CSPS layer corresponding to layer k.
[0147] In, each CSPS layer of each sub-picture is stored in a separate track. Each track may or may not have a parsing dependency or decoding dependency on one or more other tracks.
[0148] In the same or another embodiment, each track may include the bitstreams of layers i to j in the CSPS layers corresponding to all or a subset of the sub-pictures, where 0 < i =< j =< k, and k is the highest layer of the CSPS.
[0149] In, a picture includes one or more associated media data, and these associated media data include depth maps, alpha maps, 3D geometric data, occupancy maps, etc. Such associated timed media data can be divided into one or more data sub-bitstreams, and each data sub-bitstream corresponds to a sub-picture.
[0150] Fig.11 An example of a video conference based on a multi-layer sub-picture method is shown. In the video bitstream, there is a base layer video bitstream corresponding to a background picture and one or more enhancement layer video bitstreams corresponding to foreground sub-pictures. Each enhancement layer video bitstream corresponds to a CSPS layer. In the display, the picture corresponding to the base layer is displayed by default. It includes picture-in-picture (PIP) of one or more users. When the client control selects a specific user, the enhanced CSPS layer corresponding to the selected user is decoded and displayed with enhanced quality or spatial resolution.
[0151] Fig.12 A block diagram showing an example illustrating the above process is shown. For example, in operation S1210, a video bitstream having multiple layers is decoded. In operation S1220, a background region and one or more foreground sub-pictures are identified. At operation S1230, it is determined whether a specific sub-picture region, such as one of the foreground sub-pictures, is selected. If a specific sub-picture region is selected (yes at operation S1240), the enhanced sub-picture is decoded and displayed. If no specific sub-picture region is selected (no at operation S1240), the background region can be decoded and displayed.
[0152] In an embodiment, a network middlebox (such as a router) can select a subset of layers to send to a user based on its bandwidth. The picture / sub-picture organization can be used for bandwidth adaptation. For example, if a user does not have bandwidth, the router strips off layers or selects some sub-pictures due to the importance of layers and sub-pictures or based on usage settings. This can be done dynamically to adapt to the bandwidth.
[0153] Fig.13 An embodiment of a use case involving 360° video is shown. When a spherical 360° picture (e.g., picture 1310) is projected onto a flat picture, the projected 360° picture can be partitioned into multiple sub-pictures as a base layer. For example, the multiple sub-pictures can include a rear sub-picture, a top sub-picture, a right sub-picture, a left sub-picture, a front sub-picture, and a bottom sub-picture. An enhancement layer for a specific sub-picture (e.g., a front sub-picture) can be encoded and transmitted to a client. The decoder is capable of decoding a base layer including all sub-pictures and decoding an enhancement layer for a selected sub-picture. When the current viewport is the same as the selected sub-picture, the displayed picture includes a decoded sub-picture with an enhancement layer, and the displayed picture has a higher quality. Otherwise, a decoded picture with a base layer can be displayed at a lower quality.
[0154] In an embodiment, any layout information for display may be present in the file as supplementary information (such as SEI messages or metadata). One or more decoded sub-pictures may be repositioned and displayed according to the layout information signaled. The layout information may be signaled by the streaming server or broadcaster, or may be regenerated by a network entity or cloud server, or may be determined by a user's customized settings.
[0155] In an embodiment, when the input image is divided into one or more (rectangular) sub-regions, each sub-region can be encoded as an independent layer. Each independent layer corresponding to a local area can have a unique layer_id value. For each independent layer, a signal can be sent to indicate the sub-image size and position information. For example, the image size (width, height), the offset information of the upper left corner (x_offset, y_offset). Fig.14 An example of the layout of the divided sub-pictures, their sub-picture size and position information, and the picture prediction structure corresponding to the sub-pictures is shown. The layout information is signaled in a high-level syntax structure (such as a parameter set, a slice header or a tile group header, or an SEI message), and the layout information includes the sub-picture size and the sub-picture position.
[0156] In an embodiment, each sub-picture corresponding to an independent layer may have its unique POC value within an AU. When indicating a reference picture in a picture stored in a DPB by using one or more syntax elements in an RPS or RPL structure, one or more POC values corresponding to each sub-picture of a layer may be used.
[0157] In an embodiment, to indicate the (inter-layer) prediction structure, layer_id may not be used and the POC (delta) value may be used.
[0158] In an embodiment, a sub-picture having a POC value equal to N corresponding to a layer (or a local area) may or may not be used as a reference picture for a sub-picture having a POC value equal to N+K, the sub-picture having a POC value equal to N+K corresponding to the same layer (or the same local area), the reference picture being used for motion compensated prediction. In most cases, the value of the number K may be equal to the maximum number of (independent) layers, and the value of the number K may be the same as the number of sub-regions.
[0159] In an embodiment, Fig.15 Shows Fig.14 When the input picture is divided into multiple (e.g., four) sub-regions, each local region can be encoded with one or more layers. In this case, the number of independent layers can be equal to the number of sub-regions, and one or more layers can correspond to one sub-region. Therefore, each sub-region can be encoded with one or more independent layers and zero or more dependent layers.
[0160] In an embodiment, Fig.15 In , the input picture can be divided into four sub-regions. As an example, the upper right sub-region can be encoded into two layers, namely layer 1 and layer 4, and the lower right sub-region can be encoded into two layers, namely layer 3 and layer 5. In this case, layer 4 can refer to layer 1 for motion compensation prediction, and layer 5 can refer to layer 3 for motion compensation.
[0161] In an embodiment, loop filtering (such as deblocking filtering, adaptive loop filtering, shaper, bilateral filtering or any deep learning based filtering) across layer boundaries may (optionally) be disabled.
[0162] In an embodiment, motion compensated prediction or intra-block copy across layer boundaries may (optionally) be disabled.
[0163] In an embodiment, boundary filling or loop filtering at sub-picture boundaries for motion compensated prediction may be performed optionally. A flag indicating whether boundary filling is performed is signaled in a high-level syntax structure such as one or more parameter sets (VPS, SPS, PPS, or APS), a slice header or a tile group header, or an SEI message.
[0164] In an embodiment, in a VPS or an SPS, layout information of one or more sub-regions (or one or more sub-pictures) is signaled. Fig.16A shows an example of a syntax element in a VPS, and Fig. 16B An example of a syntax element in an SPS is shown. In this example, vps_sub_picture_dividing_flag is signaled in the VPS. This flag may indicate whether one or more input pictures are divided into multiple sub-regions. When the value of vps_sub_picture_dividing_flag is equal to 0, one or more input pictures in one or more encoded video sequences corresponding to the current VPS may not be divided into multiple sub-regions. In this case, the input picture size may be equal to the encoded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples), and the input picture size is signaled in the SPS. When the value of vps_sub_picture_dividing_flag is equal to 1, one or more input pictures may be divided into multiple sub-regions. In this case, the syntax elements vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples are signaled in the VPS. The values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may be equal to the width and height of one or more input pictures, respectively.
[0165] In an embodiment, the values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may not be used for decoding, but may be used for synthesis and display.
[0166] In an embodiment, when the value of vps_sub_picture_dividing_flag is equal to 1, the syntax elements pic_offset_x and pic_offset_y may be signaled in the SPS corresponding to one or more specific layers. In this case, the coded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples) signaled in the SPS may be equal to the width and height of the sub-region corresponding to the specific layer. Moreover, the position of the upper left corner of the sub-region (pic_offset_x, pic_offset_y) may be signaled in the SPS.
[0167] In an embodiment, the position information (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region may not be used for decoding, but may be used for synthesis and display.
[0168] In an embodiment, layout information (size and position) of all or a subset of sub-regions of an input picture, dependency information between layers is signaled in a parameter set or SEI message. Fig.17 An example of a syntax element is shown, which is used to indicate layout information of a sub-region, dependency information between layers, and information about the relationship between a sub-region and one or more layers. In this example, the syntax element num_sub_region indicates the number of (rectangular) sub-regions in the current encoded video sequence. The syntax element num_layers indicates the number of layers in the current encoded video sequence. The value of num_layers may be equal to or greater than the value of num_sub_region. When any sub-region is encoded as a single layer, the value of num_layers may be equal to the value of num_sub_region. When one or more sub-regions are encoded as multiple layers, the value of num_layers may be greater than the value of num_sub_region. The syntax element direct_dependency_flag[i][j] indicates a dependency from the jth layer to the ith layer. num_layers_for_region[i] indicates the number of layers associated with the i-th sub-region. sub_region_layer_id[i][j] indicates the layer_id of the jth layer associated with the i-th sub-region. sub_region_offset_x[i] and sub_region_offset_y[i] indicate the horizontal position and vertical position of the upper left corner of the i-th sub-region, respectively. sub_region_width[i] and sub_region_height[i] indicate the width and height of the i-th sub-region, respectively.
[0169] In an embodiment, one or more syntax elements are signaled in a high-level syntax structure (e.g., VPS, DPS, SPS, PPS, APS, or SEI message). The one or more syntax elements specify an output layer set to indicate one of a plurality of layers to be output with or without profile level information. Fig.18 , a syntax element num_output_layer_sets is signaled in the VPS, which indicates the number of output layer sets (OLS) in the coded video sequence that references the VPS. For each output layer set, as many output_layer_flags as the number of output layers may be signaled.
[0170] In an embodiment, output_layer_flag[i] equal to 1 specifies that the i-th layer is output. vps_output_layer_flag[i] equal to 0 specifies that the i-th layer is not output.
[0171] In an embodiment, one or more syntax elements may be signaled in a high-level syntax structure (e.g., VPS, DPS, SPS, PPS, APS, or SEI message) that specify profile level information for each output layer set. Fig.18 , a syntax element num_profile_tile_level may be signaled in the VPS, indicating the number of profile level information for each OLS in the coded video sequence referring to the VPS. For each output layer set, a set of syntax elements for profile level information may be signaled with the same number as the number of output layers, or an index indicating a specific profile level information in an entry in the profile level information may be signaled.
[0172] In an embodiment, profile_tier_level_idx[i][j], in the list of profile_tier_level() syntax structures in the VPS, specifies the index of the profile_tier_level() syntax structure that applies to the j-th tier of the i-th OLS.
[0173] In the embodiments, reference Fig.19 , when the maximum number of layers is greater than 1 (vps_max_layers_minus1>0), the syntax elements num_profile_tile_level and / or num_output_layer_sets may be signaled.
[0174] In the embodiments, reference Fig.19, the syntax element vps_output_layers_mode[i] may be present in the VPS, indicating the mode of output layer signaling for the i-th output layer set.
[0175] In an embodiment, vps_output_layers_mode[i] equal to 0 specifies that only the highest layer with the i-th output layer set is output. vps_output_layer_mode[i] equal to 1 specifies that all layers with the i-th output layer set are output. vps_output_layer_mode[i] equal to 2 specifies that the layers output are the layers with vps_output_layer_flag[i][j] equal to 1 and with the i-th output layer set. More values may be reserved.
[0176] In an embodiment, output_layer_flag[i][j] may or may not be signaled depending on the value of vps_output_layers_mode[i] for the i-th output layer set.
[0177] In the embodiments, reference Fig.19 , there is a flag vps_ptl_signal_flag[i] for the i-th output layer set. Depending on the value of vps_ptl_signal_flag[i], the profile level information for the i-th output layer set may be signaled or not.
[0178] In the embodiments, reference Fig. 20 , the number of sub-pictures in the current CVS, max_subpics_minus1, can be signaled in a high-level syntax structure (eg, VPS, DPS, SPS, PPS, APS, or SEI message).
[0179] In the embodiments, reference Fig. 20 , when the number of sub-pictures is greater than 1 (max_subpics_minus1>0), the sub-picture identifier sub_pic_id[i] for the i-th sub-picture may be signaled.
[0180] In an embodiment, in the VPS, one or more syntax elements are signaled which indicate the sub-picture identifiers of each layer belonging to each output layer set. Fig. 20 , sub_pic_id_layer[i][j][k] indicates the kth sub-picture present in the jth layer of the i-th output layer set. Using this information, the decoder can identify which sub-picture can be decoded and output for each layer of a specific output layer set.
[0181] In an embodiment, a picture header (PH) is a syntax structure containing syntax elements that apply to all slices of a coded picture. A picture unit (PU) is a group of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and contain exactly one coded picture. A PU may contain a picture header (PH) and one or more video coding layer (VCL) NAL units that make up a coded picture.
[0182] In an embodiment, an adaptation parameter set (APS) may be a syntax structure containing syntax elements that may apply to zero or more slices determined by zero or more syntax elements found in a slice header.
[0183] In an embodiment, the adaptation_parameter_set_id signaled in the APS (e.g., as u(5) signaled) may provide an identifier of the APS for reference by other syntax elements. When ps_params_type is equal to ALF_APS or SCALING_APS, the value of adaptation_parameter_set_id may be in the range of 0 to 7, inclusive. When aps_params_type is equal to LMCS_APS, the value of adaptation_parameter_set_id may be in the range of 0 to 3, inclusive.
[0184] In an embodiment, the adaptation_parameter_set_id signaled as ue(v) in APS may provide an identifier of the APS for reference by other syntax elements. When ps_params_type is equal to ALF_APS or SCALING_APS, the value of adaptation_parameter_set_id may be in the range of 0 to 7*(maximum number of layers in the current CVS), including the endpoints. When aps_params_type is equal to LMCS_APS, the value of adaptation_parameter_set_id may be in the range of 0 to 3*(maximum number of layers in the current CVS), including the endpoints.
[0185] In an embodiment, each APS (RBSP) may be available to the decoding process before it is referenced, included in at least one AU, or provided by an external device. Wherein the temporal identifier (e.g., TemporalId) of the at least one AU is less than or equal to the temporal identifier (e.g., TemporalId) of the coded slice NAL unit that references the APS. When the APS NAL unit is included in the AU, the value of the temporal identifier (e.g., TemporalId) of the APS NAL unit may be equal to the value of the temporal identifier (e.g., TemporalId) of the AU that includes the APS NAL unit.
[0186] In an embodiment, each APS (RBSP) may be available to the decoding process before it is referenced, the APS is included in at least one AU whose temporal identifier (e.g., TemporalId) is equal to 0, or is provided by an external device. When an APS NAL unit is included in an AU, the value of the temporal identifier (e.g., TemporalId) of the APS NAL unit may be equal to 0.
[0187] In an embodiment, the APS (RBSP) may be available to the decoding process before it is referenced, included in at least one AU, or provided by an external device, wherein the temporal identifier (e.g., TemporalId) of the at least one AU is equal to the temporal identifier (e.g., TemporalId) of the APS NAL unit in the CVS, and the CVS contains one or more PHs referencing the APS or one or more coded slice NAL units referencing the APS.
[0188] In an embodiment, the APS (RBSP) may be available to the decoding process before it is referenced, the APS is included in at least one AU with a temporal identifier (e.g., TemporalId) equal to 0 in the CVS, or provided by an external device. The CVS contains one or more PHs or one or more coded slice NAL units that reference the APS.
[0189] In an embodiment, when the flag no_temporal_sublayer_switching_flag is signaled in a DPS, VPS, SPS or PPS, the temporal identifier (e.g., TemporalId) value of the APS of the parameter set (referring to the parameter set containing the flag equal to 1) may be equal to 0, while the temporal identifier (e.g., TemporalId) value of the APS (referring to the parameter set containing the flag equal to 1) may be equal to or greater than the temporal identifier (e.g., TemporalId) value of the parameter set.
[0190] In an embodiment, each APS (RBSP) may be available to the decoding process before it is referenced, included in at least one AU, or provided by an external device. The temporal identifier (e.g., TemporalId) of the at least one AU is less than or equal to the temporal identifier (e.g., TemporalId) of the coded slice NAL unit (or PH NAL unit) that references each APS. When an APS NAL unit is included in the AU before the AU contains a coded slice NAL unit that references the APS, a VCL NAL unit that enables temporal upper layer switching or a VCL ANL unit with nal_unit_type equal to STSA_NUT may not exist after the APS NAL unit and before the coded slice NAL unit that references the APS. The VCL NAL unit that enables temporal upper layer switching or the VCL ANL unit with nal_unit_type equal to STSA_NUT indicates that the picture in the VCL NAL unit can be a step-wise temporal sublayer access (STSA) picture. Fig.21 An example regarding this constraint is shown.
[0191] In an embodiment, the APS NAL unit and the coded slice NAL units referencing the APS (and its PH NAL units) may be included in the same AU.
[0192] In an embodiment, an APS NAL unit and a STSA NAL unit may be included in the same AU, which may precede the coded slice NAL units (and their PH NAL units) that reference the APS.
[0193] In an embodiment, the STSA NAL unit referencing the APS, the APS NAL unit referencing the APS, and the coded slice NAL unit referencing the APS (and its PH NAL unit) may exist in the same AU.
[0194] In an embodiment, the temporal identifier (eg, TemporalId) value of the VCL NAL unit containing the APS may be equal to the temporal identifier (eg, TemporalId) value of the previous STSA NAL unit.
[0195] In an embodiment, the picture order count (POC) value of the APS NAL unit may be equal to or greater than the POC value of the STSA NAL unit. Fig.21 In the example, the value of POC M of the APS can be equal to or greater than the value of POC L of the STSA NAL unit.
[0196] In an embodiment, the picture order count (POC) value of the coded slice or PH NAL unit that references the APS NAL unit may be equal to or greater than the POC value of the referenced APS NAL unit. Fig.21 In the example, the value of POC M of the VCL NAL unit that references the APS can be equal to or greater than the value of POC L of the APS NAL unit.
[0197] In an embodiment, the APS (RBSP) may be used by the decoding process before it is referenced by one or more PHs or one or more coded slice NAL units. The APS is contained in at least one PU in the CVS or provided by an external device. The nuh_layer_id of the at least one PU is equal to the lowest nuh_layer_id value of the coded slice NAL unit that references the APS NAL unit. The CVS includes one or more PHs that reference the APS or an APS NAL unit that references one or more coded slice NAL units of the APS.
[0198] In an embodiment, the APS (RBSP) may be used by the decoding process before it is referenced by one or more PHs or one or more coded slice NAL units, the APS is included in at least one PU in the CVS, or provided by an external device. The TemporalId of the at least one PU is equal to the TemporalId of the APS NAL unit, and the nuh_layer_id of the at least one PU is equal to the lowest nuh_layer_id value of the coded slice NAL unit that references the APS NAL unit. The CVS contains one or more PHs that reference the PPS or one or more coded slice NAL units that reference the PPS.
[0199] In an embodiment, the APS (RBSP) may be used by the decoding process before it is referenced by one or more PHs or one or more coded slice NAL units, the APS is included in at least one PU in the CVS, or provided by an external component. The TemporalId of the at least one PU is equal to 0, and the nuh_layer_id is equal to the lowest nuh_layer_id value of the coded slice NAL unit that references the APS NAL unit. The CVS contains one or more PHs that reference the PPS or one or more coded slice NAL units that reference the PPS.
[0200] In an embodiment, all APS NAL units within a PU with a specific value of adaptation_parameter_set_id and a specific value of aps_params_type may have the same content regardless of whether the APS NAL unit is a prefix APS NAL unit or a suffix APS NAL unit.
[0201] In an embodiment, regardless of the nuh_layer_id value, APS NAL units may share the same value space of adaptation_parameter_set_id and aps_params_type.
[0202] In an embodiment, the nuh_layer_id value of an APS NAL unit may be equal to the lowest nuh_layer_id value of the coded slice NAL units that reference the NAL unit that references the APS NAL unit.
[0203] In an embodiment, when an APS with nuh_layer_id equal to m is referenced by one or more coded slice NAL units with nuh_layer_id equal to n, the layer with nuh_layer_id equal to m may be the same as the layer with nuh_layer_id equal to n, or the same as the (direct or indirect) reference layer of the layer with nuh_layer_id equal to m.
[0204] Fig. 22 is a flow chart of an example method 2200 for decoding an encoded video stream. In some embodiments, Fig. 22 At least one method block of may be performed by decoder 210. In some embodiments, Fig. 22 At least one method block of may be performed by another device or a group of devices (such as the video encoder 203 ) separate from the decoder 210 or including the decoder 210 .
[0205] like Fig. 22 As shown, the method 2200 may include obtaining an encoded video sequence from an encoded video code stream, wherein the encoded video sequence includes picture units corresponding to the encoded pictures (block 2210).
[0206] like Fig. 22 As further shown, the method 2200 may include obtaining a picture header (PH) network abstraction layer (NAL) unit included in the picture unit (block 2220).
[0207] like Fig. 22 As further shown, the method 2200 may include obtaining at least one video coding layer (VCL) NAL unit included in the picture unit (block 2230).
[0208] like Fig. 22As further shown, method 2200 may include decoding the encoded picture based on the PH NAL unit, the at least one VCL NAL unit and the adaptive parameter set APS to obtain a decoded picture, wherein the APS is included in the APS NAL unit obtained from the CVS, and the APS NAL unit is used for the decoding earlier than the at least one VCL NAL unit (box 2240).
[0209] like Fig. 22 As further shown, method 2200 may include outputting the decoded picture (block 2250).
[0210] In an embodiment, the APS NAL unit is available for decoding before being referenced by one or more picture headers PH or one or more coded slice NAL units, wherein the APS NAL unit is contained in at least one prediction unit PU, and the nuh_layer_id value of the at least one PU is equal to the lowest nuh_layer_id value of the one or more coded slice NAL units that reference the APS NAL unit in the CVS, and the CVS includes the one or more PH or the one or more coded slice NAL units that reference the APS.
[0211] In an embodiment, the temporal identifier of at least one VCL NAL unit is greater than or equal to the temporal identifier of the APS NAL unit.
[0212] In an embodiment, the temporal identifier of the APS NAL unit may be equal to 0.
[0213] In an embodiment, the POC of at least one VCL NAL unit may be greater than or equal to the POC of the APS NAL unit.
[0214] In an embodiment, the layer identification of the PH NAL unit and the layer identification of at least one VCL NAL unit is greater than or equal to the layer identification of the APS NAL unit.
[0215] In an embodiment, the PH NAL unit, at least one VCL NAL unit, and the APS NAL unit are included in a single access unit.
[0216] In an embodiment, the encoded video sequence further comprises a STSA NAL unit corresponding to the STSA picture, and the STSA NAL unit is not located between the APS NAL unit and the at least one VCL NAL unit.
[0217] In an embodiment, at least one VCL NAL unit, an APS NAL unit, and a STSA NAL unit may be included in a single access unit.
[0218] In an embodiment, the temporal identifier of the APS NAL unit may be greater than or equal to the temporal identifier of the STSA NAL unit.
[0219] In an embodiment, the picture order count POC of the APS NAL unit may be greater than or equal to the POC of the STSA NAL unit.
[0220] although Fig. 22 An example block diagram of method 2200 is shown, but in some embodiments, method 2200 may include comparing Fig. 22 Additional blocks in addition to, fewer blocks than, different blocks from, or blocks arranged differently from those depicted in the method 2200. Additionally or alternatively, two or more of the blocks in the method 2200 may be executed in parallel.
[0221] Furthermore, the proposed methods may be implemented by a processing circuit (eg, at least one processor or at least one integrated circuit). In one example, at least one processor executes a program stored in a non-transitory computer-readable medium to perform at least one of the proposed methods.
[0222] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable storage media. Fig.23 A computer system 2300 suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0223] The computer software may be encoded using any suitable machine code or computer language that may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions that may be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., either directly or through interpretation, microcode execution, etc.
[0224] The instructions may be executed on various types of computers or computer components, including, for example, personal computers, tablets, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0225] Fig.23 The components for computer system 2300 shown in the example are exemplary in nature and are not intended to imply any limitation on the scope of use or functionality of computer software implementing embodiments of the present application. Nor should the configuration of components be interpreted as having any dependency or requirement on any one or combination of components shown in the exemplary embodiment of computer system 2300.
[0226] Computer system 2300 may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input from one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, taps), visual input (e.g., gestures), olfactory input (not depicted). Human-machine interface devices may also be used to capture certain media that may not be directly related to a person's conscious input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0227] Input human interface devices may include one or more of the following (only one of each is depicted): keyboard 2301 , mouse 2302 , track pad 2303 , touch screen 2310 and associated graphics adapter 2350 , data gloves, joystick 2305 , microphone 2306 , scanner 2307 , camera 2308 .
[0228] The computer system 2300 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen 2310, a data glove, or a joystick 2305, but there may also be tactile feedback devices that do not serve as input devices), audio output devices (e.g., speakers 2309, headphones (not depicted)), visual output devices (e.g., touch screens 2310, including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light emitting diode (OLED) screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or output greater than three dimensions through, for example, stereographic output; virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)), and printers (not depicted).
[0229] The computer system 2300 may also include human-accessible storage devices and associated media for the storage devices, such as optical media, including CD / DVD ROM / RW 2320 with CD / DVD etc. media 2321, thumb drive 2322, removable hard drive or solid state drive 2323, legacy magnetic media such as tapes and floppy disks (not depicted), ROM / ASIC / PLD based specialized devices such as security protection devices (not depicted), and the like.
[0230] Those skilled in the art will also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.
[0231] The computer system 2300 may also include an interface to one or more communication networks. The network 2355 may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan, vehicle-mounted and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including global mobile communication systems (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), long term evolution (LTE), etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicle-mounted networks and industrial networks including CAN buses, etc. Some networks typically require an external network interface 2354 attached to some general data ports or peripheral buses 2349 (e.g., a USB port of the computer system 2300); other networks are typically integrated into the core of the computer system 2300 by attaching to the system bus as described below (e.g., integrated into a PC computer system via an Ethernet interface, or integrated into a smart phone computer system via a cellular network interface). As an example, the network 2355 can be connected to the peripheral bus 2349 using the external network interface 2354. By using any of these networks, the computer system 2300 can communicate with other entities. Such communication can be one-way receive only (such as broadcast TV), one-way send only (such as a CAN bus connected to certain CAN bus devices), or bidirectional, such as connecting to other computer systems using a local area digital network or a wide area digital network. Certain protocols and protocol stacks can be used on each of those networks and external network interfaces (2354) as described above.
[0232] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the core 2340 of the computer system 2300 .
[0233] The core 2340 may include one or more central processing units (CPUs) 2341, graphics processing units (GPUs) 2342, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 2343, hardware accelerators for certain tasks 2344, and the like. These devices, along with read-only memory (ROM) 2345, random access memory 2346, and internal mass storage devices 2347 such as internal non-user accessible hard disk drives, solid-state drives (SSDs), etc., may be connected via a system bus 2348. In some computer systems, the system bus 2348 may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, and the like. Peripheral devices may be attached to the core's system bus 2348 directly or via a peripheral bus 2349. Architectures for peripheral buses include peripheral component interconnects (PCI), USB, and the like.
[0234] The CPU 2341, GPU 2342, FPGA 2343, and accelerator 2344 may execute certain instructions, which in combination may constitute the above-mentioned computer code. The computer code may be stored in ROM 2345 or RAM 2346. Transitional data may also be stored in RAM 2346, while permanent data may be stored, for example, in an internal mass storage device 2347. Fast storage and retrieval of any memory device may be achieved by using a cache memory, which may be closely associated with one or more CPUs 2341, GPUs 2342, internal mass storage devices 2347, ROM 2345, RAM 2346, etc.
[0235] The computer readable medium may have computer codes for performing various computer-implemented operations. The media and computer codes may be media and computer codes designed and constructed specifically for the purpose of this application, or may be of a type well known and available to technicians in the computer software field.
[0236] By way of example and not limitation, a computer system having architecture 2300 and in particular core 2340 may provide functionality resulting from a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be user-accessible mass storage devices as described above, as well as certain storage devices of a non-transitory nature of core 2340 (e.g., internal mass storage device 2347 or ROM). 2345 associated media. Software implementing various embodiments of the present application may be stored in such devices and executed by core 2340. Depending on specific needs, computer-readable media may include one or more memory devices or chips. Software may enable core 2340 and specifically the processors therein (including CPU, GPU, FPGA, etc.) to perform specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM 2346 and modifying such data structures according to processes defined by software. In addition or as an alternative, a computer system may provide functions generated by logic hardwired or otherwise embodied in a circuit (e.g., accelerator 2344), which may replace or operate together with software to perform specific processes or specific parts of specific processes described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits embodying logic for execution, or both circuits. The present application covers any suitable combination of hardware and software.
[0237] Although the present application describes several exemplary embodiments, various modifications, permutations and combinations, and various alternative equivalents are possible within the scope of the present application. Therefore, it should be understood that within the spirit and scope of the application, those skilled in the art can design various systems and methods that are not explicitly shown or described herein but can embody the principles of the present application.
Claims
1. A method for decoding an encoded video bitstream, characterized in that, the method comprises: obtaining a coded video sequence (CVS) from the encoded video bitstream, the CVS including picture units corresponding to coded pictures; obtaining a picture header (PH) network abstraction layer (NAL) unit included in the picture unit; obtaining at least one video coding layer (VCL) NAL unit included in the picture unit; decoding the coded picture based on the picture header PH network abstraction layer NAL unit, the at least one VCL NAL unit, and an adaptive parameter set (APS), to obtain a decoded picture, wherein the APS is included in an APS NAL unit obtained from the CVS, and the APS NAL unit is available for the decoding earlier than the at least one VCL NAL unit; and outputting the decoded picture; wherein, before being referenced by one or more picture headers PH or one or more coded slice NAL units, the APS NAL unit is available for the decoding, wherein the APS NAL unit is included in at least one prediction unit (PU), and the nuh_layer_id value of the at least one PU is equal to the lowest nuh_layer_id value of the one or more coded slice NAL units that reference the APS NAL unit in the CVS.
2. The method according to claim 1, wherein, the CVS includes the one or more PHs that reference the APS or the one or more coded slice NAL units.
3. The method according to claim 1, wherein, the temporal identifier of the at least one VCL NAL unit is greater than or equal to the temporal identifier of the APS NAL unit.
4. The method according to claim 1, wherein, the picture order count (POC) of the at least one VCL NAL unit is greater than or equal to the POC of the APS NAL unit.
5. The method according to claim 1, wherein, the layer identifier of the picture header PH network abstraction layer NAL unit and the layer identifier of the at least one VCL NAL unit are greater than or equal to the layer identifier of the APS NAL unit.
6. The method according to claim 1, wherein, the picture header PH network abstraction layer NAL unit, the at least one VCL NAL unit, and the APS NAL unit are included in a single access unit.
7. The method according to any one of claims 1 - 6, wherein, the encoded video sequence further includes a stepwise temporal sub-layer access (STSA) NAL unit corresponding to an STSA picture, and wherein the STSA NAL unit is not located between the APS NAL unit and the at least one VCL NAL unit.
8. The method according to claim 7, wherein, the picture header PH network abstraction layer NAL unit, the at least one VCL NAL unit, the APS NAL unit, and the STSA NAL unit are included in a single access unit.
9. The method according to claim 7, wherein, the time identifier of the APS NAL unit is greater than or equal to the time identifier of the STSA NAL unit.
10. The method according to claim 7, wherein, the picture order count (POC) of the APS NAL unit is greater than or equal to the POC of the STSA NAL unit.
11. A method for video coding, characterized in that, the method comprises: obtaining video data corresponding to a picture unit; generating a picture header (PH) network abstraction layer (NAL) unit corresponding to the picture unit; generating at least one video coding layer (VCL) NAL unit corresponding to the picture unit; generating an encoded video sequence (CVS) based on the PH NAL unit, the at least one VCL NAL unit, and an adaptive parameter set (APS), wherein the APS is included in an APS NAL unit, and the APS NAL unit precedes the at least one VCL NAL unit for use in decoding the CVS; and wherein, before being referenced by one or more picture headers (PH) or one or more encoded slice NAL units, the APS NAL unit is available for use in decoding the CVS, wherein the APS NAL unit is included in at least one prediction unit (PU), and the nuh_layer_id value of the at least one PU is equal to the lowest nuh_layer_id value of the one or more encoded slice NAL units that reference the APS NAL unit in the CVS.
12. A method for storing or transmitting a video bitstream, characterized in that, the video bitstream is generated according to the encoding method of claim 11.
13. An electronic device, characterized in that, comprising: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code to execute the method according to any one of claims 1-12.
14. A device for decoding an encoded video bitstream, characterized in that, the device comprises: a first obtaining module configured to obtain an encoded video sequence (CVS) from the encoded video bitstream, the CVS comprising picture units corresponding to encoded pictures; a second obtaining module configured to obtain a picture header (PH) network abstraction layer (NAL) unit included in the picture unit; a third obtaining module configured to obtain at least one video coding layer (VCL) NAL unit included in the picture unit; a decoding module configured to decode the encoded picture based on the PH NAL unit, the at least one VCL NAL unit, and an adaptive parameter set (APS) to obtain a decoded picture, wherein the APS is included in an APS NAL unit obtained from the CVS, and the APS NAL unit precedes the at least one VCL NAL unit for use in the decoding; and an output module configured to output the decoded picture; Among them, before being referenced by one or more picture headers PH or one or more encoded slice NAL units, the APS NAL unit is available for decoding; Among them, the APS NAL unit is included in at least one prediction unit PU, and the value of nuh_layer_id of the at least one PU is equal to the lowest value of nuh_layer_id of the one or more encoded slice NAL units that reference the APS NAL unit in the CVS.
15. A non-volatile computer-readable medium storing instructions, Characterized in that, The instructions include: one or more instructions that, when executed by a processor, cause the one or more processors to: execute the method according to any one of claims 1-12.
Citation Information
Patent Citations
Method and apparatus for video coding and decoding
US20140218473A1
Advanced screen content coding with improved palette table and index map coding methods
US20150341643A1
Method for Palette Table Initialization and Management
US20170026641A1
Method for decoding a video bitstream
US20170324981A1