Method and apparatus for decoding an encoded video bitstream
By signaling sub-picture identifiers in the picture parameter set or sequence parameter set, the problem of low video encoding and decoding efficiency in the prior art is solved, and efficient decoding under adaptive resolution changes is achieved.
Patent Information
- Application Number
- CN202011606590.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-27
- Filing Date
- 2020-12-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-12-30
AI Technical Summary
Existing video encoding and decoding technologies are not efficient when processing image size changes, especially when adaptive resolution changes, and cannot efficiently decode encoded video bitstreams.
By determining the way that sub-picture identifiers are signaled in the picture parameter set (PPS) or sequence parameter set (SPS), different syntax elements are used for decoding, improving the efficiency of video encoding and decoding.
It realizes more efficient decoding of encoded video bitstreams when adaptive resolution changes, improving the overall efficiency of video encoding and decoding.
Smart Images

Figure CN113453011B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority of U.S. Provisional Application No. 63 / 001,087, filed on March 27, 2020, and U.S. Application No. 17 / 081,135, filed on October 27, 2020, the entire contents of which are incorporated herein by reference. Technical field
[0003] This application relates to video encoding and decoding, and more particularly, to methods and devices, apparatuses, and computer - readable media for decoding an encoded video bitstream. Background art
[0004] It is well - known to use inter - picture prediction with motion compensation for video encoding and decoding. Uncompressed digital video may include a series of pictures, each picture having a spatial dimension such as 1920×1080 luminance samples and associated chrominance samples. The series of pictures has a fixed or variable picture rate (also informally called the frame rate), such as 60 pictures per second or 60 Hz. Uncompressed video has very high bit - rate requirements. For example, a 1080p60 4:2:0 video with 8 bits per sample (1920x1080 luminance sample resolution, 60 Hz frame rate) requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video would require more than 600 GB of storage space.
[0005] One purpose of video encoding and decoding is to reduce redundant information in the input video signal through compression. Video compression can help reduce the requirements for the above - mentioned bandwidth or storage space, and in some cases can reduce by two or more orders of magnitude. Both lossless and lossy compression, as well as combinations of the two, can be employed. Lossless compression is a technique for reconstructing an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal can be used for the intended application. Lossy compression is widely used in video. The amount of allowable distortion depends on the application. For example, users of some consumer streaming applications can tolerate higher distortion compared to users of television applications. The achievable compression ratio reflects that higher allowed / tolerated distortion can result in a higher compression ratio.
[0006] Video encoders and decoders can utilize several major categories of techniques, such as including motion compensation, transformation, quantization, and entropy coding, some of which will be introduced below.
[0007] Historically, video encoders and decoders have tended to operate on a given picture size which, in most cases, is defined and kept constant for an encoded video sequence (CVS), a group of pictures (GOP), or a similar multi-picture time frame. For example, in MPEG-2, it is known that the system design changes the horizontal resolution (and thus the picture size) only at I pictures depending on factors such as scene activity, and thus is typically used for GOPs. For example, according to Appendix P of ITU-T Rec. H.263, it is known to resample reference pictures at different resolutions within a CVS. However, the picture size here does not change, only the reference pictures are resampled, which may result in only part of the picture canvas being used (in the case of downsampling), or only part of the scene being captured (in the case of upsampling). Further, Appendix Q of H.263 allows resampling of individual macroblocks by a factor of two up or down (in each dimension). Again, the picture size remains unchanged. The size of the macroblocks is fixed in H.263 and thus does not need to be signaled.
[0008] In modern video coding and decoding, predicting changes in picture size in a predicted picture has become mainstream. For example, VP9 allows reference picture resampling and changing the resolution of the entire picture. Similarly, certain proposals for Versatile Video Coding (VVC) (including, for example, "On adaptive resolution change (ARC) for VVC" by Hendry et al., Joint Video Team document JVET-M0135-v1, January 9 - 19, 2019, which is incorporated herein by reference in its entirety) allow resampling the entire reference picture to a different (higher or lower) resolution. In that document, it is proposed to encode different candidate resolutions in the sequence parameter set and reference them by each picture syntax element in the picture parameter set. SUMMARY OF THE INVENTION
[0009] Embodiments of the present application provide a method and device, apparatus, and computer-readable medium for decoding an encoded video bitstream, aiming to solve the problem of low video coding and decoding efficiency in the prior art.
[0010] In an embodiment, a method for decoding an encoded video bitstream is provided, including: obtaining a first flag, where the first flag indicates that a sub-picture identifier of a sub-picture of a current picture is signaled explicitly; based on the first flag, obtaining a second flag, where the second flag indicates whether the sub-picture identifier is signaled in a picture parameter set (PPS) or in a sequence parameter set (SPS); when the second flag indicates that the sub-picture identifier is signaled in the picture parameter set (PPS), determining the sub-picture identifier based on a first syntax element included in the picture parameter set (PPS); when the second flag indicates that the sub-picture identifier is signaled in the sequence parameter set (SPS), determining the sub-picture identifier based on a second syntax element included in the sequence parameter set (SPS); and decoding the current picture based on the determined sub-picture identifier.
[0011] In an embodiment, a device for decoding an encoded video bitstream is provided, including: at least one memory for storing program code; and at least one processor for reading the program code and operating according to the instructions of the program code to execute the method for decoding an encoded video bitstream in the embodiment.
[0012] In an embodiment, a device for decoding an encoded video bitstream is provided, including: a first obtaining module for obtaining a first flag, where the first flag indicates that a sub-picture identifier of a sub-picture of a current picture is signaled explicitly; a second obtaining module for, based on the first flag, obtaining a second flag, where the second flag indicates whether the sub-picture identifier is signaled in a picture parameter set PPS or in a sequence parameter set SPS; a first determining module for, when the second flag indicates that the sub-picture identifier is signaled in the picture parameter set PPS, determining the sub-picture identifier based on a first syntax element included in the picture parameter set PPS; a second determining module for, when the second flag indicates that the sub-picture identifier is signaled in the sequence parameter set SPS, determining the sub-picture identifier based on a second syntax element included in the sequence parameter set SPS; and a decoding module for decoding the current picture based on the determined sub-picture identifier.
[0013] In an embodiment, a non-volatile computer-readable medium is provided for storing instructions, where the instructions include one or more instructions that, when executed by one or more processors of a device for decoding an encoded video bitstream, cause the one or more processors to execute the method for decoding an encoded video bitstream in the embodiment.
[0014] In an embodiment of the present application, by determining whether to signal the sub-picture identifier in the PPS or in the SPS, the sub-picture identifier can be determined based on different syntax elements, and the current picture can be decoded based on the determined sub-picture identifier, thereby improving the efficiency of video encoding and decoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The other features, properties, and various advantages of the present application will become further apparent from the following detailed description and the accompanying drawings, in which
[0016] Figure 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment;
[0017] Figure 2 is a schematic diagram of a simplified block diagram of a communication system according to another embodiment;
[0018] Figure 3 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment;
[0019] Figure 4 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment;
[0020] Figures 5A - 5E is a schematic diagram of an option for signaling ARC parameters according to an embodiment;
[0021] Figures 6A - 6B is a schematic diagram of an example of a syntax table according to an embodiment;
[0022] Figure 7 is an example of a prediction structure with scalability having adaptive resolution change according to an embodiment;
[0023] Figure 8 is an example of a syntax table according to an embodiment;
[0024] Figure 9 is a schematic diagram of a simplified block diagram for parsing and decoding the POC cycle and access unit count value of each access unit according to an embodiment;
[0025] Figure 10 is a schematic diagram of a video bitstream structure including multiple layers of sub-pictures according to an embodiment;
[0026] Figure 11 is a schematic diagram of the display of a selected sub-picture with enhanced resolution according to an embodiment;
[0027] Figure 12 is a block diagram of the decoding and display process of a video bitstream including multiple layers of sub-pictures according to an embodiment;
[0028] Figure 13Schematic diagram of 360 video display with enhanced layer having sub - pictures according to an embodiment;
[0029] Figure 14 Example of layout information of sub - pictures and its corresponding layer and picture prediction structure according to an embodiment;
[0030] Figure 15 Example of layout information of sub - pictures and its corresponding layer and picture prediction structure of spatial scalability mode with local region according to an embodiment;
[0031] Figures 16A - 16B Example of syntax table for sub - picture layout information according to an embodiment;
[0032] Figure 17 Example of syntax table of SEI message for sub - picture layout information according to an embodiment;
[0033] Figure 18 Example of syntax table according to an embodiment, which is used to indicate profile / layer / level information of output layer and each output layer set;
[0034] Figure 19 Example of syntax table according to an embodiment, which is used to indicate output layer mode enabled for each output layer set;
[0035] Figure 20 Example of syntax table according to an embodiment, which is used to indicate current sub - pictures of each layer of each output layer set;
[0036] Figure 21 Example of syntax table including a part of picture parameter set according to an embodiment;
[0037] Figure 22 Schematic diagram of an example process of decoding an encoded video bitstream according to an embodiment;
[0038] Figure 23 Schematic diagram of a computer system according to an embodiment. Detailed implementation manners
[0039] Figure 1is a simplified block diagram of a communication system (100) according to an embodiment of the present application. The communication system (100) includes at least two terminal devices (110, 120), and the terminal devices can communicate with each other through a network (150). For unidirectional data transmission, the first terminal device (110) can encode video data at a local location for transmission through the network (150) to the second terminal device (120). The second terminal device (120) can receive the encoded video data of another terminal from the network (150), decode the encoded video data to recover the video data, and display the recovered video data. Unidirectional data transmission is more common in applications such as media services.
[0040] Figure 1 Shows a second pair of terminal devices (130, 140) that support two-way transmission of encoded video, and the two-way transmission can occur, for example, during a video conference. For two-way data transmission, each of the third terminal device (130) and the fourth terminal device (140) can encode video data collected at a local location for transmission through the network (150) to the other of the third terminal device (130) and the fourth terminal device (140). Each of the third terminal device (130) and the fourth terminal device (140) can also receive the encoded video data transmitted by the other of the third terminal device (130) and the fourth terminal device (140), and can decode the encoded video data and display the recovered video data on a local display device.
[0041] In Figure 1 the first terminal device (110), the second terminal device (120), the third terminal device (130), and the fourth terminal device (140) can be servers, personal computers, and smart phones, but the principles disclosed in the present application are not limited thereto. The embodiments disclosed in the present application are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (150) represents any number of networks that transfer encoded video data between the first terminal device (110), the second terminal device (120), the third terminal device (130), and the fourth terminal device (140), including, for example, wired and / or wireless communication networks. The communication network (150) can exchange data in circuit-switched and / or packet-switched channels. The network can include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of the present application, unless otherwise explained below, the architecture and topology of the network (150) may be irrelevant to the operations disclosed in the present application.
[0042] As an application embodiment of the present application, Figure 2Shows the placement of video decoders and encoders in a streaming environment. This application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0043] A streaming system (200) may include an acquisition subsystem (213), which may include a video source (201) such as a digital camera that creates, for example, an uncompressed video sample stream (202). The video sample stream (202) is depicted as a thick line as compared to the encoded video bitstream to emphasize that it is a high-data volume video sample stream. The video sample stream (202) may be processed by an encoder (203) coupled to the camera (201). The encoder (203) may include hardware, software, or a combination of both to implement or carry out aspects of the present application described in more detail below. The encoded video bitstream (204) is depicted as a thin line as compared to the video sample stream (202) to emphasize the lower-data volume encoded video bitstream, which may be stored on a streaming server (205) for future use. One or more streaming clients (206, 208) may access the streaming server (205) to retrieve copies (207) and (209) of the encoded video bitstream (204). Client (206) may include a video decoder (210). The video decoder (210) decodes an incoming copy (207) of the encoded video bitstream and produces an output video sample stream (211) that may be presented on a display (212) or another rendering device (not shown). In some streaming systems, the video bitstreams (204, 207, 209) may be encoded according to certain video codec / compression standards. Examples of such standards include ITU-T Recommendation H.265. A video codec standard under development is informally referred to as Versatile Video Coding (VVC). The present application may be used in the context of VVC.
[0044] Figure 3 Is a functional block diagram of a video decoder (210) according to an embodiment of the present application.
[0045] A receiver (310) may receive one or more encoded video sequences to be decoded by a video decoder (210); in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel (312), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data as well as other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not shown). The receiver (310) may separate the encoded video sequences from the other data. To guard against network jitter, a buffer memory (315) may be coupled between the receiver (310) and an entropy decoder / parser (320) (hereinafter referred to as "parser"). When the receiver (310) receives data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, it may also be unnecessary to configure the buffer memory (315), or the buffer memory may be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (315) may also be required, which may be relatively large and may have an adaptive size.
[0046] The video decoder (210) may include a parser (320) to reconstruct symbols (321) from the entropy-encoded video sequences. The classes of these symbols include information for managing the operation of the video decoder (210), as well as potential information for controlling a display device such as the display 212, which is not part of the decoder but may be coupled to the decoder, as Figure 3As shown in []. The control information for the display device can be a parameter set segment (not labeled) of Supplementary Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (320) can parse and / or entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (320) can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroup can include Group of Pictures (GOP), picture, sub-picture, tile, slice, brick, macroblock, Coding Tree Unit (CTU), Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. A tile can indicate a rectangular area of CUs / CTUs within a specific tile column and row in a picture. A brick can indicate a rectangular area of CU / CTU rows within a specific tile. A slice can indicate one or more bricks in a picture, which are contained in a NAL unit. A sub-picture can indicate a rectangular area of one or more slices in a picture. The entropy decoder / parser can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0047] The parser (320) can perform entropy decoding and / or parsing operations on the video sequence received from the buffer memory (315) to create symbols (321).
[0048] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (321) can involve multiple different units. Which units are involved and the way they are involved can be controlled by the subgroup control information parsed by the parser (320) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (320) and multiple units below are not described.
[0049] In addition to the functional blocks already mentioned, the video decoder (210) can conceptually be subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purposes of this application, it is appropriate to conceptually subdivide into the functional units below.
[0050] The first unit can be a scaler and / or inverse transform unit (351). The scaler and / or inverse transform unit (351) can receive quantized transform coefficients as symbols (321) and control information from the parser (320), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (351) can output a block including sample values, and the sample values can be input into the aggregator (355).
[0051] In some cases, the output samples of the scaler and / or inverse transform unit (351) can belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed pictures but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed information extracted from the (partially reconstructed) current picture (358). In some cases, the aggregator (355) adds the predictive information generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler and / or inverse transform unit (351) based on each sample.
[0052] In other cases, the output samples of the scaler and / or inverse transform unit (351) can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit (353) can access the reference picture memory (357) to extract samples for prediction. After motion-compensating the extracted samples according to the symbol (321), these samples can be added by the aggregator (355) to the output of the scaler and / or inverse transform unit (351) (which is called the residual sample or residual signal in this case), thereby generating output sample information. The motion compensation prediction unit (353) obtaining the prediction samples from the address in the reference picture memory (357) can be controlled by a motion vector, and the motion vector is in the form of the symbol (321) for use by the motion compensation prediction unit (353), and the symbol (321) includes, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory (357), a motion vector prediction mechanism, etc. when using sub-sample accurate motion vectors.
[0053] The output samples of the aggregator (355) can be employed by various loop filtering techniques in the loop filter unit (356). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream, and the parameters can be available to the loop filter unit (356) as symbols (321) from the parser (320). However, in other embodiments, video compression techniques can also respond to meta-information obtained during the decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0054] The output of the loop filter unit (356) can be a sample stream that can be output to the display device (212) and stored in the reference picture memory (357) for subsequent inter-picture prediction.
[0055] Once fully reconstructed, certain encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture is identified (e.g., by the parser (320)) as a reference picture, the current picture (358) can become part of the reference picture memory (357), and a new current picture memory can be reallocated before starting the reconstruction of subsequent encoded pictures.
[0056] The video decoder (210) can perform decoding operations according to, for example, the predefined video compression techniques recorded in the ITU-T H.265 standard. The encoded video sequence can conform to the syntax specified by the video compression technique or standard used in the sense that the encoded video sequence follows the syntax of the video compression technique or standard, particularly the profile, specified in the video compression technique literature or standard. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, and so on. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.
[0057] In an embodiment, the receiver (310) may receive additional (redundant) data together with the encoded video. The additional data may be part of an encoded video sequence. The additional data may be used by the video decoder (210) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.
[0058] Figure 4 May be a functional block diagram of a video encoder (203) according to an embodiment of the present application.
[0059] The video encoder (203) may receive video samples from a video source (201) (not part of the decoder), and the video source may capture video images to be encoded by the video encoder (203).
[0060] The video source (201) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (203), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 Y CrCb, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (201) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (201) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed in sequence. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on the description of samples.
[0061] According to an embodiment, the video encoder (203) may encode and compress pictures of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed is a function of the controller (450). The controller (450) controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, the couplings are not labeled in the figures. Parameters set by the controller (450) may include rate control related parameters (e.g., picture skipping, quantizer, λ value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller (450) as they may be related to the video encoder (203) optimized for a specific system design.
[0062] Some video encoders operate in a manner of an "encoding loop" that is easily understood by those skilled in the art. As a simple description, the encoding loop may include an encoding part of an encoder (430) (hereinafter referred to as a "source coder" which is responsible for creating symbols based on an input picture to be encoded and reference pictures), and a (local) decoder (433) embedded in the video encoder (203). The "local" decoder (433) reconstructs symbols in a manner similar to how a (remote) decoder creates sample data to create sample data (since in the video compression techniques considered in this application, any compression between symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input into the reference picture memory (434). Since the decoding of the symbol stream produces a bit-exact result regardless of the decoder location (local or remote), the content in the reference picture memory is also bit-exactly corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs in cases where synchronization cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.
[0063] The operation of the "local" decoder (433) may be the same as that of the "remote" decoder that has been described in detail above in connection with Figure 3 the video decoder (210). However, briefly referring additionally to Figure 4 , when symbols are available and the entropy encoder (445) and the parser (320) can encode and / or decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the video decoder (210) (including the channel (312), the receiver (310), the buffer memory (315), and the parser (320)) may not be fully implemented in the local decoder (433).
[0064] At this point, it can be observed that any decoder technology other than parsing and / or entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, the present application focuses on decoder operations. The description of encoder technology can be abbreviated because they are inverse to the fully described decoder technology. A more detailed description is only needed in certain areas and is provided below.
[0065] As part of the operation, the source encoder (430) may perform motion compensation predictive coding. Referring to one or more previously encoded frames designated as "reference frames" in the video sequence, the motion compensation predictive coding performs predictive coding on the input frame. In this way, the coding engine (432) encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, and the reference frame can be selected as the prediction reference for the input frame.
[0066] The local video decoder (433) may decode the encoded video data of the frame that can be designated as a reference frame based on the symbols created by the source encoder (430). The operation of the coding engine (432) may be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 4 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder (433) replicates the decoding process that can be performed by the video decoder on the reference frame and can store the reconstructed reference frame in the reference picture memory (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame, which has the same content (in the absence of transmission errors) as the reconstructed reference frame to be obtained by the remote video decoder.
[0067] The predictor (435) may perform a prediction search for the coding engine (432). That is, for a new frame to be encoded, the predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or some metadata that can be used as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor (435) may operate on a per-pixel basis of the sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (435), it can be determined that the input picture may have a prediction reference obtained from multiple reference pictures stored in the reference picture memory (434).
[0068] The controller (450) may manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.
[0069] The outputs of all the above functional units can be entropy encoded in an entropy encoder (445). The entropy encoder (445) can perform lossless compression on the symbols generated by various functional units according to techniques well known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., so as to convert the symbols into an encoded video sequence.
[0070] The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) to prepare for transmission through a communication channel (460), which can be a hardware / software link leading to a storage device that will store the encoded video data. The transmitter (440) can merge the encoded video data from the source encoder (430) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0071] The controller (450) can manage the operation of the video encoder (203). During encoding, the controller (450) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, a picture can generally be assigned to any of the following picture types.
[0072] An intra picture (I picture), which can be a picture that can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, an Independent Decoder Refresh (IDR) picture. Those skilled in the art know those variants of the I picture and their respective applications and characteristics.
[0073] A predictive picture (P picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, and the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0074] A bi - predictive picture (B picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, and the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0075] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, and the other blocks are determined according to the coding assignments of the corresponding pictures applied to the blocks. For example, blocks of an I picture can be non-predictively encoded, or the blocks can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.
[0076] The video encoder (203) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (203) can perform various compression operations, including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0077] In an embodiment, the transmitter (440) can transmit additional data and the encoded video. The source encoder (430) can include such data as part of, for example, an encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0078] Recently, aggregating or extracting multiple semantically independent picture parts in the compression domain into a single video picture has drawn some attention. Specifically, in the context of, for example, 360 codec or certain surveillance applications, multiple semantically independent source pictures (e.g., the six cube surfaces of a cube-projected 360 scene, or the inputs of a single camera in the case of a multi-camera surveillance setup) may require separate adaptive resolution settings to handle different scene activities at a given point in time. In other words, the encoder can choose to use different resampling factors for different semantically independent pictures that make up the entire 360 or surveillance scene at a given point in time. When combined into a single picture, this in turn requires reference picture resampling to be performed on a portion of the encoded pictures and requires adaptive resolution codec signaling to be available.
[0079] Below, some terms will be introduced that will be referenced in the remainder of this specification.
[0080] In some cases, a sub - picture refers to a rectangular arrangement of samples, blocks, macro - blocks, coding units, or similar entities that are semantically grouped and can be independently coded at varying resolutions. One or more sub - pictures can form a picture. One or more coded sub - pictures can form a coded picture. One or more sub - pictures can be combined into a picture, and one or more sub - pictures can be extracted from a picture. In certain environments, one or more coded sub - pictures can be assembled in the compressed domain without transcoding to the sample level to form a coded picture, and in the same or other cases, one or more coded sub - pictures can be extracted from a coded picture in the compressed domain.
[0081] Adaptive Resolution Change (ARC) refers to a mechanism that allows changing the resolution of a picture or sub - picture in a coded video sequence, for example, by resampling a reference picture. ARC parameters hereinafter refer to the control information required to perform adaptive resolution change, which can include, for example, filter parameters, scaling factors, the resolution of the output and / or reference pictures, various control flags, etc.
[0082] In an embodiment, encoding and decoding can be performed on a single, semantically independent coded video picture. Before describing the meaning and the implied additional complexity of encoding / decoding multiple sub - pictures with independent ARC parameters, the options for signaling ARC parameters will be described.
[0083] Reference Figures 5A - 5E , which shows several embodiments for signaling ARC parameters. As indicated by each embodiment, these embodiments have certain advantages and disadvantages from the perspectives of codec efficiency, complexity, and architecture. A video codec standard or technology can select one or more of these embodiments, or select options known in the related art, to signal ARC parameters. These embodiments are not mutually exclusive and can be interchanged based on application requirements, the standard technology involved, or the choice of encoder.
[0084] The categories of ARC parameters can include:
[0085] - Upsampling factor and / or downsampling factor, which are separated or combined in the X and Y dimensions;
[0086] - Upsampling factor and / or downsampling factor, which, when combined with the temporal dimension, represent the zoom - in / zoom - out at a constant rate for a given number of pictures;
[0087] - Either of the above may involve encoding and decoding of one or more short syntax elements, which may point to a table containing one or more factors.
[0088] - The resolution of the input picture, output picture, reference picture, and coded picture, combined or separate, in the X or Y dimension, in units of samples, blocks, macroblocks, coding units (CUs), or any other suitable granularity. If there are more than one resolution (e.g., one for the input picture and another for the reference picture), in some cases, one set of values can be inferred from the other set. This can be controlled, for example, by using flags. For more detailed examples, see below.
[0089] - "Warping" coordinates are similar to those used in H.263 Annex P and are also in the above-mentioned suitable granularity. H.263 Annex P defines an efficient way to encode and decode such warping coordinates, but other potentially more efficient ways can also be envisioned. For example, the variable-length reversible "Huffman" encoding and decoding of the warping coordinates in Annex P can be replaced by binary encoding and decoding of a suitable length, where the length of the binary codeword can be derived, for example, from the maximum picture size, possibly multiplied by a certain factor and offset by a certain value to allow "warping" outside the boundaries of the maximum picture size.
[0090] - Upsampling filter parameters and / or downsampling filter parameters. In an embodiment, there may be only a single filter for upsampling and / or downsampling. However, in an embodiment, it may be desirable to allow more flexibility in filter design, which may require signaling the filter parameters. Such parameters can be selected by an index in a list of possible filter designs, the filter can be fully specified (e.g., by a list of filter coefficients, using suitable entropy coding techniques), or the filter can be implicitly selected by the upsampling and / or downsampling ratio and then signaled sequentially according to any of the above mechanisms.
[0091] Hereafter, this specification assumes the encoding and decoding of a finite set of upsampling factors and / or downsampling factors indicated by a codeword (using the same factors in the X and Y dimensions). The codeword can be variable-length encoded, for example, using the Ext-Golomb code common to certain syntax elements in video coding specifications (such as H.264 and H.265). For example, according to Table 1, values can be appropriately mapped to upsampling factors and / or downsampling factors.
[0092] Table 1
[0093] Codeword Ext - Golomb Code Original / Target Resolution 0 1 1 / 1 1 010 1 / 1.5 (50% Magnification) 2 011 1.5 / 1 (50% Reduction) 3 00100 1 / 2 (100% Magnification) 4 00101 2 / 1 (100% Reduction)
[0094] Many similar mappings can be designed according to the requirements of the application and the functionality of the upscaling and downscaling mechanisms available in video compression techniques or standards. The table can be extended to more values. These values can also be represented by entropy coding and decoding mechanisms other than the Ext-Golomb code, such as using binary coding. This may have certain advantages when the resampling factor is of concern outside the video processing engine (most importantly, the encoder and decoder) itself, e.g., via the MANE. It should be noted that for cases where the resolution does not need to be changed, a shorter Ext-Golomb code can be selected. In the above table, there is only 1 bit. This can have the advantage of improving the coding and decoding efficiency compared to using binary codes in most cases.
[0095] The number of entries in the table and their semantics can be fully or partially configurable. For example, the basic outline of the table can be conveyed in a "high" parameter set such as a sequence or decoder parameter set. In an embodiment, one or more such tables can be defined in a video coding and decoding technique or standard and can be selected via, for example, a decoder or sequence parameter set.
[0096] The following describes how to include the encoded upsampling factor and / or downsampling factor (ARC information) as described above in the syntax of a video coding and decoding technique or standard. Similar considerations can be applied to one or several codewords that control the upsampling filter and / or downsampling filter. When the filter or other data structure requires a relatively large amount of data, see the discussion below.
[0097] As Figure 5A shown, H.263 Annex P includes the ARC information (502) in the picture header (501) in the form of four warping coordinates, specifically in the H.263PLUSPTYPE (503) header extension. This may be a reasonable design choice when there is an available picture header and the ARC information is expected to change frequently. However, when using H.263 format signaling, the overhead may be very high, and since the picture header may have an instantaneous nature, the scaling factor may not fall within the picture boundaries.
[0098] As Figure 5BAs shown, JVCET-M135-v1 includes ARC reference information (505) (index) located in the picture parameter set (504), and this ARC reference information (505) indexes a table (506) including the target resolution, and this table (506) is located within the sequence parameter set (507). By using the SPS as an interoperability negotiation point during capability exchange, the placement of possible resolutions in the table (506) of the sequence parameter set (507) can be demonstrated. By referring to the appropriate picture parameter set (504), the resolution can vary with the picture within the limits set by the values in the table (506).
[0099] Reference Figures 5C - 5E , there can be the following embodiments for transmitting ARC information in a video bitstream. Each of these options has certain advantages over the above embodiments. These embodiments can exist simultaneously in the same video coding technology or standard.
[0100] In an embodiment, for example Figure 5C in the embodiment shown, ARC information (509) such as a resampling (scaling) factor can exist in a slice header, GOP header, tile header, or tile group header. Figure 5C An embodiment using a tile group header (508) is shown. If the ARC information is small, such as a single variable length ue(v) or a fixed length codeword of several bits as shown above, this is sufficient. Placing the ARC information directly in the tile group header has the additional advantage that the ARC information can be applied to, for example, the sub-picture represented by the tile group rather than the entire picture, see also below. Additionally, even if the video compression technology or standard only contemplates the change of adaptive resolution of the entire picture (e.g., compared with tile group-based adaptive resolution change), from the perspective of error recovery, placing the ARC information in the tile group header has certain advantages compared to placing it in the picture header in the H.263 format.
[0101] In an embodiment, for example Figure 5D in the embodiment shown, the ARC information (512) itself can exist in an appropriate parameter set, such as a picture parameter set, header parameter set, tile parameter set, adaptive parameter set, etc. Figure 5D An embodiment using an adaptive parameter set (511) is shown. More advantageously, the scope of this parameter set can be no larger than a picture, such as a tile group. By activating the relevant parameter set, the use of the ARC information is implicit. For example, when the video coding technology or standard only considers picture-based ARC, then the picture parameter set or equivalent parameters may be applicable.
[0102] In an embodiment, for example Figure 5EIn the illustrated embodiment, the ARC reference information (513) may be present in the slice group header (514) or a similar data structure. This reference information (513) may refer to a subset of the ARC information (515) available in a parameter set (516) whose scope extends beyond a single picture, such as a sequence parameter set or a decoder parameter set.
[0103] As Figure 6A shown, as an exemplary syntax structure of a header that may apply to a (possibly rectangular) portion of a picture, the slice group header (601) may conditionally include a variable-length, Exp-Golomb coded syntax element Dec_pic_size_idx (602) (shown in bold). Adaptive resolution (603) may be used to control the presence of this syntax element in the slice group header (here, the value of the flag is not shown in bold), which means that the flag appears in the bitstream at the point where it appears in the syntax diagram. Whether the adaptive resolution is applied to the picture or a portion of the picture may be signaled in any high-level syntax structure, either inside or outside the bitstream. In the illustrated example, it is signaled in the sequence parameter set as described below.
[0104] Referring Figure 6B to it also shows an excerpt of the sequence parameter set (610). The first syntax element shown is the adaptive_pic_resolution_change_flag (611). When true, this flag may indicate the use of adaptive resolution, which may in turn require some control information. In this example, the value of this flag is based on an if statement in the parameter set (612) and the slice group header (601), and this control information appears conditionally based on the value of this flag.
[0105] In this example, when adaptive resolution is used, the output resolution in samples (613) is encoded. The label 613 refers to output_pic_width_in_luma_sample and output_pic_height_in_luma_sample, which together define the resolution of the output picture. Elsewhere in the video coding and decoding technology or standard, certain restrictions on either value may be defined. For example, the level definition may limit the total number of output samples, which may be the product of the values of these two syntax elements. Additionally, certain video coding and decoding technologies or standards, or external technologies or standards (such as system standards), may limit the number range (e.g., one or both dimensions must be divisible by an integer power of 2) or the aspect ratio (e.g., the width and height must have a relationship such as 4:3 or 16:9). Such restrictions may be introduced to facilitate hardware implementation or for other reasons.
[0106] In some applications, it is recommended that the encoder indicate to the decoder to use a certain reference picture size instead of implicitly assuming that size to be the output picture size. In this example, the syntax element reference_pic_size_present_flag(614) controls the conditional presence of the reference picture size(615) (again, this label refers to both the width and the height).
[0107] Finally, a table of possible decoded picture widths and heights is shown. Such a table can be represented, for example, by the table indication (num_dec_pic_size_in_luma_samples_minus1)(616). "minus1" refers to the interpretation of the value of this syntax element. For example, if the encoded value is 0, there is one table entry. If the value is 5, there are six table entries. For each "line" in the table, the width and height of the decoded picture are included in the syntax(617).
[0108] The syntax element dec_pic_size_idx(602) in the tile group header can be used to index the presented table entries(617), allowing each tile group to have a different decoded size (effectively a scaling factor).
[0109] Certain forms of reference picture resampling are achieved by combining temporal scalability (signaled in a completely different manner than in this application), such that some video coding technologies or standards (e.g., VP9) support spatial scalability in order to achieve spatial scalability. Specifically, certain reference pictures can be upsampled to a higher resolution using ARC format technology to form the basis of a spatial enhancement layer. Those upsampled pictures can be refined using conventional prediction mechanisms at the higher resolution to increase detail.
[0110] The embodiments discussed herein can be used in such an environment. In some cases, in the same or another embodiment, the value in the NAL unit header (e.g., the temporal ID field) can be used not only to indicate the temporal layer but also to indicate the spatial layer. Doing so can have certain advantages for some system designs. For example, for a scalable environment, existing selective forwarding units (SFUs) created and optimized for temporal layer selected forwarding based on the temporal ID value in the NAL unit header can be used without modification. To achieve this, it may be necessary to indicate the mapping between the encoded picture size and the temporal layer through the temporal ID field in the NAL unit header.
[0111] In some video coding and decoding techniques, an access unit (AU) may refer to one or more coded pictures, one or more slices, one or more tiles, and / or one or more NAL unit bitstreams captured at a given time instance and combined into a corresponding picture, slice, tile, and / or NAL unit bitstream. For example, the given time instance may be a composition time.
[0112] In HEVC and certain other video coding and decoding techniques, a picture order count (POC) value may be used to indicate a reference picture selected from a plurality of reference pictures stored in a decoded picture buffer (DPB). When an access unit (AU) includes one or more pictures, slices, or tiles, each picture, slice, or tile belonging to the same AU may carry the same POC value, from which it can be derived that they are created based on the content of the same composition time. In other words, in the case where two pictures / slices / tiles carry the same given POC value, the POC value may indicate two pictures / slices / tiles belonging to the same AU and having the same composition time. Conversely, two pictures / slices / tiles with different POC values may indicate those pictures / slices / tiles belonging to different AUs and having different composition times.
[0113] In an embodiment, since an access unit may include pictures, slices, or tiles with different POC values, the rigid relationship can be relaxed. By allowing different POC values to be used within one AU, the POC value can be used to identify pictures / slices / tiles that may be independently decoded with the same presentation time. This can in turn support multiple scalable layers without changing the reference picture selection signaling (e.g., reference picture set signaling or reference picture list signaling), as described in more detail below.
[0114] However, for other pictures / slices / tiles with different POC values, it is still desirable to be able to identify the AU to which the pictures / slices / tiles belong based only on the POC value. This can be achieved as described below.
[0115] In an embodiment, an access unit count (AUC) may be signaled in a high-level syntax structure (e.g., NAL unit header, slice header, tile group header, SEI message, parameter set, or AU delimiter). The AUC value can be used to identify which NAL units, pictures, slices, or tiles belong to a given AU. The AUC value may correspond to different composition time instances. The AUC value may be equal to a multiple of the POC value. The AUC value can be calculated by dividing the POC value by an integer value. In some cases, the division operation may impose a certain burden on the decoder implementation. In such cases, making some minor restrictions in the numbering space of the AUC value can allow a shift operation to replace the division operation. For example, the AUC value may be equal to the most significant bit (MSB) value within the range of the POC value.
[0116] In an embodiment, the value of poc_cycle_au for each AU may be signaled in a high-level syntax structure (e.g., NAL unit header, slice header, tile group header, SEI message, parameter set, or AU delimiter). The poc_cycle_au may indicate how many different and consecutive POC values may be associated with the same AU. For example, if the value of poc_cycle_au is equal to 4, pictures, slices, or tiles with POC values equal to 0 to 3 (inclusive of 0 and 3) may be associated with the AU with AUC value equal to 0, and pictures, slices, or tiles with POC values equal to 4 to 7 (inclusive of 4 and 7) may be associated with the AU with AUC value equal to 1. Thus, the value of AUC may be inferred by dividing the POC value by the value of poc_cycle_au.
[0117] In an embodiment, the value of poc_cycle_au may be derived from information, such as that located in a video parameter set (VPS), that identifies the number of spatial or SNR layers in an encoded video sequence. An example of such a possible relationship is briefly described below. Although the derivation as described above may save a few bits in the VPS and thus may improve coding / decoding efficiency, in some embodiments, the poc_cycle_au may be explicitly encoded in a high-level syntax structure at a level lower than the video parameter set in order to be able to minimize the poc_cycle_au for a given small portion of the bitstream (e.g., a picture). Since the POC value and / or the value of a syntax element that indirectly references the POC may be encoded in a low-level syntax structure, this optimization may save more bits compared to the bits that can be saved by the above-described derivation process.
[0118] In an embodiment, Figure 8An example of a syntax table is shown. This syntax table is used to signal the syntax elements of vps_poc_cycle_au in the VPS (or SPS), where vps_poc_cycle_au indicates the poc_cycle_au for all pictures / slices in the encoded video sequence. This syntax table is also used to signal the syntax elements of slice_poc_cycle_au, where slice_poc_cycle_au indicates the poc_cycle_au of the current slice in the slice header. If the POC values of each AU increase uniformly, vps_contant_poc_cycle_per_au in the VPS can be set to 1, and vps_poc_cycle_au can be signaled in the VPS. In this case, slice_poc_cycle_au may not be signaled explicitly, and the AUC value of each AU can be calculated by dividing the POC value by vps_poc_cycle_au. If the POC values of each AU do not increase uniformly, vps_contant_poc_cycle_per_au in the VPS can be set to 0. In this case, vps_access_unit_cnt may not be signaled, but slice_access_unit_cnt can be signaled in the slice header for each slice or picture. Each slice or picture can have a different slice_access_unit_cnt value. The AUC value of each AU can be calculated by dividing the POC value by slice_poc_cycle_au.
[0119] Figure 9 A block diagram showing an example of the above process is shown. For example, in step S910, the VPS (or SPS) can be parsed, and in step S920, it can be determined whether the POC period of each AU is constant within the encoded video sequence. If the POC period of each AU is constant (yes at step S920), then in step S930, the value of the access unit count of the specific access unit can be calculated based on the poc_cycle_au signaled for the encoded video sequence and the POC value of the specific access unit. If the POC period of each AU is not constant (no at step S920), then in step S940, the value of the access unit count of the specific access unit can be calculated based on the poc_cycle_au signaled at the picture level and the POC value of the specific access unit. In step S950, a new VPS (or SPS) can be parsed.
[0120] In an embodiment, even if the POC values of pictures, slices, or tiles may be different, the pictures, slices, or tiles corresponding to an AU having the same AUC value can be associated with the same decoding or output time instance. Therefore, without any inter-picture parsing / decoding dependency among the pictures, slices, or tiles in the same AU, all or a subset of the pictures, slices, or tiles associated with the same AU can be decoded in parallel and output at the same time instance.
[0121] In an embodiment, even if the POC values of pictures, slices, or tiles may be different, the pictures, slices, or tiles corresponding to an AU having the same AUC value can be associated with the same composition / display time instance. When the composition time is included in the container format, even if pictures correspond to different AUs, but if the pictures have the same composition time, they can be displayed at the same time instance.
[0122] In an embodiment, each picture, slice, or tile can have the same temporal identifier (temporal_id) in the same AU. All or a subset of the pictures, slices, or tiles corresponding to a time instance can be associated with the same time sublayer. In an embodiment, each picture, slice, or tile can have the same or different spatial layer id (layer_id) in the same AU. All or a subset of the pictures, slices, or tiles corresponding to a time instance can be associated with the same or different spatial layers.
[0123] Figure 7 An example of a video sequence structure with a combination of temporal_id, layer_id, POC value, and AUC value having adaptive resolution change is shown. In this example, the pictures, slices, or tiles in the first AU with AUC = 0 can have temporal_id = 0 and layer_id = 0 or 1, while the pictures, slices, or tiles in the second AU with AUC = 1 can have temporal_id = 1 and layer_id = 0 or 1 respectively. Regardless of the values of temporal_id and layer_id, the POC value of each picture increases by 1. In this example, the value of poc_cycle_au can be equal to 2. In an embodiment, the value of poc_cycle_au can be set to be equal to the number of (spatially scalable) layers. Therefore, in this example, the POC value increases by 2 while the AUC value increases by 1.
[0124] In the above embodiments, all or a subset of the inter-picture or inter-layer prediction structures and reference picture indications can be supported by using existing reference picture set (RPS) signaling or reference picture list (RPL) signaling in HEVC. In the RPS or RPL, the selected reference picture can be indicated by signaling the POC value or the incremental value of the POC between the current picture and the selected reference picture. In an embodiment, the RPS and RPL can be used to indicate inter-picture or inter-layer prediction structures without changing the signaling, but with the following limitations. If the value of the temporal_id of the reference picture is greater than the value of the temporal_i of the current picture, the current picture may not use the reference picture for motion compensation or other prediction. If the value of the layer_id of the reference picture is greater than the value of the layer_id of the current picture, the current picture may not use the reference picture for motion compensation or other prediction.
[0125] In an embodiment, the use of POC difference-based motion vector scaling for temporal motion vector prediction can be prohibited between multiple pictures within an access unit. Thus, although each picture within the access unit may have a different POC value, the motion vectors are not scaled and used for temporal motion vector prediction within the access unit. This is because reference pictures with different POCs within the same AU are considered to have the same temporal instance. Thus, in this embodiment, when the reference picture belongs to the AU associated with the current picture, the motion vector scaling function can return 1.
[0126] In an embodiment, when the spatial resolution of the reference picture is different from the spatial resolution of the current picture, optionally, the use of POC difference-based motion vector scaling for temporal motion vector prediction can be prohibited between multiple pictures. When motion vector scaling is allowed, the motion vectors are scaled based on the POC difference and the ratio of the spatial resolutions between the current picture and the reference picture.
[0127] In an embodiment, for temporal motion vector prediction, especially when poc_cycle_au has non-uniform values (e.g., when vps_contant_poc_cycle_per_au == 0), the motion vectors can be scaled based on the AUC difference instead of the POC difference. Otherwise (e.g., when vps_contant_poc_cycle_per_au == 1), the motion vector scaling based on the AUC difference may be the same as the motion vector scaling based on the POC difference.
[0128] In an embodiment, when the motion vectors are scaled based on the AUC difference, the reference motion vectors within the same AU (with the same AUC value) as the current picture are not scaled based on the AUC difference, but are used for motion vector prediction without scaling or are scaled based on the ratio of the spatial resolutions between the current picture and the reference picture and then used for motion vector prediction.
[0129] In an embodiment, the AUC value can be used to identify the boundaries of AUs and for Hypothetical Reference Decoder (HRD) operations, which require timing with AU granularity for both input and output. In an embodiment, the decoded picture with the highest layer in the AU can be output for display. The AUC value and the layer_id value can be used to identify the output picture.
[0130] In an embodiment, a picture can include one or more sub - pictures. Each sub - picture can cover a partial or the entire area of the picture. The area supported by one sub - picture can overlap or not overlap with the area supported by another sub - picture. The area covered by one or more sub - pictures can cover or not cover the entire area of the picture. If the picture includes sub - pictures, the area supported by the sub - picture can be the same as the area supported by the picture.
[0131] In an embodiment, sub - pictures can be encoded by a similar encoding method as that used for the encoded pictures. Sub - pictures can be encoded independently or can be encoded based on another sub - picture or an encoded picture. Sub - pictures can have or not have any parsing dependencies on another sub - picture or an encoded picture.
[0132] In an embodiment, the encoded sub - pictures can be included in one or more layers. The encoded sub - pictures in a layer can have different spatial resolutions. The original sub - pictures can be spatially resampled (e.g., upsampled or downsampled), encoded with different spatial resolution parameters, and included in the bit - stream corresponding to the layer.
[0133] In an embodiment, a sub - picture with (W, H) can be encoded and included in the encoded bit - stream corresponding to layer 0, where W indicates the width of the sub - picture and H indicates the height of the sub - picture. A sub - picture resampled (or downsampled) from a sub - picture with the original spatial resolution (with (W*S w,k ,H*S h,k )) can be encoded and included in the encoded bit - stream corresponding to layer k, where S w,k 、S h,k indicate the resampling rates in the horizontal and vertical directions. If the values of S w,k 、S h,k are greater than 1, the resampling can be upsampling. However, if the values of S w,k 、S h,k are less than 1, the resampling can be downsampling.
[0134] In an embodiment, the visual quality of the encoded sub - pictures in one layer may be different from that of the encoded sub - pictures in another layer for the same sub - picture or a different sub - picture. For example, sub - picture i in layer n can be encoded with a quantization parameter Q i,nis encoded, and the sub - picture j in layer m can be encoded with quantization parameter Q j,m is encoded.
[0135] In an embodiment, an encoded sub - picture in one layer can be independently decodable, having no parsing or decoding dependency on an encoded sub - picture in another layer of the same local region. A sub - picture layer that can be independently decoded without referring to another sub - picture layer of the same local region can be an independent sub - picture layer. An encoded sub - picture in an independent sub - picture layer may or may not have a decoding or parsing dependency on a previously encoded sub - picture in the same sub - picture layer, but the encoded sub - picture has no dependency on an encoded picture in another sub - picture layer.
[0136] In an embodiment, an encoded sub - picture in one layer can be dependently decodable, having any parsing or decoding dependency on an encoded sub - picture in another layer of the same local region. A sub - picture layer that can be dependently decoded by referring to another sub - picture layer of the same local region can be a dependent sub - picture layer. An encoded sub - picture in a dependent sub - picture can refer to an encoded sub - picture belonging to the same sub - picture, a previously encoded sub - picture in the same sub - picture layer, or both reference sub - pictures.
[0137] In an embodiment, an encoded sub - picture can include one or more independent sub - picture layers and one or more dependent sub - picture layers. However, for an encoded sub - picture, there can be at least one independent sub - picture layer. The value of the layer identifier (layer_id) of an independent sub - picture layer can be equal to 0, and this layer identifier can be present in the NAL unit header or another high - level syntax structure. A sub - picture layer with layer_id equal to 0 can be a base sub - picture layer.
[0138] In an embodiment, a picture can include one or more foreground sub - pictures and a background sub - picture. The region supported by the background sub - picture can be equal to the region of the picture. The region supported by the foreground sub - pictures can overlap with the region supported by the background sub - picture. The background sub - picture can be a base sub - picture layer, and the foreground sub - pictures can be non - base (enhanced) sub - picture layers. One or more non - base sub - picture layers can be decoded by referring to the same base layer. Each non - base sub - picture layer with layer_id equal to a can refer to a non - base sub - picture layer with layer_id equal to b, where a is greater than b.
[0139] In an embodiment, a picture may include one or more foreground sub - pictures with or without a background sub - picture. Each sub - picture may have its own base sub - picture layer and one or more non - base (enhanced) layers. Each base sub - picture layer may be referenced by one or more non - base sub - picture layers. Each non - base sub - picture layer with layer_id equal to a may reference a non - base sub - picture layer with layer_id equal to b, where a is greater than b.
[0140] In an embodiment, a picture may include one or more foreground sub - pictures with or without a background sub - picture. Each encoded sub - picture in a (base or non - base) sub - picture layer may be referenced by one or more non - base layer sub - pictures belonging to the same sub - picture and one or more non - base layer sub - pictures not belonging to the same sub - picture.
[0141] In an embodiment, a picture may include one or more foreground sub - pictures with or without a background sub - picture. The sub - pictures in layer a may be further divided into multiple sub - pictures in the same layer. One or more encoded sub - pictures in layer b may reference the divided sub - pictures in layer a.
[0142] In an embodiment, an encoded video sequence (CVS) may be a set of encoded pictures. The CVS may include one or more encoded sub - picture sequences (CSPS), where a CSPS may be a set of encoded sub - pictures covering the same local region of a picture. The CSPS may have the same or different temporal resolution as the encoded video sequence.
[0143] In an embodiment, a CSPS may be encoded and included in one or more layers. A CSPS may include one or more CSPS layers. Decoding one or more CSPS layers corresponding to a CSPS may reconstruct a sub - picture sequence corresponding to the same local region.
[0144] In an embodiment, the number of CSPS layers corresponding to a CSPS may be the same as or different from the number of CSPS layers corresponding to another CSPS.
[0145] In an embodiment, a CSPS layer may have a different temporal resolution (e.g., frame rate) from another CSPS layer. The original (uncompressed) sub - picture sequence may be resampled in time (e.g., upsampled or downsampled), encoded with different temporal resolution parameters, and included in the bitstream corresponding to the layer.
[0146] In an embodiment, a sub - picture sequence with a frame rate F may be encoded and included in the encoded bitstream corresponding to layer 0, while the temporally upsampled (or downsampled) sub - picture sequence with F*S in the original sub - picture sequence t,k may be encoded and included in the encoded bitstream corresponding to layer k, where St,k Indicates the temporal sampling rate of layer k. If S t,k has a value greater than 1, the temporal resampling process can be frame rate up conversion. However, if S t,k has a value less than 1, the temporal resampling process can be frame rate down conversion.
[0147] In an embodiment, when a sub-picture with CSPS layer a is referenced by a sub-picture with CSPS layer b for motion compensation or any inter-layer prediction, if the spatial resolution of CSPS layer a is different from that of CSPS layer b, the decoded pixels in CSPS layer a are resampled and used as a reference. The resampling process can use upsampling filtering or downsampling filtering.
[0148] Figure 10 An example video stream is shown, which includes a background video CSPS with layer_id equal to 0 and multiple foreground CSPS layers. Although an encoded sub-picture can include at least one CSPS layer, the background region that does not belong to any foreground CSPS layer can include a base layer. The base layer can contain both the background region and the foreground region, while the enhanced CSPS layer can contain the foreground region. In the same region, the enhanced CSPS layer may have better visual quality than the base layer. The enhanced CSPS layer can reference the reconstructed pixels in the corresponding same region and the motion vectors of the base layer.
[0149] In an embodiment, in a video file, the video bitstream corresponding to the base layer is contained in a track, while the CSPS layer corresponding to each sub-picture is contained in a separate track.
[0150] In an embodiment, the video bitstream corresponding to the base layer is contained in a track, while the CSPS layers with the same layer_id are contained in separate tracks. In this example, the track corresponding to layer k only includes the CSPS layer corresponding to layer k.
[0151] In an embodiment, each CSPS layer of each sub-picture is stored in a separate track. Each track may or may not have any parsing or decoding dependencies on one or more other tracks.
[0152] In an embodiment, each track can contain the bitstream corresponding to layers i to j of the CSPS layer of all or a subset of the sub-pictures, where 0 < i =< j =< k and k is the highest layer of the CSPS.
[0153] In an embodiment, the picture includes one or more associated media data, and these associated media data include depth maps, alpha (α) maps, 3D geometric data, occupancy maps, etc. Such associated timed media data can be divided into one or more data sub-streams, and each data sub-stream corresponds to a sub-picture.
[0154] Figure 11 An example of a video conference based on a multi-layer sub-picture method is shown. In a video stream, it includes a base layer video bitstream corresponding to a background picture and one or more enhancement layer video bitstreams corresponding to foreground sub-pictures. Each enhancement layer video bitstream can correspond to a CSPS layer. On a display, the picture corresponding to the base layer is displayed by default. It includes a picture-in-picture (PIP) of one or more users. When a specific user is selected through the controls of the client, the enhanced CSPS layer corresponding to the selected user can be decoded and displayed with enhanced quality or spatial resolution.
[0155] Figure 12 A block diagram showing an example of the above process is shown. For example, in step S1210, a video bitstream with multiple layers can be decoded. In step S1220, the background region and one or more foreground sub-pictures can be identified. In step 1230, it can be determined whether a specific sub-picture region is selected, such as one of the foreground sub-pictures. If a specific sub-picture region is selected (Yes at step S1240), the enhanced sub-picture can be decoded and displayed. If a specific sub-picture region is not selected (No at step S1240), the background region can be decoded and displayed.
[0156] In an embodiment, a network intermediate box (such as a router) can select a subset of the layers to be sent to a user according to its bandwidth. Picture / sub-picture organization can be used for bandwidth adaptation. For example, if a user has no bandwidth, the router will strip layers or select some sub-pictures due to their importance or based on the settings used, and can do this dynamically to adapt to the bandwidth.
[0157] Figure 13Embodiments related to the usage of 360-degree videos are shown. When a spherical 360-degree picture (e.g., picture 1310) is projected onto a planar picture, the projected 360-degree picture can be divided into multiple sub-pictures as a base layer. For example, the multiple sub-pictures can include a rear sub-picture, an upper sub-picture, a right sub-picture, a left sub-picture, a front sub-picture, and a lower sub-picture. The enhancement layer of a specific sub-picture (e.g., the front sub-picture) can be encoded and sent to the client. The decoder is capable of decoding the base layer including all sub-pictures and the enhancement layer of the selected sub-picture. When the current viewport is the same as the selected sub-picture, the displayed picture may have a higher quality compared to the decoded sub-picture with the enhancement layer. Otherwise, the decoded picture with the base layer can be displayed with a lower quality.
[0158] In an embodiment, any layout information for display can exist in a file as supplementary information (e.g., SEI message or metadata). One or more decoded sub-pictures can be repositioned and displayed according to the signaled layout information. The layout information can be signaled by a streaming server or a broadcast device, or can be regenerated by a network entity or a cloud server, or can be determined by the user's customized settings.
[0159] In an embodiment, when an input picture is divided into one or more (rectangular) sub-regions, each sub-region can be encoded as an independent layer. Each independent layer corresponding to a local region can have a unique layer_id value. For each independent layer, sub-picture size and position information can be signaled, e.g., picture size (width, height), offset information (x_offset, y_offset) at the upper left corner. Figure 14 An example of the layout of the divided sub-pictures, their sub-picture size and position information, and their corresponding picture prediction structure is shown. The layout information including one or more sub-picture sizes and one or more sub-picture positions can be signaled in a high-level syntax structure (e.g., one or more parameter sets, slice headers, or tile group headers, or SEI messages).
[0160] In an embodiment, each sub-picture corresponding to an independent layer can have a unique POC value within an AU. When indicating reference pictures in the pictures stored in the DPB by using syntax elements in the RPS or RPL structure, the POC value of each sub-picture corresponding to the layer can be used.
[0161] In an embodiment, in order to indicate the (inter-layer) prediction structure, the layer_id can not be used and the POC (delta) value can be used.
[0162] In an embodiment, a sub-picture corresponding to a layer (or local region) with a POC value equal to N may or may not be used as a reference picture for a sub-picture corresponding to the same layer (or the same local region) with a POC value equal to N+K for motion compensation prediction. In most cases, the value of the quantity K may be equal to the maximum number of (independent) layers, which may be the same as the number of sub-regions.
[0163] In an embodiment, Figure 15 shows Figure 14 an extended case. When an input picture is divided into multiple (e.g., four) sub-regions, each local region may be encoded with one or more layers. In this case, the number of independent layers may be equal to the number of sub-regions, and one or more layers may correspond to the sub-regions. Thus, each sub-region may be encoded with one or more independent layers and zero or more dependent layers.
[0164] In an embodiment, in Figure 15 the input picture may be divided into four sub-regions. As an example, the upper-right sub-region may be encoded with two layers, namely layer 1 and layer 4, while the lower-right sub-region may be encoded with two layers, namely layer 3 and layer 5. In this case, layer 4 may perform motion compensation prediction with reference to layer 1, and layer 5 may perform motion compensation with reference to layer 3.
[0165] In an embodiment, in-loop filtering (e.g., deblocking filtering, adaptive in-loop filtering, shaper, bilateral filtering, or any deep learning-based filtering) across layer boundaries may be (optionally) disabled.
[0166] In an embodiment, motion compensation prediction or intra-block copy across layer boundaries may be (optionally) disabled.
[0167] In an embodiment, boundary filling for motion compensation prediction or in-loop filtering at sub-picture boundaries may be optionally processed. A flag may be signaled in a high-level syntax structure (e.g., one or more parameter sets (VPS, SPS, PPS, or APS), slice header or tile group header, or SEI message) to indicate whether the boundary filling is processed.
[0168] In an embodiment, layout information of one or more sub-regions (or one or more sub-pictures) may be signaled in the VPS or SPS. Figure 16A Shows an example of syntax elements in the VPS, Figure 16BAn example of a syntax element in SPS is shown. In this example, vps_sub_picture_dividing_flag is signaled in the VPS. This flag can indicate whether one or more input pictures are divided into multiple sub-regions. When the value of vps_sub_picture_dividing_flag is equal to 0, one or more input pictures in one or more encoded video sequences corresponding to the current VPS may not be divided into multiple sub-regions. In this case, the input picture size may be equal to the encoded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples), which is signaled in the SPS. When the value of vps_sub_picture_dividing_flag is equal to 1, the one or more input pictures may be divided into multiple sub-regions. In this case, the syntax elements vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples are signaled in the VPS. The values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may be equal to the width and height of the one or more input pictures, respectively.
[0169] In an embodiment, the values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may not be used for decoding, but for synthesis and display.
[0170] In an embodiment, when the value of vps_sub_picture_dividing_flag is equal to 1, the syntax elements pic_offset_x and pic_offset_y corresponding to one or more specific layers may be signaled in the SPS. In this case, the encoded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples) signaled in the SPS may be equal to the width and height of the sub-region corresponding to the specific layer. Moreover, the position (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region may be signaled in the SPS.
[0171] In an embodiment, the position information (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region may not be used for decoding, but for synthesis and display.
[0172] In an embodiment, the layout information (size and position) of all or a subset of sub-regions of one or more input pictures, and the dependency information between layers may be signaled in a parameter set or SEI message. Figure 17 Examples of syntax elements are shown for information indicating the layout of sub-regions, the dependency between layers, and the relationship between sub-regions and one or more layers. In this example, the syntax element num_sub_region indicates the number of (rectangular) sub-regions in the current encoded video sequence. The syntax element num_layers indicates the number of layers in the current encoded video sequence. The value of num_layers may be equal to or greater than the value of num_sub_region. When any sub-region is encoded in a single layer, the value of num_layers may be equal to the value of num_sub_region. When one or more sub-regions are encoded in multiple layers, the value of num_layers may be greater than the value of num_sub_region. The syntax element direct_dependency_flag[i][j] indicates the dependency from layer j to layer i. num_layers_for_region[i] indicates the number of layers associated with the i-th sub-region. sub_region_layer_id[i][j] indicates the layer_id of the j-th layer associated with the i-th sub-region. sub_region_offset_x[i] and sub_region_offset_y[i] indicate the horizontal and vertical positions of the upper left corner of the i-th sub-region, respectively. sub_region_width[i] and sub_region_height[i] indicate the width and height of the i-th sub-region, respectively.
[0173] In an embodiment, one or more syntax elements may be signaled in a high-level syntax structure (such as VPS, DPS, SPS, PPS, APS, or SEI message), the one or more syntax elements specifying an output layer set to indicate one of the multiple layers with or without profile-level information to be output. Refer to Figure 18 , the syntax element num_output_layer_sets may be signaled in the VPS, the syntax element num_output_layer_sets indicating the number of output layer sets (OLSs) in the encoded video sequence referring to the VPS. For each output layer set, as many output_layer_flag as the number of output layers may be signaled.
[0174] In an embodiment, output_layer_flag[i] being equal to 1 specifies the output of the i-th layer. vps_output_layer_flag[i] being equal to 0 specifies no output of the i-th layer.
[0175] In an embodiment, one or more syntax elements that specify the profile tier level information for each output layer set may be signaled in a high-level syntax structure (such as a VPS, DPS, SPS, PPS, APS, or SEI message). Still referring to Figure 18 , the syntax element num_profile_tile_level may be signaled in the VPS, and the syntax element num_profile_tile_level indicates the number of profile tier level information for each OLS in the encoded video sequence with reference to the VPS. For each output layer set, a set of syntax elements for the profile tier level information as many as the number of output layers or an index indicating the specific profile tier level information of the entries in the profile tier level information may be signaled.
[0176] In an embodiment, profile_tier_level_idx[i][j] specifies the index of the profile_tier_level() syntax structure applied to the j-th layer of the i-th OLS in the list of profile_tier_level() syntax structures in the VPS.
[0177] In an embodiment, referring to Figure 19 , when the maximum number of layers is greater than 1 (vps_max_layers_minus1>0), the syntax elements num_profile_tile_level and / or num_output_layer_sets may be signaled.
[0178] In an embodiment, referring to Figure 19 , the syntax element vps_output_layers_mode[i] that indicates the mode of the output layer signaling for the i-th output layer set may be present in the VPS.
[0179] In an embodiment, vps_output_layers_mode[i] being equal to 0 specifies using the i-th output layer set to output only the highest layer. vps_output_layer_mode[i] being equal to 1 specifies using the i-th output layer set to output all layers. vps_output_layer_mode[i] being equal to 2 specifies that the layers to be output are the layers where vps_output_layer_flag[i][j] is equal to 1 and the i-th output layer set is used. More values may be reserved.
[0180] In an embodiment, according to the value of vps_output_layers_mode[i] for the i-th output layer set, output_layer_flag[i][j] may or may not be signaled.
[0181] In an embodiment, referring to Figure 19 , for the i-th output layer set, there may be a flag vps_ptl_signal_flag[i]. According to the value of vps_ptl_signal_flag[i], the profile level information of the i-th output layer set may or may not be signaled.
[0182] In an embodiment, referring to Figure 20 , the number of sub-pictures max_subpics_minus1 in the current CVS may be signaled in a high-level syntax structure (such as VPS, DPS, SPS, PPS, APS, or SEI message).
[0183] In an embodiment, referring to Figure 20 , when the number of sub-pictures is greater than 1 (max_subpics_minus1 > 0), the sub-picture identifier sub_pic_id[i] of the i-th sub-picture may be signaled.
[0184] In an embodiment, one or more syntax elements indicating the sub-picture identifiers of each layer belonging to each output layer set may be signaled in the VPS. Referring to Figure 20 , sub_pic_id_layer[i][j][k] indicates the k-th sub-picture present in the j-th layer of the i-th output layer set. Using this information, the decoder can identify which sub-pictures are decoded and output for each layer of a specific output layer set.
[0185] In an embodiment, the picture header (PH) is a syntax structure that contains the syntax elements applied to all slices of an encoded picture. A picture unit (PU) is a group of NAL units that are related to each other according to specified classification rules, are consecutive in decoding order, and contain only one encoded picture. A PU may contain a picture header (PH) and one or more VCL NAL units that make up the encoded picture.
[0186] In an embodiment, the SPS (raw byte sequence payload (RBSP)) may be used for the decoding process before being referred to, and it is included in at least one AU with TemporalId equal to 0 or provided by an external device.
[0187] In an embodiment, the SPS (RBSP) can be used for the decoding process before being referenced, which includes being provided in at least one AU in the CVS with a TemporalId equal to 0 or by an external device, and the at least one AU contains one or more PPSs that reference the SPS.
[0188] In an embodiment, the SPS (RBSP) can be used for the decoding process before being referenced by one or more PPSs, which includes being provided in at least one PU in the CVS or by an external device, the nuh_layer_id of the at least one PU being equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit, and the at least one PU contains one or more PPSs that reference the SPS.
[0189] In an embodiment, the SPS (RBSP) can be used for the decoding process before being referenced by one or more PPSs, which includes being provided in at least one PU or by an external device, the TemporalId of the at least one PU being equal to 0 and the nuh_layer_id being equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit.
[0190] In an embodiment, the SPS (RBSP) can be used for the decoding process before being referenced by one or more PPSs, which includes being provided in at least one PU in the CVS or by an external device, the TemporalId of the at least one PU being equal to 0, the nuh_layer_id being equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit, and the at least one PU contains one or more PPSs that reference the SPS.
[0191] In the same or another embodiment, the pps_seq_parameter_set_id can specify the value of the sps_seq_parameter_set_id for the referenced SPS. The value of the pps_seq_parameter_set_id can be the same for all PPSs referenced by the coded pictures in the coded layer video sequence (CLVS).
[0192] In the same or another embodiment, all SPS NAL units in the CVS with a specific value of the sps_seq_parameter_set_id can have the same content.
[0193] In the same or another embodiment, regardless of the nuh_layer_id value, the SPS NAL units can share the same value space of the sps_seq_parameter_set_id.
[0194] In the same or another embodiment, the nuh_layer_id value of the SPS NAL unit may be equal to the lowest nuh_layer_id value of the PPS NAL unit that references the reference SPS NAL unit.
[0195] In an embodiment, when an SPS with nuh_layer_id equal to m is referenced by one or more PPSs with nuh_layer_id equal to n, the layer with nuh_layer_id equal to m may be the same as the layer with nuh_layer_id equal to n or the (direct or indirect) reference layer of the layer with nuh_layer_id equal to m.
[0196] In an embodiment, the PPS (RBSP) may be used for the decoding process before being referenced, which includes being in at least one AU with a TemporalId equal to the TemporalId of the PPS NAL unit or being provided by an external device.
[0197] In an embodiment, the PPS (RBSP) may be used for the decoding process before being referenced, which includes being in at least one AU in the CVS with a TemporalId equal to the TemporalId of the PPS NAL unit or being provided by an external device, and the at least one AU contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0198] In an embodiment, the PPS (RBSP) may be used for the decoding process before being referenced by one or more PHs (or coded slice NAL units), which includes being in at least one PU in the CVS or being provided by an external device, the nuh_layer_id of the at least one PU is equal to the lowest nuh_layer_id value of the coded slice NAL unit that references the PPS NAL unit, and the at least one PU contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0199] In an embodiment, the PPS (RBSP) may be used for the decoding process before being referenced by one or more PHs (or coded slice NAL units), which includes being in at least one PU in the CVS or being provided by an external device, the TemporalId of the at least one PU is equal to the TemporalId of the PPS NAL unit, the nuh_layer_id is equal to the lowest nuh_layer_id value of the coded slice NAL unit that references the PPS NAL unit, and the at least one PU contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0200] In the same or another embodiment, the ph_pic_parameter_set_id in PH may be the value of pps_pic_parameter_set_id specified for the used referenced PPS. The value of pps_seq_parameter_set_id may be the same in all PPSs referenced by the coded pictures in CLVS.
[0201] In the same or another embodiment, all PPS NAL units having a specific value of pps_pic_parameter_set_id within a PU may have the same content.
[0202] In the same or another embodiment, regardless of the value of nuh_layer_id, PPS NAL units may share the same value space of pps_pic_parameter_set_id.
[0203] In the same or another embodiment, the nuh_layer_id value of a PPS NAL unit may be equal to the lowest nuh_layer_id value of the coded slice NAL unit of the referenced NAL unit that references the PPS NAL unit.
[0204] In an embodiment, when a PPS with nuh_layer_id equal to m is referenced by one or more coded slice NAL units with nuh_layer_id equal to n, the layer with nuh_layer_id equal to m may be the same as the layer with nuh_layer_id equal to n or the (direct or indirect) reference layer of the layer with nuh_layer_id equal to m.
[0205] In an embodiment, PPS (RBSP) may be used for the decoding process before being referenced, which includes being provided in at least one AU with TemporalId equal to the TemporalId of the PPS NAL unit or by an external device.
[0206] In an embodiment, PPS (RBSP) may be used for the decoding process before being referenced, which includes being provided in at least one AU in CVS with TemporalId equal to the TemporalId of the PPS NAL unit or by an external device, and the at least one AU contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0207] In an embodiment, the PPS (RBSP) can be used for the decoding process before being referenced by one or more PHs (or coded slice NAL units), which includes being provided in at least one PU in the CVS or by an external device. The nuh_layer_id of the at least one PU is equal to the lowest nuh_layer_id value of the coded slice NAL units that reference the PPS NAL unit, and the at least one PU contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0208] In an embodiment, the PPS (RBSP) can be used for the decoding process before being referenced by one or more PHs (or coded slice NAL units), which includes being provided in at least one PU in the CVS or by an external device. The TemporalId of the at least one PU is equal to the TemporalId of the PPS NAL unit, the nuh_layer_id is equal to the lowest nuh_layer_id value of the coded slice NAL units that reference the PPS NAL unit, and the at least one PU contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0209] In the same or another embodiment, the ph_pic_parameter_set_id in the PH can be the value of pps_pic_parameter_set_id specified for the referenced PPS. The value of pps_seq_parameter_set_id can be the same for all PPSs referenced by the coded pictures in the CLVS.
[0210] In the same or another embodiment, all PPS NAL units having a specific value of pps_pic_parameter_set_id within a PU can have the same content.
[0211] In the same or another embodiment, regardless of the nuh_layer_id value, the PPS NAL units can share the same value space of pps_pic_parameter_set_id.
[0212] In the same or another embodiment, the nuh_layer_id value of the PPS NAL unit can be equal to the lowest nuh_layer_id value of the coded slice NAL units of the reference NAL unit that references the PPS NAL unit.
[0213] In an embodiment, when a PPS with nuh_layer_id equal to m is referenced by one or more coded slice NAL units with nuh_layer_id equal to n, the layer with nuh_layer_id equal to m may be the same as the layer with nuh_layer_id equal to n or the (direct or indirect) reference layer of the layer with nuh_layer_id equal to m.
[0214] In an embodiment, as Figure 21 shown, pps_subpic_id[i] in the picture parameter set may specify the sub-picture ID of the i-th sub-picture. The syntax element pps_subpic_id[i] has a length of pps_subpic_id_len_minus1 + 1 bits.
[0215] For each i value in the range from 0 to sps_num_subpics_minus1 (including 0 and sps_num_subpics_minus1), the variable SubpicIdVal[i] may be derived as follows:
[0216] for(i = 0; i <= sps_num_subpics_minus1; i++)
[0217] if(subpic_id_mapping_explicitly_signalled_flag)
[0218] SubpicIdVal[i] = subpic_id_mapping_in_pps_flag? pps_subpic_id[i] :
[0219] sps_subpic_id[i]
[0220] else
[0221] SubpicIdVal[i] = i
[0222] In the same or another embodiment, for any two different values of i and j in the range from 0 to sps_num_subpics_minus1 (including 0 and sps_num_subpics_minus1), SubpicIdVal[i] may not be equal to SubpicIdVal[j].
[0223] In the same or another embodiment, when the current picture is not the first picture of CLVS, for each i value in the range from 0 to sps_num_subpics_minus1 (including 0 and sps_num_subpics_minus1), if the value of SubpicIdVal[i] is not equal to the value of SubpicIdVal[i] of the previous picture in decoding order in the same layer, then the nal_unit_type of all the encoded slice NAL units of the subpicture in the current picture with subpicture index i may be equal to a specific value in the range from IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT).
[0224] In the same or another embodiment, when the current picture is not the first picture of CLVS, for each i value in the range from 0 to sps_num_subpics_minus1 (including 0 and sps_num_subpics_minus1), if the value of SubpicIdVal[i] is not equal to the value of SubpicIdVal[i] of the previous picture in decoding order in the same layer, then sps_independent_subpics_flag may be equal to 1.
[0225] In the same or another embodiment, when the current picture is not the first picture of CLVS, for each i value in the range from 0 to sps_num_subpics_minus1 (including 0 and sps_num_subpics_minus1), if the value of SubpicIdVal[i] is not equal to the value of SubpicIdVal[i] of the previous picture in decoding order in the same layer, then subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i] may be equal to 1.
[0226] In the same or another embodiment, when the current picture is not the first picture of CLVS, for each i value in the range from 0 to sps_num_subpics_minus1 (including 0 and sps_num_subpics_minus1), if the value of SubpicIdVal[i] is not equal to the value of SubpicIdVal[i] of the previous picture in decoding order in the same layer, then sps_independent_subpics_flag may be equal to 1, or subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i] may be equal to 1.
[0227] In the same or another embodiment, when a sub-picture is independently encoded without reference to another sub-picture, the value of the sub-picture identifier of a region can be changed within the encoded video sequence.
[0228] Figure 22 is a schematic diagram of an example process 2200 for decoding an encoded video bitstream. In some embodiments, Figure 22 one or more steps of can be performed by decoder 210. In some implementations, Figure 22 one or more steps of can be performed by another device or a set of devices separate from or including decoder 210, such as encoder 203.
[0229] As Figure 22 shown, process 2200 may include obtaining a first flag that indicates the sub-picture identifier of the sub-pictures of the current picture being signaled explicitly (step 2210). In an embodiment, the first flag may correspond to subpic_id_explicitly_signalled_flag.
[0230] As Figure 22 further shown, process 2200 may include obtaining, based on the first flag, a second flag that indicates whether the sub-picture identifier is signaled in the Picture Parameter Set (PPS) or in the Sequence Parameter Set (SPS) (step 2220). In an embodiment, the second flag may correspond to subpic_id_mapping_in_pps_flag.
[0231] As Figure 22 further shown, process 2200 may include determining, based on the second flag, whether the sub-picture identifier is signaled in the PPS (step 2230).
[0232] As Figure 22 further shown, when the second flag indicates that the sub-picture identifier is signaled in the PPS (being "yes" at step 2230), process 2200 may include determining the sub-picture identifier based on a first syntax element included in the PPS (step 2240). In an embodiment, the first syntax element may correspond to pps_subpic_id.
[0233] As Figure 22Further shown, when the second flag indicates that the sub-picture identifier is signaled in the SPS (being "no" at step 2230), process 2200 may include determining the sub-picture identifier based on a second syntax element included in the SPS (step 2250). In an embodiment, the second syntax element may correspond to sps_subpic_id.
[0234] As Figure 22 Further shown, process 2200 may include decoding the current picture based on the determined sub-picture identifier (step 2260).
[0235] In an embodiment, the second flag is signaled in the PPS.
[0236] In an embodiment, the length of the first syntax element is signaled in the PPS.
[0237] In an embodiment, based on the sub-picture identifier being different from the previous sub-picture identifier of the corresponding previous sub-picture of a previous picture, the second flag indicates that the sub-picture identifier is signaled in the PPS.
[0238] In an embodiment, based on the sub-picture identifier being different from the previous sub-picture identifier of the corresponding previous sub-picture of a previous picture, a third flag indicates that the sub-picture of the current picture is independent. In an embodiment, the third flag may correspond to sps_independent_subpics_flag.
[0239] In an embodiment, based on the sub-picture identifier being different from the previous sub-picture identifier of the corresponding previous sub-picture of a previous picture, a fourth flag indicates that the sub-picture of the current picture is treated as a picture. In an embodiment, the fourth flag may correspond to subpic_treated_as_pic_flag.
[0240] In an embodiment, based on the sub-picture identifier being different from the previous sub-picture identifier of the corresponding previous sub-picture of a previous picture, a fifth flag indicates that loop filtering is disabled at the sub-picture boundary in the current picture. In an embodiment, the fifth flag may correspond to loop_filter_across_subpic_enabled_flag.
[0241] While Figure 22 example steps of process 2200 are shown, in some implementations, the process 2200 may include more steps, fewer steps, different steps, or steps in a different arrangement than Figure 22 shown. Additionally or alternatively, two or more steps of the process 2200 may be executed in parallel.
[0242] In addition, the proposed method can be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-volatile computer-readable medium to perform one or more of the proposed methods.
[0243] The above techniques can be implemented as computer software by computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 23 FIG. 2300 shows a computer system, which is suitable for implementing certain embodiments of the present application.
[0244] The computer software can be encoded by any suitable machine code or computer language, and code including instructions is created through mechanisms such as assembly, compilation, and linking. The instructions can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through methods such as decoding and microcode.
[0245] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0246] Figure 23 The components shown for the computer system 2300 are exemplary in nature and are not used to impose any limitations on the scope of use or functions of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependence on or requirement for any one component or combination thereof shown in the exemplary embodiments of the computer system 2300.
[0247] The computer system 2300 may include certain human-machine interface input devices. Such human-machine interface input devices can respond to inputs from one or more human users through tactile inputs (such as keyboard inputs, swipes, data glove movements), audio inputs (such as sounds, applause), visual inputs (such as gestures), and olfactory inputs (not shown). The human-machine interface device can also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional videos, three-dimensional videos including stereoscopic videos).
[0248] The human-machine interface input device may include one or more of the following (only one of which is drawn): keyboard 2301, mouse 2302, touchpad 2303, touch screen 2310 and associated graphics adapter 2350, data glove, joystick 2305, microphone 2306, scanner 2307, camera 2308.
[0249] The computer system 2300 may also include certain human - machine interface output devices. Such human - machine interface output devices can stimulate the senses of one or more human users through, for example, haptic output, sound, light, and smell / taste. Such human - machine interface output devices may include haptic output devices (e.g., haptic feedback through the touch screen 2310, data gloves, or joystick 2305, but there may also be haptic feedback devices that do not serve as input devices), audio output devices (e.g., speakers 2309, headphones (not shown)), visual output devices (e.g., the screen 2310 including a cathode ray tube (CRT) screen, a liquid crystal (LCD) screen, a plasma screen, an organic light - emitting diode (OLED) screen, each of which may or may not have touch - screen input functionality, each of which may or may not have haptic feedback functionality - some of which may output two - dimensional visual output or output above three - dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke - emitting boxes (not shown)), and printers (not shown).
[0250] The computer system 2300 may also include human - accessible storage devices and their associated media, such as optical media including high - density read - only / rewritable compact discs (CD / DVD ROM / RW) 2320 or similar media 2321, thumb drives 2322, removable hard disk drives or solid - state drives 2323, traditional magnetic media such as tapes and floppy disks (not shown), ROM / ASIC / PLD - based dedicated devices such as security software protectors (not shown), and so on.
[0251] Those skilled in the art should also understand that the term "computer - readable medium" as used in conjunction with this application does not include transmission media, carrier waves, or other transient signals.
[0252] The computer system 2300 may also include an interface to one or more communication networks (2355). For example, the network may be wireless, wired, optical. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicle network and an industrial network, a real-time network, a delay-tolerant network, and so on. The network also includes Ethernet, wireless local area network, cellular network (Global System for Mobile Communications (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), Long Term Evolution (LTE), etc.), local area networks such as these, television cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicle and industrial networks (including CANBus), and so on. Some networks typically require an external network interface adapter (2354) for connection to certain general-purpose data ports or peripheral buses (2349) (e.g., the Universal Serial Bus (USB) port of the computer system 2300); other systems are typically integrated into the core of the computer system 2300 by connecting to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). As an example, the network 2355 may be connected to the peripheral bus 2349 using the network interface 2354. By using any of these networks, the computer system 2300 can communicate with other entities. The communication may be unidirectional, only for receiving (e.g., wireless television), unidirectional only for sending (e.g., CAN bus to certain CAN bus devices), or bidirectional, e.g., to other computer systems via a local or wide area digital network. Each of the above-mentioned networks and network interfaces (2354) may use certain protocols and protocol stacks.
[0253] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be connected to the core 2340 of the computer system 2300.
[0254] The core 2340 may include one or more central processing units (CPUs) 2341, a graphics processing unit (GPU) 2342, a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) 2343, a hardware accelerator 2344 for specific tasks, and so on. These devices, as well as read-only memory (ROM) 2345, random access memory (RAM) 2346, internal mass storage (e.g., internal non-user-accessible hard disk drive, solid state drive (SSD), etc.) 2347, and so on, may be connected via the system bus 2348. In some computer systems, the system bus 2348 may be accessed in the form of one or more physical plugs for expansion by additional central processing units, graphics processing units, etc. Peripheral devices may be directly attached to the system bus 2348 of the core or connected via the peripheral bus 2349. The architecture of the peripheral bus includes Peripheral Component Interconnect (PCI), Universal Serial Bus USB, and so on.
[0255] The CPU 2341, GPU 2342, FPGA 2343, and accelerator 2344 can execute certain instructions, which, when combined, can form the aforementioned computer code. This computer code can be stored in the ROM 2345 or RAM 2346. Intermediate data can also be stored in the RAM 2346, while permanent data can be stored, for example, in the internal mass storage 2347. Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more of the CPU 2341, GPU 2342, mass storage 2347, ROM 2345, RAM 2346, etc.
[0256] The computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this application, or can be of the kind well-known and available to those skilled in the computer software art.
[0257] By way of example and not limitation, a computer system having the architecture 2300, and in particular the core 2340, can provide the functionality of a processor (including CPU, GPU, FPGA, accelerator, etc.) to execute software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the above-described user-accessible mass storage, as well as specific memories of the non-volatile core 2340, such as the core internal mass storage 2347 or ROM 2345. The software implementing the various embodiments of this application can be stored in such devices and executed by the core 2340. Depending on specific requirements, the computer-readable medium can include one or more storage devices or chips. The software can cause the core 2340, and in particular the processors therein (including CPU, GPU, FPGA, etc.), to execute the specific processes or specific portions of the specific processes described herein, including defining data structures stored in the RAM 2346 and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality that is logically hardwired or otherwise included in circuitry (e.g., accelerator 2344) that can operate in place of or in conjunction with the software to execute the specific processes or specific portions of the specific processes described herein. In appropriate instances, reference to software can include logic, and vice versa. In appropriate instances, reference to a computer-readable medium can include circuitry (such as an integrated circuit (IC)) that stores the software for execution, circuitry that contains the logic for execution, or both. This application encompasses any suitable combination of hardware and software.
[0258] Although the present application has described multiple exemplary embodiments, various changes, permutations, and various equivalent substitutions of the embodiments fall within the scope of the present application. Therefore, it should be understood that those skilled in the art can design various systems and methods, which, although not explicitly shown or described herein, embody the principles of the present application and thus fall within the spirit and scope of the present application.
Claims
1. A method for decoding an encoded video bitstream, characterized in that, The method includes: Obtaining a first flag, where the first flag indicates a sub-picture identifier that explicitly signals a sub-picture of the current picture; Based on the first flag, obtaining a second flag, where the second flag indicates whether the sub-picture identifier is signaled in a Picture Parameter Set (PPS) or in a Sequence Parameter Set (SPS); When the second flag indicates that the sub-picture identifier is signaled in the Picture Parameter Set (PPS), determining the sub-picture identifier based on a first syntax element included in the Picture Parameter Set (PPS); When the second flag indicates that the sub-picture identifier is signaled in the Sequence Parameter Set (SPS), determining the sub-picture identifier based on a second syntax element included in the Sequence Parameter Set (SPS); Based on the sub-picture identifier being different from the sub-picture identifier of the corresponding previous sub-picture of the previous picture in decoding order in the same layer, a third flag indicates that the sub-picture of the current picture is independent; and Decoding the current picture based on the determined sub-picture identifier.
2. The method according to claim 1, wherein Signaling the second flag in the Picture Parameter Set (PPS).
3. The method according to claim 1, wherein Signaling the length of the first syntax element in the Picture Parameter Set (PPS).
4. The method according to claim 1, wherein Based on the sub-picture identifier being different from the sub-picture identifier of the corresponding previous sub-picture of the previous picture in decoding order in the same layer, a fourth flag indicates that the sub-picture of the current picture is regarded as a picture.
5. The method according to claim 1, characterized in that Based on the sub-picture identifier being different from the sub-picture identifier of the corresponding previous sub-picture of the previous picture in decoding order in the same layer, a fifth flag indicates that loop filtering is disabled at the sub-picture boundary in the current picture.
6. A method for encoding video data, applied to an encoder including a local decoder, characterized in that, The local decoder performs the following method: Obtaining a first flag, where the first flag indicates a sub-picture identifier that explicitly signals a sub-picture of the current picture; Based on the first flag, obtaining a second flag, where the second flag indicates whether the sub-picture identifier is signaled in a Picture Parameter Set (PPS) or in a Sequence Parameter Set (SPS); When the second flag indicates that the sub-picture identifier is signaled in the Picture Parameter Set (PPS), determining the sub-picture identifier based on a first syntax element included in the Picture Parameter Set (PPS); When the second flag indicates that the sub-picture identifier is signaled in the Sequence Parameter Set (SPS), determining the sub-picture identifier based on a second syntax element included in the Sequence Parameter Set (SPS); Based on the sub-picture identifier being different from the sub-picture identifier of the corresponding previous sub-picture of the previous picture in decoding order in the same layer, a third flag indicates that the sub-picture of the current picture is independent; And Decoding the current picture based on the determined sub-picture identifier.
7. A device for decoding an encoded video bitstream, characterized in that, The device includes: At least one memory for storing program code; and At least one processor for reading the program code and operating according to the instructions of the program code to perform the method according to any one of claims 1-5.
8. An apparatus for decoding an encoded video bitstream, characterized in that, The apparatus includes: A first obtaining module, configured to obtain a first flag, where the first flag indicates a sub-picture identifier that explicitly signals a sub-picture of the current picture; A second obtaining module, configured to obtain a second flag based on the first flag, where the second flag indicates whether the sub-picture identifier is signaled in a picture parameter set (PPS) or in a sequence parameter set (SPS); A first determining module, configured to, when the second flag indicates that the sub-picture identifier is signaled in the PPS, determine the sub-picture identifier based on a first syntax element included in the PPS; A second determining module, configured to, when the second flag indicates that the sub-picture identifier is signaled in the SPS, determine the sub-picture identifier based on a second syntax element included in the SPS; A third determining module, configured to, based on that the sub-picture identifier is different from a sub-picture identifier of a corresponding previous sub-picture of a previous picture in decoding order in the same layer, a third flag indicates that the sub-picture of the current picture is independent; and A decoding module, configured to decode the current picture based on the determined sub-picture identifier.
9. A non - volatile computer - readable medium, characterized in that, Instructions for storage, the instructions including one or more instructions that, when executed by one or more processors of a device for decoding an encoded video bitstream, cause the one or more processors to execute the method according to any one of claims 1-5.
10. A method for storing a video bitstream, characterized in that, The video bitstream is decoded according to the method according to any one of claims 1-5, or the video bitstream is generated according to the method according to claim 6.