Method and apparatus for video decoding
By identifying and processing sub-picture areas in multi-layer video bitstreams in video encoding and decoding, the inter-layer alignment problem is solved, and video decoding and display with adaptive resolution is realized, and technical flexibility and efficiency are improved.
Patent Information
- Application Number
- CN202080036505.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-05
- Filing Date
- 2020-10-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-10-19
AI Technical Summary
Existing video encoding and decoding technologies are difficult to effectively deal with the problem of alignment between layers in encoded video data, especially under the requirement of adaptive resolution settings between multiple semantic independent source images.
By decoding a video bitstream with multiple layers, a sub-picture area is identified, and a foreground sub-picture area is selected for decoding and displaying as needed, or a background area is selected for decoding and displaying.
Cross-layer aligned video decoding and display are realized, adapting to the adaptive resolution requirements of different scene activities, and improving the flexibility and efficiency of video encoding and decoding.
Smart Images

Figure CN114127800B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 954,844, filed on December 30, 2019, and U.S. Patent Application No. 17 / 063,025, filed on October 5, 2020, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to the field of video encoding and decoding, and more particularly, to parameter set references and ranges in an encoded video stream. Background Art
[0004] For decades, video encoding and decoding using inter - picture prediction with motion compensation has been well - known. Uncompressed digital video can include a series of pictures, each picture having a spatial dimension such as 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (also informally called the frame rate), such as 60 pictures per second or 60 Hz. Uncompressed video has high bit - rate requirements. For example, a 1080p60 4:2:0 video (1920x1080 luminance sample resolution at 60 Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video would require more than 600 GB of storage space.
[0005] One goal of video encoding and decoding is to reduce the redundancy of the input video signal through compression. Compression can help reduce the requirements for the above - mentioned bandwidth or storage space, and in some cases, can reduce by two or more orders of magnitude. Both lossless compression and lossy compression, as well as combinations of both, can be employed. Lossless compression is a technique for reconstructing an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal can be used for the intended application. Lossy compression is widely used in video. The amount of allowable distortion depends on the application. For example, users of some consumer streaming applications can tolerate higher distortion compared to users of television applications. The achievable compression ratio reflects that higher allowed / tolerated distortion can result in a higher compression ratio.
[0006] Video encoders and decoders can utilize several broad categories of techniques, such as including motion compensation, transformation, quantization, and entropy coding, some of which will be introduced below.
[0007] Historically, video encoders and decoders tended to operate on a given picture size, which in most cases was defined for an encoded video sequence (CVS), group of pictures (GOP), or similar multi-picture time frame and remained unchanged. For example, in MPEG-2, it was known for system designs to change the horizontal resolution (and thus the picture size) based on factors such as the activity of the scene, but only on I pictures, and thus typically for GOPs. For example, according to ITU-T Rec. H.263 Annex P, it was known to resample reference pictures at different resolutions within a CVS. However, here the picture size did not change, only the reference pictures were resampled, resulting in potentially only part of the picture canvas being used (in the case of downsampling), or only part of the scene being captured (in the case of upsampling). Additionally, H.263 Annex Q allows a single macroblock to be resampled up or down by a factor of 2 (in each dimension). Again, the picture size remains unchanged. The size of the macroblock is fixed in H.263 and thus does not need to be signaled.
[0008] In modern video coding and decoding, changes in the picture size in predicted pictures are increasingly becoming mainstream. For example, VP9 allows reference picture resampling and changing the resolution of the entire picture. Similarly, certain proposals have been made for VVC (including, for example, "On adaptive resolution change (ARC) for VVC" by Hendry et al., Joint Video Team document JVET-M0135-v1, January 9 - 19, 2019, which is incorporated herein by reference in its entirety), allowing the entire reference picture to be resampled to a different higher or lower resolution. In this document, it is proposed to encode different candidate resolutions in the sequence parameter set and reference them by per-picture syntax elements in the picture parameter set. Summary of the Invention
[0009] An embodiment relates to a method, system, and computer-readable medium for cross-layer alignment in encoded video data. According to one aspect, a method for cross-layer alignment in encoded video data is provided. The method may include: decoding a video bitstream having multiple layers; identifying, from the multiple layers of the decoded video bitstream, one or more sub-picture regions, the sub-picture regions including a background region and one or more foreground sub-picture regions; when it is determined to select a foreground sub-picture region, decoding and displaying an enhanced sub-picture; and when it is determined not to select a foreground sub-picture region, decoding and displaying the background region.
[0010] According to another aspect, a computer system for cross-layer alignment in encoded video data is provided. The computer system may include: one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored on at least one of the one or more storage devices for execution by at least one of the one or more processors via at least one of the one or more memories, whereby the computer system is capable of performing a method. The method may include: decoding a video bitstream having multiple layers; identifying one or more sub-picture regions from the multiple layers of the decoded video bitstream, the sub-picture regions including a background region and one or more foreground sub-picture regions; when it is determined to select a foreground sub-picture region, decoding and displaying an enhanced sub-picture; and when it is determined not to select a foreground sub-picture region, decoding and displaying the background region.
[0011] According to yet another aspect, a computer-readable medium for cross-layer alignment in encoded video data is provided. The computer-readable medium may include one or more computer-readable storage devices and program instructions stored on at least one of the one or more tangible storage devices, the program instructions being executable by a processor. The program instructions may be executable by a processor to perform a method. The method may correspondingly include: decoding a video bitstream having multiple layers; identifying one or more sub-picture regions from the multiple layers of the decoded video bitstream, the sub-picture regions including a background region and one or more foreground sub-picture regions; when it is determined to select a foreground sub-picture region, decoding and displaying an enhanced sub-picture; and when it is determined not to select a foreground sub-picture region, decoding and displaying the background region. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] These and other objects, features, and advantages will become apparent from the following detailed description of exemplary embodiments read in conjunction with the accompanying drawings. The various features of the drawings are not drawn to scale as the illustrations are for the convenience of those skilled in the art to clearly understand in conjunction with the detailed description. In the drawings:
[0013] Figure 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.
[0014] Figure 2 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.
[0015] Figure 3 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.
[0016] Figure 4 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.
[0017] Figure 5Schematic diagram of options for signaling ARC parameters according to an embodiment.
[0018] Figure 6 Example of a syntax table according to an embodiment.
[0019] Figure 7 Schematic diagram of a computer system according to an embodiment.
[0020] Figure 8 Example of a prediction structure with scalability of adaptive resolution change.
[0021] Figure 9 Example of a syntax table according to an embodiment.
[0022] Figure 10 Schematic diagram of a simplified block diagram for parsing and decoding the POC cycle and access unit count value of each access unit.
[0023] Figure 11 Schematic diagram of a video bitstream structure including multiple layers of sub-pictures.
[0024] Figure 12 Schematic diagram of the display of a selected sub-picture with enhanced resolution.
[0025] Figure 13 Block diagram of the decoding and display process of a video bitstream including multiple layers of sub-pictures.
[0026] Figure 14 Schematic diagram of a 360 video display with an enhanced layer of sub-pictures.
[0027] Figure 15 Example of the layout information of a sub-picture and its corresponding layer and the picture prediction structure.
[0028] Figure 16 Example of the layout information of a sub-picture and its corresponding layer and the picture prediction structure with spatial scalability mode of a local area.
[0029] Figure 17 Example of a syntax table for sub-picture layout information.
[0030] Figure 18 Example of a syntax table of an SEI message for sub-picture layout information.
[0031] Figure 19 Example of a syntax table for indicating the output layer and profile / tier / level information of each output layer set.
[0032] Figure 20 Example of a syntax table for indicating the output layer mode on of each output layer set.
[0033] Figure 21 It is an example of a syntax table for indicating the current sub-picture of each layer of each output layer set. Detailed implementation manners
[0034] This disclosure provides detailed embodiments of the claimed structures and methods; however, it is understood that the disclosed embodiments are merely illustrative of the claimed structures and methods that may be embodied in various forms. However, these structures and methods may be embodied in many different forms and should not be construed as limited to the exemplary embodiments described herein. Instead, these exemplary embodiments are provided to make the present disclosure more thorough and complete and to fully convey the scope to those skilled in the art. In the specification, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments.
[0035] The embodiments generally relate to the field of data processing, and more particularly, to media processing. The exemplary embodiments described below provide a system, a method, and a computer program to allow alignment across multiple layers of encoded video data. Thus, some embodiments have the ability to improve the computing field through improved video encoding and decoding.
[0036] As previously mentioned, video encoders and decoders tend to operate on a given picture size, which in most cases is defined and remains constant for an encoded video sequence (CVS), a group of pictures (GOP), or a similar multi-picture time frame. For example, in MPEG-2, it is known that system designs change the horizontal resolution (and thus the picture size) only at I pictures based on factors such as scene activity, and thus it is typically used for GOPs. For example, according to ITU-T Rec.H.263 Annex P, it is known to resample reference pictures at different resolutions within a CVS. However, the picture size here does not change, only the reference pictures are resampled, resulting in only part of the picture canvas being used (in the case of downsampling), or only part of the scene being captured (in the case of upsampling). Additionally, H.263 Annex Q allows resampling of individual macroblocks by a factor of two up or down (in each dimension). Again, the picture size remains unchanged. The size of the macroblocks is fixed in H.263, so there is no need to signal it.
[0037] However, in the context of, for example, 360 coding or certain surveillance applications, multiple semantically independent source pictures (e.g., the six cube faces of a cube-projected 360 scene, or the individual camera inputs in the case of a multi-camera surveillance setup) may require separate adaptive resolution settings to cope with different scene activities at a given point in time. In other words, the encoder may choose to use different resampling factors for different semantically independent pictures that make up the entire 360 scene or surveillance scene at a given point in time. When combined into a single picture, this in turn requires reference picture resampling to be performed on parts of the encoded pictures, and adaptive resolution coding signaling to be available. Therefore, it may be advantageous to use the available adaptive resolution coding signaling data to better align, encode, decode, and display video layers.
[0038] Figure 1 A simplified block diagram of a communication system (100) according to an embodiment of the present disclosure is shown. The communication system (100) may include at least two terminal devices (110, 120) interconnected by a network (150). For unidirectional data transmission, a first terminal device (110) may encode video data at a local location for transmission over the network (150) to another terminal device (120). The second terminal device (120) may receive the encoded video data of the other terminal device from the network (150), decode the encoded video data, and display the recovered video data. Unidirectional data transmission is more common in applications such as media services.
[0039] Figure 1 A second pair of terminal devices (130, 140) that supports two-way transmission of encoded video is shown, which may occur, for example, during a video conference. For two-way data transmission, each terminal device (130, 140) may encode video data collected at a local location for transmission over the network (150) to another terminal device. Each terminal device (130, 140) may also receive the encoded video data transmitted by the other terminal device, and may decode the encoded video data and display the recovered video data on a local display device.
[0040] In Figure 1Among them, the terminal devices (110-140) can be servers, personal computers, and smart phones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (150) represents any number of networks that transfer encoded video data between the terminal devices (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise explained below, the architecture and topology of the network (150) may be immaterial to the operation of the present disclosure.
[0041] As an application embodiment of the disclosed subject matter, Figure 2 shows the placement of video decoders and encoders in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0042] A streaming system may include an acquisition subsystem (213), which may include a video source (201) such as a digital camera, and the video source creates a stream of uncompressed video samples (202). The video sample stream (202) is depicted as a thick line as compared to the encoded video bitstream to emphasize that it is a high-data-volume video sample stream. The video sample stream (202) may be processed by a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video bitstream (204) is depicted as a thin line as compared to the video sample stream (202) to emphasize the lower-data-volume encoded video bitstream, which may be stored on a streaming server (205) for future use. One or more streaming clients (206, 208) may access the streaming server (205) to retrieve copies (207, 209) of the encoded video bitstream (204). The streaming client (206) may include a video decoder (210). The video decoder (210) decodes an incoming copy (207) of the encoded video bitstream and produces an output video sample stream (211) that can be presented on a display (212) or another presentation device (not shown). In some streaming systems, the video bitstreams (204, 207, 209) may be encoded according to certain video codec / compression standards. Examples of such standards include ITU-T Recommendation H.265. A video codec standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0043] Figure 3 is a functional block diagram of a video decoder (210) according to one or more embodiments.
[0044] A receiver (310) may receive one or more encoded video sequences to be decoded by a video decoder (210); in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel (312), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data as well as other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not shown). The receiver (310) may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory (315) may be coupled between the receiver (310) and an entropy decoder / parser (320) (hereinafter referred to as "parser"). When the receiver (310) receives data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, it may also be possible not to configure the buffer memory (315), or the buffer memory may be made smaller. Of course, for use on packet networks such as the Internet, a buffer memory (315) may be required, which may be relatively large and may advantageously have an adaptive size.
[0045] The video decoder (210) may include a parser (320) to reconstruct symbols (321) from the entropy-coded video sequences. The classes of these symbols include information for managing the operation of the video decoder (210), as well as potential information for controlling a display device such as display 212, which is not part of the decoder but may be coupled to the decoder, as Figure 3As shown in. The control information for the display device may be in the form of a supplementary enhancement information (SEI message) or a parameter set segment (not labeled) of video usability information (VUI). The parser (320) may perform parsing / entropy decoding on the received encoded video sequence. The encoding and decoding of the encoded video sequence may be performed according to video coding techniques or standards and may follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (320) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroups may include group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), and so on. The entropy decoder / parser may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0046] The parser (320) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (315) to create symbols (321).
[0047] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (321) may involve multiple different units. Which units are involved and the way they are involved may be controlled by the subgroup control information parsed by the parser (320) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (320) and multiple units below are not described.
[0048] In addition to the functional blocks already mentioned, the video decoder (210) may be conceptually divided into several functional units as described below. In practical implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide into the functional units below.
[0049] The first unit is a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives the quantized transform coefficients as symbols (321) and control information from the parser (320), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output a block including sample values, and the sample values may be input into an aggregator (355).
[0050] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed part of the current picture. Such predictive information may be provided by an intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) generates a block having the same size and shape as the block being reconstructed using the surrounding reconstructed information extracted from the (partially reconstructed) current picture (358). In some cases, the aggregator (355) adds the prediction information generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) based on each sample.
[0051] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to an inter-coded and potentially motion-compensated block. In this case, the motion compensation prediction unit (353) may access a reference picture memory (357) to extract samples for prediction. After motion compensation of the extracted samples according to the symbol (321), these samples may be added by the aggregator (355) to the output of the scaler / inverse transform unit (which is referred to as residual samples or a residual signal in this case), thereby generating output sample information. The motion compensation prediction unit obtaining prediction samples from an address within the reference picture memory may be controlled by a motion vector, and the motion vector is in the form of the symbol (321) for use by the motion compensation prediction unit, and the symbol (321) includes, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory when using sub-sample accurate motion vectors, a motion vector prediction mechanism, etc.
[0052] The output samples of the aggregator (355) may be adopted by various loop filtering techniques in a loop filter unit (356). The video compression technique may include an in-loop filter technique, and the in-loop filter technique is controlled by parameters included in the encoded video bitstream, and the parameters may be used for the loop filter unit (356) as symbols (321) from the parser (320). However, the video compression technique may also respond to meta-information obtained during decoding of a previously (in decoding order) part of the encoded picture or the encoded video sequence, and respond to previously reconstructed and loop-filtered sample values.
[0053] The output of the loop filter unit (356) can be a sample stream, which can be output to the display (212) and stored in the reference picture memory (357) for subsequent inter-picture prediction.
[0054] Once fully reconstructed, some of the encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by a parser (320)), the current picture (358) can become part of the reference picture memory (357), and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0055] The video decoder (210) can perform decoding operations according to a predetermined video compression technique recorded, for example, in the ITU-T H.265 standard. The encoded video sequence can conform to the syntax specified by the video compression technique or standard used in the sense that the encoded video sequence follows the video compression technique or standard syntax specified in the video compression technique literature or standard, especially the profile. For compliance, it is also required that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, mega samples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0056] In an embodiment, the receiver (310) can receive additional (redundant) data together with the encoded video. The additional data can be part of the encoded video sequence. The additional data can be used by the video decoder (210) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data can be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0057] Figure 4 is a functional block diagram of a video encoder (203) according to an embodiment of the present disclosure.
[0058] The video encoder (203) can receive video samples from a video source (201) (not part of the encoder), and the video source can capture video images to be encoded by the video encoder (203).
[0059] A video source (201) may provide a source video sequence in the form of a digital video sample stream to be encoded by a video encoder (203). The digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (201) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (201) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as spatial arrays of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. A person skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0060] According to an embodiment, the video encoder (203) may encode and compress pictures of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by an application. Enforcing an appropriate encoding speed is a function of a controller (450). The controller controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, the couplings are not labeled in the figure. Parameters set by the controller may include rate control related parameters (e.g., picture skip, quantizer, λ value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. A person skilled in the art can easily identify other functions of the controller (450) as they may belong to the video encoder (203) optimized for a specific system design.
[0061] Some video encoders operate in a manner that is readily recognizable to those skilled in the art as an "encoding loop". As a simple description, the encoding loop may include an encoder (hereinafter referred to as "source coder (430)") that is responsible for creating symbols based on the input picture to be encoded and reference pictures), and a (local) decoder (433) embedded in the video encoder (203). The "local" decoder (433) reconstructs the symbols in a manner similar to how the (remote) decoder creates sample data to create sample data (since in the video compression techniques contemplated by the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input into the reference picture memory (434). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well known to those skilled in the art.
[0062] The operation of the "local" decoder (433) may be the same as that of the "remote" video decoder (210) described in detail above in conjunction with Figure 3 However, briefly referring additionally to Figure 3 , when the symbols are available and the entropy encoder (445) and the parser (320) can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding part of the video decoder (210) (including the channel (312), the receiver (310), the buffer memory (315), and the parser (320)) may not be fully implemented in the local decoder (433).
[0063] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technologies can be abbreviated since they are reciprocal to the fully described decoder technologies. More detailed descriptions are only needed in certain areas and are provided below.
[0064] As part of the operation, the source encoder (430) may perform motion compensation predictive coding. Referring to one or more previously encoded frames designated as "reference frames" in the video sequence, the motion compensation predictive coding performs predictive coding on the input frame. In this way, the coding engine (432) encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, and the reference frame can be selected as the prediction reference for the input frame.
[0065] The local decoder (433) can decode the encoded video data of the frame that can be specified as a reference frame based on the symbols created by the source encoder (430). The operation of the encoding engine (432) can advantageously be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 4 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local decoder (433) replicates the decoding process that can be performed by the video decoder on the reference frame and can store the reconstructed reference frame in the reference picture memory (434). In this way, the video encoder (203) can locally store a copy of the reconstructed reference frame, which has the same content (in the absence of transmission errors) as the reconstructed reference frame to be obtained by the remote video decoder.
[0066] The predictor (435) can perform a prediction search for the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) can search the reference picture memory (434) for sample data (as a candidate reference pixel block) or some metadata that can be used as an appropriate prediction reference for the new picture, such as a reference picture motion vector, block shape, etc. The predictor (435) can operate block by block based on sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (435), it can be confirmed that the input picture can have a prediction reference obtained from multiple reference pictures stored in the reference picture memory (434).
[0067] The controller (450) can manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.
[0068] The outputs of all the above functional units can be entropy encoded in the entropy encoder (445). The entropy encoder can perform lossless compression on the symbols generated by various functional units according to techniques well known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., so as to convert the symbols into an encoded video sequence.
[0069] The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) to prepare for transmission through the communication channel (460), which can be a hardware / software link leading to a storage device that will store the encoded video data. The transmitter (440) can merge the encoded video data from the source encoder (430) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0070] The controller (450) can manage the operation of the video encoder (203). During encoding, the controller (450) can assign a certain type of encoded picture to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures can generally be assigned to any of the following picture types.
[0071] An intra picture (I picture), which can be a picture that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh (IDR) pictures. Those skilled in the art know those variants of I pictures and their respective applications and characteristics.
[0072] A predictive picture (P picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0073] A bi - predictive picture (B picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0074] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block - by - block. These blocks can be prediction - encoded with reference to other (encoded) blocks, and the other blocks are identified according to the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non - prediction - encoded, or the blocks can be prediction - encoded with reference to already - encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be prediction - encoded with reference to a previously encoded reference picture either through spatial prediction or through temporal prediction. Blocks of a B picture can be prediction - encoded with reference to one or two previously encoded reference pictures either through spatial prediction or through temporal prediction.
[0075] The video encoder (203) can perform encoding operations according to a predetermined video encoding technique or standard such as the ITU - T H.265 recommendation. In operation, the video encoder (203) can perform various compression operations, including prediction - encoding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard used.
[0076] In an embodiment, a transmitter (440) may transmit additional data and encoded video. A source encoder (430) may include such data, such as a portion of an encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set segments, etc.
[0077] Before describing certain aspects of the disclosed subject matter in more detail, some terms need to be introduced that will be referred to in the remainder of this specification.
[0078] Hereinafter, a sub-picture in some cases refers to a rectangular arrangement of samples, blocks, macroblocks, coding units, or similar entities that are semantically grouped and can be independently encoded at varying resolutions. One or more sub-pictures may form a picture. One or more encoded sub-pictures may form an encoded picture. One or more sub-pictures may be combined into a picture, and one or more sub-pictures may be extracted from a picture. In certain environments, one or more encoded sub-pictures may be combined in the compressed domain without transcoding to the sample level to become an encoded picture, and in the same or certain other cases, one or more encoded sub-pictures may be extracted from an encoded picture in the compressed domain.
[0079] Hereinafter, adaptive resolution change (ARC) refers to a mechanism that allows the resolution of a picture or sub-picture within an encoded video sequence to be changed, for example, by resampling a reference picture. Hereinafter, ARC parameters refer to the control information required to perform adaptive resolution change, which may include, for example, filter parameters, scaling factors, the resolution of the output and / or reference pictures, various control flags, etc.
[0080] The above description focuses on encoding and decoding a single semantically independent encoded video picture. Before describing the meaning of encoding / decoding multiple sub-pictures with independent ARC parameters and the additional complexity implied thereby, options for signaling ARC parameters will be described.
[0081] Reference Figure 5 , several new options for signaling ARC parameters are shown. As indicated by each option, they all have certain advantages and disadvantages from the perspectives of codec efficiency, complexity, and architecture. A video codec standard or technology may select one or more of these options, or options known in the prior art, to signal ARC parameters. These options may not be mutually exclusive and may be interchanged based on application requirements, the standard technology involved, or the choice of encoder.
[0082] Categories of ARC parameters may include:
[0083] - Upsampling / downsampling factors separated or combined in the X and Y dimensions
[0084] - Upsampling / downsampling factors with a time dimension, indicating constant rate magnification / reduction of a given number of pictures
[0085] - Either of the two above categories may involve the encoding of one or more possibly short syntax elements that may point to a table containing one or more factors.
[0086] - The combined or individual resolutions of the input picture, output picture, reference picture, and encoded picture in the X or Y dimension, in units of samples, blocks, macroblocks, CUs, or any other suitable granularity. If there are more than one resolution (e.g., one for the input picture and one for the reference picture), then in some cases, one set of values may be inferred from the other set. This may be gated, for example, using a flag. See the examples below for more details.
[0087] - "Warping" coordinates are similar to those used in H.263 Annex P, which are also in units of the above - mentioned suitable granularity. H.263 Annex P defines an efficient way to encode such warping coordinates, but other possibly more efficient ways can also be envisioned. For example, the variable - length reversible "Huffman"-type encoding of the warping coordinates of Annex P can be replaced by a binary encoding of appropriate length, where the length of the binary codeword can be derived, for example, from the maximum picture size, possibly multiplied by a certain factor and offset by a certain value to allow "warping" outside the boundaries of the maximum picture size.
[0088] - Upsampling or downsampling filter parameters. In the simplest case, there may be only a single filter for upsampling and / or downsampling. However, in some cases, it may be beneficial to allow more flexibility in filter design, and this may require signaling of the filter parameters. These parameters can be selected by an index in a list of possible filter designs, the filter can be fully specified (e.g., by a list of filter coefficients, using appropriate entropy - coding techniques), the filter can be implicitly selected by the upsampling / downsampling rate, which is in turn signaled according to any of the above - mentioned mechanisms, etc.
[0089] Thereafter, this specification assumes the encoding of a finite set of upsampling / downsampling factors (using the same factor in the X and Y dimensions), represented by a codeword. The codeword can be advantageously variable - length encoded, for example, using the Exponential - Golomb code common to certain syntax elements in video coding / decoding specifications such as H.264 and H.265.
[0090] Many similar mappings can be designed according to the needs of the application and the capabilities of the upscaling and downscaling mechanisms available in the video compression technology or standard. The table can be extended to more values. The values can also be represented by entropy coding mechanisms other than the Ext-Golomb code, such as using binary coding. This may have certain advantages when, for example, the MANE is interested in resampling factors outside of the video processing engine (most importantly, the encoder and decoder) itself. It should be noted that for the (presumably) most common cases where the resolution does not need to be changed, a shorter Ext-Golomb code can be selected; in the table above, there is only one bit. In the most common cases, this has an advantage in encoding and decoding efficiency compared to using binary codes.
[0091] The number of entries in the table and their semantics can be fully or partially configurable. For example, the basic outline of the table can be conveyed in a "high" parameter set such as a sequence or decoder parameter set. Optionally or additionally, one or more such tables can be defined in the video coding and decoding technology or standard and can be selected, for example, by a decoder or sequence parameter set.
[0092] Below, we describe how to include the upsampling / downsampling factors (ARC information) encoded as described above in the video coding and decoding technology or standard syntax. Similar considerations can be applied to one or more codewords that control the upsampling / downsampling filters. When the filter or other data structure requires a relatively large amount of data, see the discussion below.
[0093] H.263 Annex P includes the ARC information 502 in the picture header 501 in the form of four warping coordinates, specifically in the H.263 PLUS PTYPE (503) header extension. This may be a sensible design choice when a) there is a picture header available and b) it is expected that the ARC information will change frequently. However, when using H.263-style signaling, the overhead may be quite high, and because the picture header may have a transient nature, the scaling factor may not be applicable to the picture boundaries.
[0094] The JVCET-M135-v1 cited above includes ARC reference information (505) (index) located in the picture parameter set (504), which indexes a table (506) including the target resolution, and the table (506) is located within the sequence parameter set (507). By using the SPS as an interoperability negotiation point during the capability exchange, the placement of possible resolutions in the table (506) of the sequence parameter set (507) can be demonstrated. By referring to the appropriate picture parameter set (504), the resolution can vary with the picture within the limits set by the values in the table (506).
[0095] Still referring to Figure 5 , the following additional options may exist to convey ARC information in a video bitstream. Each of these options has certain advantages over the above-mentioned prior art. These options may coexist in the same video codec technology or standard.
[0096] In an embodiment, ARC information (509) such as a resampling (scaling) factor may be present in a slice header, a GOP header, a tile header, or a tile group header (hereinafter tile group header) (508). For example, as described above, if the ARC information is small, such as a single variable length ue(v) or a fixed length codeword of a few bits, this is sufficient. Placing the ARC information directly in the tile group header has the additional advantage that the ARC information can be applied to, for example, a sub-picture represented by the tile group rather than the entire picture. See also below. Additionally, even if the video compression technology or standard only contemplates changes in the adaptive resolution of the entire picture (e.g., compared to tile group-based adaptive resolution changes), from the perspective of error recovery, placing the ARC information in the tile group header has certain advantages compared to placing it in the picture header in the H.263 format.
[0097] In the same or another embodiment, the ARC information (512) itself may be present in a suitable parameter set (511), such as a picture parameter set, a header parameter set, a tile parameter set, an adaptive parameter set, etc. (the adaptive parameter set shown). More advantageously, the scope of the parameter set may not be larger than a picture, such as a tile group. By activating the relevant parameter set, the use of the ARC information is implicit. For example, when the video codec technology or standard only considers picture-based ARC, the picture parameter set or an equivalent parameter may be applicable.
[0098] In the same or another embodiment, the ARC reference information (513) may be present in the tile group header (514) or a similar data structure. The reference information (513) may refer to a subset of the ARC information (515) available in a parameter set (516) whose scope extends beyond a single picture, such as a sequence parameter set or a decoder parameter set.
[0099] Since the picture parameter set, like the sequence parameter set, can (and in some standards such as RFC3984 already does) be used for capability negotiation or announcement, it seems unnecessary to implicitly activate an additional level of the PPS indirectly from the tile group header, PPS, and SPS (as used in JVET-M0135-v1). However, if the ARC information should also apply to sub-pictures represented by, for example, tile groups, it may be a better choice to activate parameter sets limited to tile groups (such as adaptive parameter sets or header parameter sets). Additionally, if the size of the ARC information exceeds a negligible size - for example, contains filter control information such as multiple filter coefficients - then from the perspective of coding efficiency, parameters may be a better choice than directly using the header (508) because these settings can be reused by future pictures or sub-pictures by referring to the same parameter set.
[0100] When using a sequence parameter set or another higher parameter set that spans multiple pictures, certain considerations may apply:
[0101] In some cases, the parameter set storing the ARC information table (516) can be the sequence parameter set, but in other cases, the decoder parameter set is more advantageous. The decoder parameter set can have an activation range for multiple CVSs, i.e., the encoded video bitstream, i.e., all the encoded video bits from the start of the session to the end of the session. Such a range may be more appropriate because the possible ARC factors can be decoder characteristics, which may be implemented in hardware, and hardware characteristics tend not to change with any CVS (which is a set of pictures, typically one second or shorter in length in at least some entertainment systems). That is, placing the table in the sequence parameter set is explicitly included in the placement options described herein.
[0102] The ARC reference information (513) can be advantageously placed directly in the picture / slice / tile / GOP / tile group header (hereinafter the tile group header) (514), rather than in the picture parameter set as in JVET-M0135-v1. The reason is as follows. When the encoder wants to change a single value in the picture parameter set (such as the ARC reference information), it has to create a new PPS and refer to that new PPS. Suppose only the ARC reference information changes while other information (such as the quantization matrix information in the PPS) remains. Such information can be large and needs to be retransmitted to make the new PPS complete. Since the ARC reference information (513) can be a single codeword, such as an index in a table, and this is the only value that changes, retransmitting all the quantization matrix information, for example, would be cumbersome and wasteful. In this regard, from the perspective of codec efficiency, it is better to avoid indirect access through the PPS as proposed in JVET-M0135-v1. Similarly, putting the ARC reference information in the PPS has the additional drawback that since the scope of activation of the picture parameter set is the picture, the ARC information referred to by the ARC reference information (513) necessarily needs to be applied to the entire picture rather than a sub-picture.
[0103] In the same or another embodiment, the signaling of the ARC parameters can follow the detailed example outlined in Figure 6 as follows. Figure 6 Syntax diagrams in the representation used in video codec standards since at least 1993 are described. The symbols of such syntax diagrams roughly follow C-style programming. Lines in bold font represent the syntax elements present in the bitstream, and lines not in bold font generally represent the control flow or variable settings.
[0104] As an exemplary syntax structure of a header applicable to a part (possibly rectangular) of a picture, the tile group header (601) can conditionally contain the variable-length, Exp-Golomb coded syntax element dec_pic_size_idx (602) (shown in bold font). The presence of this syntax element in the tile group header can be gated based on the use of adaptive resolution (603) - here, the value of the flag is not in bold font, which means that the flag is present in the bitstream at the point where it appears in the syntax diagram. Whether adaptive resolution is used for this picture or a part thereof can be signaled in any higher-level syntax structure either inside or outside the bitstream. In the example shown, it is signaled in the sequence parameter set as described below.
[0105] Still referring to Figure 6, an excerpt of the sequence parameter set (610) is also shown. The first syntax element shown is the adaptive_pic_resolution_change_flag (611). When true, this flag may indicate the use of adaptive resolution, which in turn may require certain control information. In this example, such control information appears conditionally based on the value of this flag, and the value of this flag is based on the parameter set (612) and the if() statement in the tile group header (601).
[0106] When using adaptive resolution, in this example, what is encoded is the output resolution in samples (613). The number 613 refers to output_pic_width_in_luma_samples and output_pic_height_in_luma_samples, which together can define the resolution of the output picture. Elsewhere in the video coding / decoding technology or standard, certain limitations on either value can be defined. For example, the level definition can limit the total number of output samples, which can be the product of the values of these two syntax elements. Also, certain video coding / decoding technologies or standards, or external technologies or standards (such as system standards) may limit the number range (e.g., one or both dimensions must be divisible by a power of 2) or the aspect ratio (e.g., the width and height must be in a relationship such as 4:3 or 16:9). Such limitations can be introduced to facilitate hardware implementation or for other reasons, and are well-known in the art.
[0107] In some applications, it may be advisable for the encoder to indicate to the decoder to use a certain reference picture size instead of implicitly assuming that size to be the output picture size. In this example, the syntax element reference_pic_size_present_flag (614) gates the conditional presence of the reference picture dimensions (615) (again, this number refers to the width and height).
[0108] Finally, a table of possible decoded picture widths and heights is shown. Such a table can be represented, for example, by a table indication (num_dec_pic_size_in_luma_samples_minus1) (616). "minus1" can refer to the interpretation of the value of this syntax element. For example, if the encoded value is zero, there is one table entry. If the value is 5, there are six table entries. For each "row" in the table, the decoded picture width and height are then included in the syntax (617).
[0109] The presented table entries (617) can be indexed using the syntax element dec_pic_size_idx (602) in the tile group header, thus allowing each tile group to have a different decoded size - effectively a scaling factor.
[0110] Some video coding and decoding techniques or standards (such as VP9) support spatial scalability by implementing certain forms of reference picture resampling (signaled in a completely different way from the disclosed subject matter) in combination with temporal scalability, thereby achieving spatial scalability. Specifically, some reference pictures can be upsampled to a higher resolution using ARC-style techniques to form the basis of the spatial enhancement layer. These upsampled pictures can be corrected using normal prediction mechanisms at high resolution to add details.
[0111] The disclosed subject matter can be used in such an environment. In some cases, in the same or another embodiment, the value in the NAL unit header (such as the temporal ID field) can be used not only to indicate the temporal layer but also to indicate the spatial layer. Doing so has certain benefits for some system designs. For example, an existing selected forwarding unit (SFU) created and optimized for forwarding selection based on the NAL unit header temporal ID value for the temporal layer can be used in a scalable environment without modification. To achieve this, it may be necessary for the mapping between the coded picture size and the temporal layer to be indicated by the temporal ID field in the NAL unit header.
[0112] In some video coding and decoding techniques, an access unit (AU) can refer to one or more coded pictures, one or more slices, one or more tiles, one or more NAL units, etc., captured at a given time instance and combined into the corresponding picture / slice / tile / NAL unit bitstream. The given time instance can be the synthesis time.
[0113] In HEVC and some other video coding and decoding techniques, the picture order count (POC) value can be used to indicate a reference picture selected from among multiple reference pictures stored in the decoded picture buffer (DPB). When an access unit (AU) includes one or more pictures, slices, or tiles, each picture, slice, or tile belonging to the same AU can carry the same POC value, from which it can be derived that they are created based on the content of the same synthesis time. In other words, in the case where two pictures / slices / tiles carry the same given POC value, the POC value can indicate two pictures / slices / tiles belonging to the same AU and having the same synthesis time. Conversely, two pictures / slices / tiles with different POC values can indicate those pictures / slices / tiles belonging to different AUs and having different synthesis times.
[0114] In an embodiment of the disclosed subject matter, since the access unit may include pictures, slices or tiles having different POC values, the above rigid relationship can be relaxed. By allowing different POC values to be used within one AU, the POC values can be used to identify potentially independently decodable pictures / slices / tiles having the same presentation time. This can in turn support multiple scalable layers without changing the reference picture selection signaling (e.g., reference picture set signaling or reference picture list signaling), as described in more detail below.
[0115] However, for other pictures / slices / tiles having different POC values, it is still desirable to be able to identify the AU to which the picture / slice / tiles belong based only on the POC value. This can be achieved as described below.
[0116] In the same or other embodiments, the access unit count (AUC) may be signaled in a high-level syntax structure (e.g., NAL unit header, slice header, tile group header, SEI message, parameter set or AU delimiter). The AUC value can be used to identify which NAL units, pictures, slices or tiles belong to a given AU. The AUC value may correspond to different composition time instances. The AUC value can be equal to a multiple of the POC value. The AUC value can be calculated by dividing the POC value by an integer value. In some cases, the division operation can impose a certain burden on the decoder implementation. In such cases, making some minor restrictions in the number space of the AUC value can allow the division operation to be replaced by a shift operation. For example, the AUC value can be equal to the most significant bit (MSB) value within the range of the POC value.
[0117] In the same embodiment, the value of the POC cycle for each AU (poc_cycle_au) may be signaled in a high-level syntax structure (e.g., NAL unit header, slice header, tile group header, SEI message, parameter set or AU delimiter). The poc_cycle_au can indicate how many different and consecutive POC values can be associated with the same AU. For example, if the value of poc_cycle_au is equal to 4, then pictures, slices or tiles with POC values equal to 0 to 3 (including 0 and 3) are associated with an AU with an AUC value equal to 0, and pictures, slices or tiles with POC values equal to 4 to 7 (including 4 and 7) are associated with an AU with an AUC value equal to 1. Thus, the AUC value can be inferred by dividing the POC value by the value of poc_cycle_au.
[0118] In the same or another embodiment, the value of poc_cycle_au may be derived from information, for example, located in a Video Parameter Set (VPS), which identifies the number of spatial or SNR layers in an encoded video sequence. This possible relationship is briefly described below. Although the derivation as described above may save a few bits in the VPS and thus improve the coding / decoding efficiency, it may be advantageous to explicitly code poc_cycle_au in a suitable high-level syntax structure at a level lower than the video parameter set, in order to be able to minimize poc_cycle_au for a given small portion of the bitstream (e.g., a picture). Since the POC values (and / or the values of syntax elements that indirectly reference the POC) may be coded in a low-level syntax structure, this optimization may save more bits compared to the bits that can be saved by the above derivation process.
[0119] In the same or another embodiment, Figure 9 An example of a syntax table is shown. This syntax table is used to signal the syntax elements of vps_poc_cycle_au in the VPS (or SPS), where vps_poc_cycle_au indicates the poc_cycle_au for all pictures / slices in an encoded video sequence. This syntax table is also used to signal the syntax elements of slice_poc_cycle_au, where slice_poc_cycle_au indicates the poc_cycle_au of the current slice in the slice header. If the POC values per AU increase uniformly, vps_contant_poc_cycle_per_au in the VPS is set to 1, and vps_poc_cycle_au is signaled in the VPS. In this case, slice_poc_cycle_au is not explicitly signaled, and the AUC value per AU is calculated by dividing the POC value by vps_poc_cycle_au. If the POC values per AU do not increase uniformly, vps_contant_poc_cycle_per_au in the VPS is set to 0. In this case, vps_access_unit_cnt is not signaled, but slice_access_unit_cnt is signaled in the slice header for each slice or picture. Each slice or picture may have a different slice_access_unit_cnt value. The AUC value per AU is calculated by dividing the POC value by slice_poc_cycle_au. Figure 10 A block diagram illustrating the related workflow is shown.
[0120] In the same or other embodiments, even if the POC values of pictures, slices or tiles may be different, the pictures, slices or tiles corresponding to an AU having the same AUC value may be associated with the same decoding or output time instance. Thus, in the absence of any inter-picture parsing / decoding dependency between the pictures, slices or tiles within the same AU, all or a subset of the pictures, slices or tiles associated with that same AU may be decoded in parallel and output at the same time instance.
[0121] In the same or other embodiments, even if the POC values of pictures, slices or tiles may be different, the pictures, slices or tiles corresponding to an AU having the same AUC value may be associated with the same composition / display time instance. When the composition time is included in the container format, even if pictures correspond to different AUs, they may be displayed at the same time instance if they have the same composition time.
[0122] In the same or other embodiments, each picture, slice or tile within the same AU may have the same temporal identifier (temporal_id). All or a subset of the pictures, slices or tiles corresponding to a time instance may be associated with the same time sub-layer. In the same or other embodiments, each picture, slice or tile within the same AU may have the same or different spatial layer id (layer_id). All or a subset of the pictures, slices or tiles corresponding to a time instance may be associated with the same or different spatial layers.
[0123] Figure 8 An example of a video sequence structure with a combination of temporal_id, layer_id, POC values and AUC values with adaptive resolution change is shown. In this example, the pictures, slices or tiles in the first AU with AUC = 0 may have temporal_id = 0 and layer_id = 0 or 1, while the pictures, slices or tiles in the second AU with AUC = 1 may have temporal_id = 1 and layer_id = 0 or 1 respectively. Regardless of the values of temporal_id and layer_id, the POC value of each picture increases by 1. In this example, the value of poc_cycle_au may be equal to 2. Preferably, the value of poc_cycle_au may be set to be equal to the number of (spatially scalable) layers. Thus, in this example, the POC value increases by 2 while the AUC value increases by 1.
[0124] In the above embodiments, all or a subset of the inter-picture or inter-layer prediction structures and reference picture indications can be supported by using existing reference picture set (RPS) signaling or reference picture list (RPL) signaling in HEVC. In the RPS or RPL, the selected reference picture is indicated by signaling the POC value or the incremental value of the POC between the current picture and the selected reference picture. For the disclosed subject matter, the RPS and RPL can be used to indicate inter-picture or inter-layer prediction structures without changing the signaling, but with the following limitations. If the value of the temporal_id of the reference picture is greater than the value of the temporal_i of the current picture, the current picture may not use the reference picture for motion compensation or other prediction. If the value of the layer_id of the reference picture is greater than the value of the layer_id of the current picture, the current picture may not use the reference picture for motion compensation or other prediction.
[0125] In the same or other embodiments, the use of POC difference-based motion vector scaling for temporal motion vector prediction can be prohibited between multiple pictures within an access unit. Thus, although each picture within the access unit may have a different POC value, the motion vectors are not scaled and used for temporal motion vector prediction within the access unit. This is because reference pictures with different POCs within the same AU are considered to have the same temporal instance. Thus, in this embodiment, when the reference picture belongs to the AU associated with the current picture, the motion vector scaling function can return 1.
[0126] In the same or other embodiments, when the spatial resolution of the reference picture is different from the spatial resolution of the current picture, the use of POC difference-based motion vector scaling for temporal motion vector prediction can optionally be prohibited between multiple pictures. When motion vector scaling is allowed, the motion vectors are scaled based on the POC difference and the ratio of the spatial resolution between the current picture and the reference picture.
[0127] In the same or other embodiments, for temporal motion vector prediction, especially when poc_cycle_au has non-uniform values (when vps_contant_poc_cycle_per_au == 0), the motion vectors can be scaled based on the AUC difference rather than the POC difference. Otherwise (when vps_contant_poc_cycle_per_au == 1), the motion vector scaling based on the AUC difference may be the same as the motion vector scaling based on the POC difference.
[0128] In the same or another embodiment, when scaling the motion vector based on the AUC difference, the reference motion vector in the same AU (with the same AUC value) as the current picture is not scaled based on the AUC difference, but is used for motion vector prediction without scaling or is scaled based on the ratio of the spatial resolution between the current picture and the reference picture and then used for motion vector prediction.
[0129] In the same or other embodiments, the AUC value is used to identify the boundaries of the AU and is used for Hypothetical Reference Decoder (HRD) operations, which require timing with AU granularity for both input and output. In most cases, the decoded picture with the highest layer in the AU can be output for display. The AUC value and the layer_id value can be used to identify the output picture.
[0130] In an embodiment, a picture may include one or more sub - pictures. Each sub - picture may cover a local area or the entire area of the picture. The area supported by one sub - picture may overlap or not overlap with the area supported by another sub - picture. The area composed of one or more sub - pictures may cover or not cover the entire area of the picture. If a picture includes one sub - picture, the area supported by the sub - picture is the same as the area supported by the picture.
[0131] In the same embodiment, the sub - pictures may be encoded by a similar encoding method as that used for the encoded pictures. The sub - pictures may be encoded independently, or may be encoded according to another sub - picture or the encoded picture. The sub - pictures may or may not have any parsing dependencies on another sub - picture or the encoded picture.
[0132] In the same embodiment, the encoded sub - pictures may be included in one or more layers. The encoded sub - pictures in a layer may have different spatial resolutions. The original sub - pictures may be spatially resampled (upsampled or downsampled), encoded with different spatial resolution parameters, and included in the bitstream corresponding to the layer.
[0133] In the same or another embodiment, a sub - picture with ( W , H ) may be encoded and included in the encoded bitstream corresponding to layer 0, where W indicates the width of the sub - picture and H indicates the height of the sub - picture. A sub - picture upsampled (or downsampled) from a sub - picture with the original spatial resolution (with ( )) may be encoded and included in the encoded bitstream corresponding to layer k, where S w,k 、 S h,k indicate the resampling rates in the horizontal and vertical directions. If S w,k 、 Sh,k If the value is greater than 1, the resampling is upsampling. However, if S w,k , S h,k the value is less than 1, the resampling is downsampling.
[0134] In the same or another embodiment, the visual quality of the encoded sub - pictures in one layer may be different from the visual quality of the encoded sub - pictures in another layer of the same sub - picture or a different sub - picture. For example, the sub - pictures in layer n are i encoded with quantization parameter Q i,n , while the sub - pictures in layer m are j encoded with quantization parameter Q j,m .
[0135] In the same or another embodiment, an encoded sub - picture in one layer can be independently decodable, having no parsing or decoding dependency on the encoded sub - pictures in another layer of the same local region. A sub - picture layer that can be independently decoded without referring to another sub - picture layer of the same local region is an independent sub - picture layer. The encoded sub - pictures in an independent sub - picture layer may or may not have a decoding or parsing dependency on the previously encoded sub - pictures in the same sub - picture layer, but the encoded sub - picture may not have any dependency on the encoded pictures in another sub - picture layer.
[0136] In the same or another embodiment, an encoded sub - picture in one layer can be dependently decodable, having any parsing or decoding dependency on the encoded sub - pictures in another layer of the same local region. A sub - picture layer that can be dependently decoded with reference to another sub - picture layer of the same local region is a dependent sub - picture layer. The encoded sub - pictures in a dependent sub - picture can refer to the encoded sub - pictures belonging to the same sub - picture, the previously encoded sub - pictures in the same sub - picture layer, or both reference sub - pictures.
[0137] In the same or another embodiment, an encoded sub - picture can include one or more independent sub - picture layers and one or more dependent sub - picture layers. However, for an encoded sub - picture, there can be at least one independent sub - picture layer. An independent sub - picture layer can have a value of a layer identifier (layer_id), which can be present in the NAL unit header or another high - level syntax structure, and its value is equal to 0. The sub - picture layer with layer_id equal to 0 is the base sub - picture layer.
[0138] In the same or another embodiment, a picture may include one or more foreground sub - pictures and a background sub - picture. The area supported by the background sub - picture may be equal to the area of the picture. The area supported by the foreground sub - pictures may overlap with the area supported by the background sub - picture. The background sub - picture may be a base sub - picture layer, while the foreground sub - pictures may be non - base (enhanced) sub - picture layers. One or more non - base sub - picture layers may be decoded with reference to the same base layer. Each non - base sub - picture layer with layer_id equal to a may be decoded with reference to a non - base sub - picture layer with layer_id equal to b , where a is greater than b .
[0139] In the same or another embodiment, a picture may include one or more foreground sub - pictures with or without a background sub - picture. Each sub - picture may have its own base sub - picture layer and one or more non - base (enhanced) layers. Each base sub - picture layer may be referenced by one or more non - base sub - picture layers. Each non - base sub - picture layer with layer_id equal to a may be decoded with reference to a non - base sub - picture layer with layer_id equal to b , where a is greater than b .
[0140] In the same or another embodiment, a picture may include one or more foreground sub - pictures with or without a background sub - picture. Each encoded sub - picture in a (base or non - base) sub - picture layer may be referenced by one or more non - base layer sub - pictures belonging to the same sub - picture and one or more non - base layer sub - pictures not belonging to the same sub - picture.
[0141] In the same or another embodiment, a picture may include one or more foreground sub - pictures with or without a background sub - picture. The sub - pictures in layer a may be further divided into multiple sub - pictures in the same layer. One or more encoded sub - pictures in layer b may be decoded with reference to the divided sub - pictures in layer a .
[0142] In the same or another embodiment, a coded video sequence (CVS) may be a set of coded pictures. The CVS may include one or more coded sub - picture sequences (CSPS), where the CSPS may be a set of coded sub - pictures covering the same local area of a picture. The CSPS may have the same or different temporal resolution as the coded video sequence.
[0143] In the same or another embodiment, the CSPS may be encoded and included in one or more layers. The CSPS may include one or more CSPS layers. Decoding one or more CSPS layers corresponding to the CSPS may reconstruct a sub-picture sequence corresponding to the same local region.
[0144] In the same or another embodiment, the number of CSPS layers corresponding to the CSPS may be the same as or different from the number of CSPS layers corresponding to another CSPS.
[0145] In the same or another embodiment, the CSPS layer may have a different temporal resolution (e.g., frame rate) from another CSPS layer. The original (uncompressed) sub-picture sequence may be resampled (upsampled or downsampled) in time, encoded with different temporal resolution parameters, and included in the bitstream corresponding to the layer.
[0146] In the same or another embodiment, a sub-picture sequence with a frame rate F may be encoded and included in the encoded bitstream corresponding to layer 0, while a temporally upsampled (or downsampled) sub-picture sequence in the original sub-picture sequence with may be encoded and included in the encoded bitstream corresponding to layer k, where S t,k indicates the temporal sampling rate of layer k. If S t,k has a value greater than 1, the temporal resampling process is frame rate up conversion. However, if S t,k has a value less than 1, the temporal resampling process is frame rate downconversion.
[0147] In the same or another embodiment, when a sub-picture with a CSPS layer a is referenced by a sub-picture with a CSPS layer b for motion compensation or any inter-layer prediction, if the spatial resolution of the CSPS layer a is different from the spatial resolution of the CSPS layer b , the decoded pixels in the CSPS layer a are resampled and used as a reference. The resampling process may require upsampling filtering or downsampling filtering.
[0148] Figure 11An example video stream is shown, which includes a background video CSPS with layer_id equal to 0 and multiple foreground CSPS layers. Although the encoded sub-pictures may include one or more CSPS layers, the background area that does not belong to any foreground CSPS layer may include a base layer. The base layer may contain the background area and the foreground area, while the enhanced CSPS layer contains the foreground area. In the same area, the enhanced CSPS layer may have better visual quality than the base layer. The enhanced CSPS layer may refer to the reconstructed pixels corresponding to the same area and the motion vectors of the base layer.
[0149] In the same or another embodiment, in a video file, the video bitstream corresponding to the base layer is contained in a track, while the CSPS layers corresponding to each sub-picture are contained in separate tracks.
[0150] In the same or another embodiment, the video bitstream corresponding to the base layer is contained in a track, while the CSPS layers with the same layer_id are contained in separate tracks. In this example, the track corresponding to the layer k only includes the CSPS layer corresponding to the layer k corresponding to it.
[0151] In the same or another embodiment, each CSPS layer of each sub-picture is stored in a separate track. Each track may or may not have any parsing or decoding dependencies on one or more other tracks.
[0152] In the same or another embodiment, each track may contain a bitstream corresponding to layers i to j of the CSPS layers of all or a subset of the sub-pictures, where 0 < i =< j =< k, and k is the highest layer of the CSPS.
[0153] In the same or another embodiment, a picture includes one or more associated media data, and these associated media data include depth maps, alpha (α) maps, 3D geometric data, occupancy maps, etc. Such associated timed media data may be divided into one or more data sub-streams, and each data sub-stream corresponds to a sub-picture.
[0154] In the same or another embodiment, Figure 12 An example of a video conference based on the multi-layer sub-picture method is shown. In the video stream, it contains one base layer video bitstream corresponding to the background picture and one or more enhanced layer video bitstreams corresponding to the foreground sub-pictures. Each enhanced layer video bitstream corresponds to a CSPS layer. On the display, the picture corresponding to the base layer is displayed by default. It contains picture-in-picture (PIP) of one or more users. When a specific user is selected through the controls of the client, the enhanced CSPS layer corresponding to the selected user is decoded and displayed with enhanced quality or spatial resolution. Figure 13 A schematic diagram of this operation is shown.
[0155] In the same or another embodiment, a network intermediate box (e.g., a router) can select a subset of layers to send to a user based on its bandwidth. Picture / sub-picture organization can be used for bandwidth adaptation. For example, if the user has no bandwidth, the router strips layers or selects some sub-pictures due to their importance or based on the settings used, and can do this dynamically to adapt to the bandwidth.
[0156] Figure 14 The usage of 360 video is shown. When a spherical 360 picture is projected onto a planar picture, the projected 360 picture can be divided into multiple sub-pictures as a base layer. Enhancement layers for specific sub-pictures can be encoded and sent to the client. The decoder is capable of decoding the base layer including all sub-pictures and the enhancement layers of the selected sub-pictures. When the current viewport is the same as the selected sub-picture, the displayed picture may have a higher quality compared to the decoded sub-picture with the enhancement layer. Otherwise, the decoded picture with the base layer can be displayed in low quality.
[0157] In the same or another embodiment, any layout information for display can exist in the file as supplementary information (e.g., SEI messages or metadata). One or more decoded sub-pictures can be repositioned and displayed according to the signaled layout information. The layout information can be signaled by a streaming server or a broadcast device, or can be regenerated by a network entity or a cloud server, or can be determined by the user's custom settings.
[0158] In an embodiment, when an input picture is divided into one or more (rectangular) sub-regions, each sub-region can be encoded as an independent layer. Each independent layer corresponding to a local region can have a unique layer_id value. For each independent layer, sub-picture size and position information can be signaled, e.g., picture size (width, height), offset information (x_offset, y_offset) at the upper left corner. Figure 15 An example of the layout of the divided sub-pictures, their sub-picture size and position information, and their corresponding picture prediction structure is shown. Layout information including one or more sub-picture sizes and one or more sub-picture positions can be signaled in a high-level syntax structure (e.g., one or more parameter sets, slice headers, or tile group headers, or SEI messages).
[0159] In the same embodiment, each sub-picture corresponding to an independent layer can have a unique POC value within an AU. When indicating reference pictures in the pictures stored in the DPB by using syntax elements in the RPS or RPL structure, the POC value of each sub-picture corresponding to the layer can be used.
[0160] In the same or another embodiment, in order to indicate the (inter-layer) prediction structure, the layer_id may not be used and the POC (delta) value may be used.
[0161] In the same embodiment, a sub-picture corresponding to a layer (or local region) with a POC value equal to N may or may not be used as a reference picture for a sub-picture corresponding to the same layer (or the same local region) with a POC value equal to N+K for motion compensation prediction. In most cases, the value of the quantity K may be equal to the maximum number of (independent) layers, which may be the same as the number of sub-regions.
[0162] In the same or another embodiment, Figure 16 is shown Figure 15 an extended case. When an input picture is divided into multiple (e.g., four) sub-regions, each local region may be encoded with one or more layers. In this case, the number of independent layers may be equal to the number of sub-regions, and one or more layers may correspond to the sub-regions. Thus, each sub-region may be encoded with one or more independent layers and zero or more dependent layers.
[0163] In the same or another embodiment, in Figure 16 the input picture may be divided into four sub-regions. The upper-right sub-region may be encoded with two layers, namely layer 1 and layer 4, while the lower-right sub-region may be encoded with two layers, namely layer 3 and layer 5. In this case, layer 4 may perform motion compensation prediction with reference to layer 1, and layer 5 may perform motion compensation with reference to layer 3.
[0164] In the same or another embodiment, in-loop filtering across layer boundaries (e.g., deblocking filtering, adaptive in-loop filtering, shaper, bilateral filtering, or any deep learning-based filtering) may be (optionally) disabled.
[0165] In the same or another embodiment, motion compensation prediction or intra-block copy across layer boundaries may be (optionally) disabled.
[0166] In the same or another embodiment, boundary padding for motion compensation prediction or in-loop filtering at sub-picture boundaries may be optionally processed. A flag may be signaled in a high-level syntax structure (e.g., one or more parameter sets (VPS, SPS, PPS, or APS), slice header or tile group header, or SEI message) to indicate whether the boundary padding is processed.
[0167] In the same or another embodiment, the layout information of one or more sub-regions (or one or more sub-pictures) may be signaled in the VPS or SPS. Figure 17Examples of syntax elements in the VPS and SPS are shown. In this example, the vps_sub_picture_dividing_flag is signaled in the VPS. This flag can indicate whether one or more input pictures are divided into multiple sub-regions. When the value of vps_sub_picture_dividing_flag is equal to 0, one or more input pictures in one or more encoded video sequences corresponding to the current VPS may not be divided into multiple sub-regions. In this case, the input picture size may be equal to the encoded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples), which is signaled in the SPS. When the value of vps_sub_picture_dividing_flag is equal to 1, the one or more input pictures may be divided into multiple sub-regions. In this case, the syntax elements vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples are signaled in the VPS. The values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may be equal to the width and height of the one or more input pictures, respectively.
[0168] In the same or another embodiment, the values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may not be used for decoding, but for synthesis and display.
[0169] In the same or another embodiment, when the value of vps_sub_picture_dividing_flag is equal to 1, the syntax elements pic_offset_x and pic_offset_y corresponding to one or more specific layers may be signaled in the SPS. In this case, the encoded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples) signaled in the SPS may be equal to the width and height of the sub-region corresponding to the specific layer. Also, the position (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region may be signaled in the SPS.
[0170] In the same or another embodiment, the position information (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region may not be used for decoding, but for synthesis and display.
[0171] In the same or another embodiment, the layout information (size and position) of all or a subset of sub-regions of one or more input pictures, as well as the dependency information between layers, may be signaled in a parameter set or SEI message. Figure 18 Examples of syntax elements are shown for information indicating the layout of sub-regions, the dependency between layers, and the relationship between sub-regions and one or more layers. In this example, the syntax element num_sub_region indicates the number of (rectangular) sub-regions in the current encoded video sequence. The syntax element num_layers indicates the number of layers in the current encoded video sequence. The value of num_layers may be equal to or greater than the value of num_sub_region. When any sub-region is encoded in a single layer, the value of num_layers may be equal to the value of num_sub_region. When one or more sub-regions are encoded in multiple layers, the value of num_layers may be greater than the value of num_sub_region. The syntax element direct_dependency_flag[ i ][ j ] indicates the dependency from layer j to layer i. num_layers_for_region[ i ] indicates the number of layers associated with the i-th sub-region. sub_region_layer_id[ i ][ j ] indicates the layer_id of the j-th layer associated with the i-th sub-region. sub_region_offset_x[ i ] and sub_region_offset_y[ i ] indicate the horizontal and vertical positions of the upper left corner of the i-th sub-region, respectively. sub_region_width [ i ] and sub_region_height[ i ] indicate the width and height of the i-th sub-region, respectively.
[0172] In one embodiment, one or more syntax elements may be signaled in a high-level syntax structure (such as a VPS, DPS, SPS, PPS, APS, or SEI message), the one or more syntax elements specifying an output layer set to indicate one of a plurality of layers with or without profile level information to be output. Refer to Figure 19, a signal can be sent in the VPS to the syntax element num_output_layer_sets, which indicates the number of output layer sets (OLSs) in the encoded video sequence that refers to the VPS. For each output layer set, as many output_layer_flag as the number of output layers can be signaled.
[0173] In the same embodiment, output_layer_flag[ i ] being equal to 1 specifies the output of the i-th layer. vps_output_layer_flag[ i ] being equal to 0 specifies not to output the i-th layer.
[0174] In the same or another embodiment, one or more syntax elements can be signaled in a high-level syntax structure (such as a VPS, DPS, SPS, PPS, APS, or SEI message), and the one or more syntax elements specify the profile tier information of each output layer set. Still referring to Figure 19 , a signal can be sent in the VPS to the syntax element num_profile_tile_level, which indicates the number of profile tier information of each OLS in the encoded video sequence that refers to the VPS. For each output layer set, as many sets of syntax elements for profile tier information as the number of output layers can be signaled or an index indicating the specific profile tier information of the entries in the profile tier information can be signaled.
[0175] In the same embodiment, profile_tier_level_idx[ i ][ j ] specifies the index of the profile_tier_level( ) syntax structure applied to the j-th layer of the i-th OLS in the list of profile_tier_level( ) syntax structures in the VPS.
[0176] In the same or another embodiment, referring to Figure 20 , when the maximum number of layers is greater than 1 (vps_max_layers_minus1>0), the syntax elements num_profile_tile_level and / or num_output_layer_sets can be signaled.
[0177] In the same or another embodiment, referring to Figure 20 , the syntax element vps_output_layers_mode[ i ], which indicates the mode of the output layer signaling of the i-th output layer set, can exist in the VPS.
[0178] In the same embodiment, vps_output_layers_mode[ i ] being equal to 0 specifies that only the top layer is output using the i-th output layer set. vps_output_layer_mode[ i ] being equal to 1 specifies that all layers are output using the i-th output layer set. vps_output_layer_mode[ i ] being equal to 2 specifies that the layers to be output are the layers for which vps_output_layer_flag[ i ][ j ] is equal to 1 and the i-th output layer set is used. More values may be reserved.
[0179] In the same embodiment, depending on the value of vps_output_layers_mode[ i ] for the i-th output layer set, output_layer_flag[ i ][ j ] may or may not be signaled.
[0180] In the same or another embodiment, referring to Figure 20 , for the i-th output layer set, there may be a flag vps_ptl_signal_flag[ i ]. Depending on the value of vps_ptl_signal_flag[ i ], the profile level information of the i-th output layer set may or may not be signaled.
[0181] In the same or another embodiment, referring to Figure 21 , the number of sub-pictures max_subpics_minus1 in the current CVS may be signaled in a high-level syntax structure (such as a VPS, DPS, SPS, PPS, APS, or SEI message).
[0182] In the same embodiment, referring to Figure 21 , when the number of sub-pictures is greater than 1 (max_subpics_minus1 > 0), the sub-picture identifier sub_pic_id[i] of the i-th sub-picture may be signaled.
[0183] In the same or another embodiment, one or more syntax elements indicating the sub-picture identifiers of each layer belonging to each output layer set may be signaled in the VPS. Referring to Figure 21 , sub_pic_id_layer[i][j][k] indicates the k-th sub-picture present in the j-th layer of the i-th output layer set. Using this information, the decoder can identify which sub-pictures are decoded and output for each layer of a specific output layer set.
[0184] In an embodiment, a picture header (PH) is a syntax structure that contains syntax elements that apply to all slices of an encoded picture. A picture unit (PU) is a set of NAL units that are related to each other according to specified classification rules, are consecutive in decoding order, and exactly contain one encoded picture. A PU may contain a picture header (PH) and one or more VCL NAL units that make up the encoded picture.
[0185] In an embodiment, the SPS (RBSP) can be used in the decoding process before being referenced, including in at least one access unit (AU) where the TemporalId is equal to 0, or provided externally.
[0186] In an embodiment, the SPS (RBSP) can be used in the decoding process before being referenced, including in at least one AU in the CVS where the TemporalId is equal to 0, or provided externally, and the CVS contains one or more picture parameter sets (PPS) that reference the SPS.
[0187] In an embodiment, the SPS (RBSP) can be used in the decoding process before being referenced by one or more PPS, including in at least one PU, or provided externally, and the nuh_layer_id of the at least one PU is equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit in the CVS, and the CVS contains one or more PPS that reference the SPS.
[0188] In an embodiment, the SPS (RBSP) can be used in the decoding process before being referenced by one or more PPS, including in at least one PU, or provided externally, and the TemporalId of the at least one PU is equal to 0 and the nuh_layer_id is equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit.
[0189] In an embodiment, the SPS (RBSP) can be used in the decoding process before being referenced by one or more PPS, including in at least one PU, or provided externally, and the TemporalId of the at least one PU is equal to 0 and the nuh_layer_id is equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit in the CVS, and the CVS contains one or more PPS that reference the SPS.
[0190] In the same or another embodiment, pps_seq_parameter_set_id is the value of sps_seq_parameter_set_id specified for the SPS being referenced. The value of pps_seq_parameter_set_id can be the same in all PPSs referenced by the coded pictures in the CLVS.
[0191] In the same or another embodiment, all SPS NAL units in the CVS having a particular sps_seq_parameter_set_id value can have the same content.
[0192] In the same or another embodiment, regardless of the value of nuh_layer_id, SPS NAL units can share the same value space of sps_seq_parameter_set_id.
[0193] In the same or another embodiment, the nuh_layer_id value of an SPS NAL unit can be equal to the lowest nuh_layer_id value of the PPS NAL unit of the reference SPS NAL unit.
[0194] In an embodiment, when the SPS with nuh_layer_id equal to m is referenced by one or more PPSs with nuh_layer_id equal to n , the layer with nuh_layer_id equal to m can be the same as the layer with nuh_layer_id equal to n or the (direct or indirect) reference layer of the layer with nuh_layer_id equal to m.
[0195] In an embodiment, the PPS (RBSP) will be used in the decoding process before being referenced, including in at least one AU, or provided externally, where the TemporalId of the at least one AU is equal to the TemporalId of the PPS NAL unit.
[0196] In an embodiment, the PPS (RBSP) can be used in the decoding process before being referenced, including in at least one AU, or provided externally, where the TemporalId of the at least one AU is equal to the TemporalId of the PPS NAL unit in the CVS, and the CVS contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0197] In an embodiment, the PPS (RBSP) can be used for the decoding process before being referenced by one or more PHs (or coded slice NAL units), including in at least one PU, or provided externally, where the nuh_layer_id of the at least one PU is equal to the lowest nuh_layer_id value of the coded slice NAL units in the CVS that reference the PPS NAL unit, and the CVS contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0198] In an embodiment, the PPS (RBSP) can be used for the decoding process before being referenced by one or more PHs (or coded slice NAL units), including in at least one PU, or provided externally, where the TemporalId of the at least one PU is equal to the TemporalId of the PPS NAL unit and the nuh_layer_id is equal to the lowest nuh_layer_id value of the coded slice NAL units in the CVS that reference the PPS NAL unit, and the CVS contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0199] In the same or another embodiment, the ph_pic_parameter_set_id in the PH specifies the value of pps_pic_parameter_set_id for the PPS being referenced in use. The value of pps_seq_parameter_set_id can be the same for all PPSs referenced by the coded pictures in the CLVS.
[0200] In the same or another embodiment, all PPS NAL units in a PU having a specific pps_pic_parameter_set_id value will have the same content.
[0201] In the same or another embodiment, regardless of the nuh_layer_id value, PPS NAL units can share the same value space for pps_pic_parameter_set_id.
[0202] In the same or another embodiment, the nuh_layer_id value of a PPS NAL unit can be equal to the lowest nuh_layer_id value of the coded slice NAL units of the reference NAL unit (which references the PPS NAL unit).
[0203] In an embodiment, when the PPS with nuh_layer_id equal to m is referenced by one or more coded slice NAL units with nuh_layer_id equal to n then the nuh_layer_id equal to mThe layer with can be the same as the (direct or indirect) reference layer of the layer with nuh_layer_id equal to n or the layer with nuh_layer_id equal to m .
[0204] In an embodiment, the PPS (RBSP) will be used for the decoding process before being referenced, including in at least one AU, or provided externally, where the TemporalId of the at least one AU is equal to the TemporalId of the PPS NAL unit.
[0205] In an embodiment, the PPS (RBSP) can be used for the decoding process before being referenced, including in at least one AU, or provided externally, where the TemporalId of the at least one AU is equal to the TemporalId of the PPS NAL unit in the CVS, and the CVS contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0206] In an embodiment, the PPS (RBSP) can be used for the decoding process before being referenced by one or more PHs (or coded slice NAL units), including in at least one PU, or provided externally, where the nuh_layer_id of the at least one PU is equal to the lowest nuh_layer_id value of the coded slice NAL units in the CVS that reference the PPS NAL unit, and the CVS contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0207] In an embodiment, the PPS (RBSP) can be used for the decoding process before being referenced by one or more PHs (or coded slice NAL units), including in at least one PU, or provided externally, where the TemporalId of the at least one PU is equal to the TemporalId of the PPS NAL unit and the nuh_layer_id is equal to the lowest nuh_layer_id value of the coded slice NAL units in the CVS that reference the PPS NAL unit, and the CVS contains one or more PHs (or coded slice NAL units) that reference the PPS.
[0208] In the same or another embodiment, the ph_pic_parameter_set_id in the PH specifies the value of pps_pic_parameter_set_id for the PPS referenced in use. The value of pps_seq_parameter_set_id can be the same for all PPSs referenced by the coded pictures in the CLVS.
[0209] In the same or another embodiment, all PPS NAL units with a specific pps_pic_parameter_set_id value in the PU will have the same content.
[0210] In the same or another embodiment, regardless of the nuh_layer_id value, PPS NAL units can share the same value space of pps_pic_parameter_set_id.
[0211] In the same or another embodiment, the nuh_layer_id value of a PPS NAL unit can be equal to the lowest nuh_layer_id value of the encoded slice NAL unit of the reference NAL unit (which references the PPS NAL unit).
[0212] In an embodiment, when the PPS with nuh_layer_id equal to m is referenced by one or more encoded slice NAL units with nuh_layer_id equal to n , the layer with nuh_layer_id equal to m can be the same as the layer with nuh_layer_id equal to n or the (direct or indirect) reference layer of the layer with nuh_layer_id equal to m .
[0213] In an embodiment, when the flag no_temporal_sublayer_switching_flag is signaled in the DPS, VPS, or SPS, the TemporalId value of the PPS that references the parameter set containing the flag equal to 1 can be equal to 0, while the TemporalId value of the PPS that references the parameter set containing the flag equal to 1 can be equal to or greater than the TemporalId value of the parameter set.
[0214] In an embodiment, each PPS (RBSP) can be used for the decoding process before being referenced, including in at least one AU, or provided externally, where the TemporalId of the at least one AU is less than or equal to the TemporalId of the encoded slice NAL unit (or PH NAL unit) that references it. When the PPS NAL unit is included in an AU before the AU containing the encoded slice NAL unit of the reference PPS, there may be no VCL NAL unit enabling temporal upper layer switching, or a VCL NAL unit with nal_unit_type equal to STSA_NUT (which indicates that the picture in the VCL NAL unit can be a stepwise temporal sublayer access (STSA) picture) between the PPS NAL unit and before the encoded slice NAL unit of the reference APS.
[0215] In the same or another embodiment, the PPS NAL unit and the encoded slice NAL unit of the reference PPS (and its PH NAL unit) can be included in the same AU.
[0216] In the same or another embodiment, the PPS NAL unit and the STSA NAL unit can be included in the same AU, which is before the encoded slice NAL unit of the reference PPS (and its PH NAL unit).
[0217] In the same or another embodiment, the STSA NAL unit, the PPS NAL unit, and the encoded slice NAL unit of the reference PPS (and its PH NAL unit) can exist in the same AU.
[0218] In the same embodiment, the TemporalId value of the VCL NAL unit containing the PPS can be equal to the TemporalId value of the previous STSA NAL unit.
[0219] In the same embodiment, the picture order count (POC) value of the PPS NAL unit can be equal to or greater than the POC value of the STSA NAL unit.
[0220] In the same embodiment, the picture order count (POC) value of the encoded slice or PH NAL unit of the reference PPS NAL unit can be equal to or greater than the POC value of the PPS NAL unit being referenced.
[0221] In an embodiment, since all VCL NAL units in an AU should have the same TemporalId value, the value of sps_max_sublayer_minus1 should be the same in all layers of the encoded video sequence. The value of sps_max_sublayers_minus1 should be the same in all SPSs that are referenced by the encoded pictures in the CVS.
[0222] In an embodiment, the chroma_format_idc value of the SPSs referenced by one or more encoded pictures in layer A should be equal to the chroma_format_idc value in the SPSs referenced by one or more encoded pictures in layer B, where layer A is the direct reference layer of layer B. This is because any encoded picture should have the same chroma_format_idc value as its reference picture. In the CVS, the chroma_format_idc value of the SPSs referenced by one or more encoded pictures in layer A should be equal to the chroma_format_idc value in the SPSs referenced by one or more encoded pictures in the direct reference layer of layer A.
[0223] In an embodiment, the subpics_present_flag and sps_subpic_id_present_flag values of the SPSs referenced by one or more encoded pictures in layer A should be equal to the subpics_present_flag and sps_subpic_id_present_flag values in the SPSs referenced by one or more encoded pictures in layer B, where layer A is the direct reference layer of layer B. This is because the sub-picture layout needs to be aligned or associated across layers. Otherwise, sub-pictures with multiple layers may not be correctly extracted. In the CVS, the subpics_present_flag and sps_subpic_id_present_flag values of the SPSs referenced by one or more encoded pictures in layer A should be equal to the subpics_present_flag and sps_subpic_id_present_flag values in the SPSs referenced by one or more encoded pictures in the direct reference layer of layer A.
[0224] In an embodiment, when an STSA picture in layer A is referenced by a picture in the direct reference layer of layer A in the same AU, the picture that references the STSA should be an STSA picture. Otherwise, the temporal sublayer up-switching cannot be synchronized across layers. When an STSA NAL unit in layer A is referenced by a VCL NAL unit in the direct reference layer of layer A in the same AU, the nal_unit_type value of the VCL NAL unit that references the STSA NAL unit should be equal to STSA_NUT.
[0225] In an embodiment, when a RASL picture in layer A is referenced by a picture in the direct reference layer of layer A within the same AU, the picture that references the RASL should be a RASL picture. Otherwise, the picture cannot be correctly decoded. When a RASL NAL unit in layer A is referenced by a VCL NAL unit in the direct reference layer of layer A within the same AU, the nal_unit_type value of the VCL NAL unit that references the RASL NAL unit should be equal to RASL_NUT.
[0226] The above techniques for signaling adaptive resolution parameters can be implemented as computer software by computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 7 FIG. 700 shows a computer system that is suitable for implementing some embodiments of the disclosed subject matter.
[0227] The computer software can be encoded in any suitable machine code or computer language, creating code including instructions through mechanisms such as assembly, compilation, and linking, and the instructions can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through decoding, microcode, etc.
[0228] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0229] Figure 7 The components shown for computer system 700 are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Nor should the configuration of the components be construed as having any dependence on or requirement for any one component or combination thereof shown in the exemplary embodiments of computer system 700.
[0230] Computer system 700 may include certain human-machine interface input devices. Such human-machine interface input devices can respond to inputs from one or more human users through tactile inputs (such as keyboard input, swiping, data glove movement), audio inputs (such as sound, applause), visual inputs (such as gestures), olfactory inputs (not shown). The human-machine interface device can also be used to capture certain media, which does not have to be directly related to human conscious input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0231] The human-machine interface input device may include one or more of the following (only one of which is drawn): keyboard 701, mouse 702, touchpad 703, touch screen 710, data glove, joystick 705, microphone 706, scanner 707, camera 708.
[0232] The computer system 700 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate one or more human users' senses through, for example, haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include haptic output devices (e.g., haptic feedback through the touch screen 710, data glove, or joystick 705, but there may also be haptic feedback devices that are not used as input devices), audio output devices (e.g., speaker 709, headphones (not shown)), visual output devices (e.g., the touch screen 710 including a cathode ray tube (CRT) screen, liquid crystal display (LCD) screen, plasma screen, organic light emitting diode (OLED) screen, each of which has or does not have touch screen input function, each of which has or does not have haptic feedback function - some of which may output two-dimensional visual output or output above three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).
[0233] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable compact discs (CD / DVD ROM / RW) 720 or similar media 721 with CD / DVD, thumb drives 722, removable hard disk drives or solid state drives 723, traditional magnetic media such as tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.
[0234] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the currently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0235] The computer system 700 may also include an interface to one or more communication networks. For example, the network may be wireless, wired, or optical. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, and so on. The network also includes local area networks such as Ethernet, wireless local area network, cellular networks (including Global System for Mobile Communications (GSM), Third Generation (3G), Fourth Generation (4G), Fifth Generation (5G), Long Term Evolution (LTE), etc.), television wired or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANBus), and so on. Some networks typically require an external network interface adapter for connection to certain common data ports or peripheral buses (749) (e.g., the Universal Serial Bus (USB) port of the computer system 700); other systems are typically integrated into the core of the computer system 700 by connecting to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks, the computer system 700 can communicate with other entities. The communication can be unidirectional, only for receiving (e.g., wireless television), unidirectional only for sending (e.g., CAN bus to certain CAN bus devices), or bidirectional, e.g., via a local or wide area digital network to other computer systems. Each of the above networks and network interfaces can use certain protocols and protocol stacks.
[0236] The above-mentioned human-machine interface device, human-accessible storage device, and network interface can be connected to the core 740 of the computer system 700.
[0237] The core 740 may include one or more central processing units (CPUs) 741, a graphics processing unit (GPU) 742, a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) 743, a hardware accelerator 744 for specific tasks, and so on. These devices, as well as read-only memory (ROM) 745, random access memory (RAM) 746, internal mass storage (e.g., internal non-user-accessible hard disk drive, solid state drive (SSD), etc.) 747, and so on, can be connected via a system bus 748. In some computer systems, the system bus 748 can be accessed in the form of one or more physical plugs for expansion with additional central processing units, graphics processing units, and so on. Peripherals can be directly attached to the system bus 748 of the core or connected via a peripheral bus 749. The architecture of the peripheral bus includes Peripheral Component Interconnect (PCI), Universal Serial Bus USB, and so on.
[0238] The CPU 741, GPU 742, FPGA 743, and accelerator 744 can execute certain instructions that, when combined, can form the aforementioned computer code. The computer code can be stored in the ROM 745 or RAM 746. Transitional data can also be stored in the RAM 746, while permanent data can be stored in, for example, the internal mass storage 747. Fast storage and retrieval of any memory device can be achieved by using a cache memory that can be closely associated with one or more of the CPU 741, GPU 742, mass storage 747, ROM 745, RAM 746, etc.
[0239] The computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure or can be of the kind well known and available to those skilled in the field of computer software.
[0240] By way of example and not limitation, a computer system having the architecture 700, and particularly the core 740, can provide the functionality of a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the aforementioned user-accessible mass storage and specific memories of the non-volatile core 740, such as the core internal mass storage 747 or ROM 745. The software implementing the various embodiments of the present application can be stored in such devices and executed by the core 740. Depending on specific requirements, the computer-readable medium can include one or more storage devices or chips. The software can cause the core 740, and particularly the processors therein (including the CPU, GPU, FPGA, etc.), to execute the specific processes or specific portions of the specific processes described herein, including defining data structures stored in the RAM 746 and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality that is logically hardwired or otherwise embodied in circuitry (e.g., the accelerator 744) that can operate in place of or in conjunction with the software to execute the specific processes or specific portions of the specific processes described herein. In appropriate instances, references to software can include logic and vice versa. In appropriate instances, references to the computer-readable medium can include circuitry (such as an integrated circuit (IC)) that stores the software for execution, circuitry that embodies the logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0241] Although the present disclosure has described multiple exemplary embodiments, various changes, permutations, and various equivalent substitutions of the embodiments fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.
Claims
1. A method for video decoding, characterized in that, the method comprises: decoding a video bitstream having multiple layers; identifying one or more sub-picture regions from the multiple layers of the decoded video bitstream; and aligning the one or more sub-pictures across the multiple layers, wherein, based on the network abstraction layer NAL unit having a nal_unit_type equal to STSA_NUT, the TemporalID of the NAL unit in the reference picture list is different from the TemporalID of the NAL unit having a nal_unit_type equal to STSA_NUT, wherein the NAL unit in the reference picture list includes a picture parameter set PPS NAL unit, and the NAL unit having a nal_unit_type equal to STSA_NUT includes an encoded slice NAL unit.
2. The method according to claim 1, characterized in that, further comprising: selecting a subset of layers from the multiple layers based on bandwidth; and transmitting the selected subset of layers.
3. The method according to claim 1, characterized in that, disabling in-loop filtering across the boundaries between the layers.
4. The method according to claim 1, characterized in that, disabling motion compensation prediction or intra-block copy across the boundaries between the layers.
5. The method according to claim 1, characterized in that, further comprising: processing boundary padding for motion compensation prediction or in-loop filtering at the boundaries of the sub-picture regions.
6. The method according to any one of claims 1-5, characterized in that, signaling layout information of the sub-picture regions through parameter set data.
7. The method according to claim 6, characterized in that, the layout information includes the size and position associated with the sub-picture regions.
8. The method according to claim 6, characterized in that, relocating and displaying one or more of the sub-picture regions based on the layout information.
9. The method according to any one of claims 1-5, characterized in that, each sub-picture region is encoded as an independent layer corresponding to a local region having a unique layer identification value.
10. The method according to claim 9, characterized in that, each sub-picture region corresponding to the independent layer has a unique picture order count value within an access unit.
11. A method for video encoding for generating a video bitstream, characterized in that, the method comprises: determining one or more sub-picture regions and encoding the one or more sub-picture regions into multiple layers; aligning the one or more sub-pictures across the multiple layers, Among them, based on the network abstraction layer NAL unit having a nal_unit_type equal to STSA_NUT, the TemporalID of the NAL unit in the reference picture list is different from the TemporalID of the NAL unit having a nal_unit_type equal to STSA_NUT, where the NAL unit in the reference picture list includes a picture parameter set PPS NAL unit, and the NAL unit having a nal_unit_type equal to STSA_NUT includes an encoded slice NAL unit.
12. A method for storing or transmitting a video bitstream, characterized in that, the video bitstream is generated according to the video coding method described in claim 11, or the video bitstream is decoded based on the video decoding method described in any one of claims 1-10.
13. A computer system, characterized in that, the computer system includes: one or more computer-readable non-volatile storage media configured to store computer program code; and one or more computer processors configured to access the computer program code and operate according to the instructions of the computer program code to perform the method described in any one of claims 1-12.
14. A video decoding device, characterized in that, the device includes: a first decoding module configured to decode a video bitstream having multiple layers; an identification module configured to identify one or more sub-picture regions from the multiple layers of the decoded video bitstream; and a second decoding and first display module configured to align the one or more sub-pictures across the multiple layers, Among them, based on the network abstraction layer NAL unit having a nal_unit_type equal to STSA_NUT, the TemporalID of the NAL unit in the reference picture list is different from the TemporalID of the NAL unit having a nal_unit_type equal to STSA_NUT, where the NAL unit in the reference picture list includes a picture parameter set PPS NAL unit, and the NAL unit having a nal_unit_type equal to STSA_NUT includes an encoded slice NAL unit.
15. A video encoder, at least including a local decoder, characterized in that, the local decoder is used to execute the video decoding method described in any one of claims 1 to 10.
16. A non-volatile computer-readable storage medium, characterized in that, stores a computer program for video decoding, and the computer program is configured to cause one or more computer processors to execute the method described in any one of claims 1-12.
17. A method for storing encoded video data, characterized in that, the method includes: reconstructing symbols to create sample data based on the video decoding method described in any one of claims 1 to 10; and inputting the reconstructed sample stream into a reference picture memory.
Citation Information
Patent Citations
Video encoding and decoding of foreground and background wherein picture is divided into slice
CN1593065A
Method and apparatus for video coding and decoding
US20140301463A1
Layered content delivery for virtual and augmented reality experiences
US20180089903A1