Systems and methods for generic video coding

By indicating whether to enable or disable surround motion compensation in the video data and performing corresponding surround offset compensation, the coding efficiency problem of existing video coding standards on multiple video formats is solved, and efficient coding of 360-degree video is achieved.

CN121967680APending Publication Date: 2026-05-01INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2020-09-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing video coding standards struggle to efficiently support the coding efficiency of various video formats, such as standard dynamic range video, high dynamic range video, and omnidirectional video, when handling surround motion compensation.

Method used

By including indications to enable or disable surround motion compensation in the video data and using surround offset for compensation, the encoding and decoding of sub-images are achieved, including the application of surround motion compensation and geometric filling.

Benefits of technology

It improves the efficiency of video encoding and supports efficient encoding of various video formats, especially the encoding performance of 360-degree video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967680A_ABST
    Figure CN121967680A_ABST
Patent Text Reader

Abstract

Systems and methods for generic video coding. Systems, methods, and tools associated with video coding are described. Signaling of certain syntax elements may be moved from a slice header to a picture header and / or a layer access unit separator (AUD). Dependencies between the AUD and one or more parameter sets may be explored. A syntax element may be signaled to implement wrap-around motion compensation for certain sub-pictures and to specify wrap-around motion compensation offsets for that sub-picture.
Need to check novelty before this filing date? Find Prior Art

Description

Systems and methods for general video coding

[0001] Cross-reference to related applications This application claims the benefit of U.S. Provisional Application No. 62 / 902,647, filed September 19, 2019, and U.S. Provisional Application No. 62 / 911,797, filed October 7, 2019, the contents of which are incorporated herein by reference. Background Technology

[0002] Video coding standards are constantly evolving to improve coding efficiency (e.g., compression efficiency) and to support various video formats, such as standard dynamic range video, high dynamic range video, omnidirectional video, and projection. Summary of the Invention

[0003] This document describes systems, methods, and tools associated with general video coding. A video decoding apparatus as described herein may include one or more processors configured to acquire video data, determine whether to apply surround motion compensation to a first sub-image of an encoded picture based on the video data, and perform surround motion compensation on the first sub-image in response to determining that surround motion compensation should be applied to the first sub-image. The video data may include information about the first sub-image, including, for example, an indication of whether surround motion compensation is enabled for the first sub-image. The encoded picture may also include a second sub-image, and the video data may include information about the second sub-image, including, for example, an indication of whether surround motion compensation is enabled for the second sub-image. For example, the video data may include information indicating that surround motion compensation is enabled for the first sub-image and disabled for the second sub-image. The video data may also include information specifying (e.g., when surround motion compensation is enabled for the first or second sub-image) a corresponding surround offset associated with the first and second sub-images, and one or more processors of the video decoding apparatus may be configured to perform surround compensation based on the surround offset associated with the first or second sub-image.

[0004] In the example, the information in the video data may include a Picture Parameter Set (PPS) syntax element indicating that surround motion compensation is enabled and a Sequence Parameter Set (SPS) syntax element indicating that a first sub-picture (or a second sub-picture) will be treated as a picture. In the example, performing surround motion compensation on the first sub-picture may include performing bilinear interpolation of luminance samples on the first sub-picture based on a surround offset associated with it. In the example, the surround motion compensation described herein may be performed in the horizontal direction, and the coded picture including the first and second sub-pictures may be associated with a 360-degree video.

[0005] The video encoding apparatus described herein may include one or more processors configured to encode a picture, obtain information indicating whether to enable surround motion compensation for a first sub-picture of the encoded picture and a surround offset associated with the first sub-picture, and form a set of encoded data including the encoded picture and the obtained information. In an example, the encoded picture may also include a second sub-picture, and the obtained information may further indicate that surround motion compensation is enabled for the first sub-picture and disabled for the second sub-picture. In an example, the obtained information may include a Picture Parameter Set (PPS) syntax element indicating that surround motion compensation is enabled and a Sequence Parameter Set (SPS) syntax element indicating that the first sub-picture will be treated as a picture. In an example, one or more processors of the video encoding apparatus may be further configured to transmit the set of encoded data to a receiving device. Attached Figure Description

[0006] Figure 1 shows an exemplary video encoder.

[0007] Figure 2 shows an exemplary video decoder.

[0008] Figure 3 shows a block diagram of an example system in which various aspects and examples are implemented.

[0009] Figure 4 is an illustration of an exemplary image divided into tiles and slices.

[0010] Figure 5 is a diagram showing an exemplary sub-image grid that can be used to indicate sub-image IDs.

[0011] Figure 6 is an illustration showing an example of applying wrapping to original and merged images.

[0012] Figure 7 is a diagram illustrating an example of applying geometric filling to an isometric projection format (ERP).

[0013] Figure 8 is an illustration of an example of sub-picture wrapping within an image.

[0014] Figure 9 is a diagram illustrating an example of sub-image wrap-around filling.

[0015] Figure 10 is a diagram showing an example of a sub-image grid.

[0016] Figure 11 is a diagram illustrating an example of a layer-based Gradual Decoding Refresh (GDR) image.

[0017] Figure 12A is a system diagram illustrating an exemplary communication system that can be implemented in one or more of the disclosed embodiments.

[0018] Figure 12B is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that can be used within the communication system shown in Figure 12A, according to one embodiment.

[0019] Figure 12C is a system diagram illustrating an exemplary radio access network (RAN) and an exemplary core network (CN) that can be used within the communication system shown in Figure 12A, according to one embodiment.

[0020] Figure 12D is a system diagram illustrating another exemplary RAN and another exemplary CN that can be used within the communication system shown in Figure 12A, according to one embodiment. Detailed Implementation

[0021] A detailed description of exemplary embodiments will now be described with reference to the various accompanying drawings. Although this specification provides detailed examples of possible specific implementations, it should be noted that the details are intended to be exemplary and in no way limit the scope of this application.

[0022] This application describes multiple aspects, including tools, features, examples, models, methods, etc. Many of these aspects are described in a particular manner, and are generally described in a way that may sound restrictive, at least to illustrate individual characteristics. However, this is for clarity and does not limit the application or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in previous filings.

[0023] The aspects described and contemplated in this patent application can be implemented in many different forms. Figures 1 through 12D described herein provide some examples, but other examples are considered, and the discussion of Figures 1 through 12D does not limit the breadth of implementations. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having a bitstream generated according to any of the methods stored thereon.

[0024] In this application, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "pixel" and "sample," and the terms "image," "picture," and "frame." Generally, but not necessarily, the term "reconstruction" is used at the encoding end, while "decoding" is used at the decoding end.

[0025] This document describes various methods, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or purpose of specific steps and / or actions may be modified or combined. Additionally, in various examples, terms such as "first," "second," etc., may be used to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply a modification of the order of operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and may occur, for example, before, during, or in overlapping time periods of the second decoding.

[0026] The various methods and other aspects described in this application can be used to modify modules (e.g., decoding modules) of the video encoder 100 and decoder 200, as shown in Figures 1 and 2. Furthermore, these aspects are not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations (whether pre-existing or future-developed) and any extensions to such standards and recommendations (including VVC and HEVC). Unless otherwise specified or technically excluded, the aspects described in this application may be used individually or in combination.

[0027] Various numerical values ​​are used in this application, such as a sub-image grid size of 4×4 and values ​​ranging from 0 to 254, and so on. Specific values ​​are for illustrative purposes, and the aspects described are not limited to these specific values.

[0028] Figure 1 shows encoder 100. Variations of this encoder 100 are envisioned, but for clarity, encoder 100 is described below without describing all anticipated variations.

[0029] Before encoding, the video sequence may undergo pre-coding (101), for example, by applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or by performing remapping of the input image components to obtain a signal distribution that is more resilient to compression (e.g., histogram equalization using one of the color components). Metadata may be associated with pre-processing and appended to the bitstream.

[0030] In encoder 100, the image is encoded by encoder elements as described below. The image to be encoded is partitioned (102) and processed in units such as CUs. For example, each unit is encoded using either an intra-frame mode or an inter-frame mode. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (160). In inter-frame mode, motion estimation (175) and compensation (170) are performed. The encoder determines (105) which of the intra-frame mode or inter-frame mode is used to encode the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (110) the prediction block from the original image block.

[0031] The predicted residual is then transformed (125) and quantized (130). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (145) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., encode the residual directly without applying the transform or quantization process.

[0032] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (155) to reconstruct the image block. A loop filter (165) is applied to the reconstructed image to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coded artifacts. The filtered image is stored in a reference image buffer (180).

[0033] Figure 2 shows a block diagram of a video decoder 200. In decoder 200, the bitstream is decoded by decoder elements, as described below. Video decoder 200 generally performs a decoding process that is the reverse of the encoding process described in Figure 1. Encoder 100 typically also performs video decoding as part of the encoding of video data.

[0034] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 100. First, entropy decoding (230) is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. Image partitioning information indicates how the image should be partitioned. Therefore, the decoder can partition (235) the image based on the decoded image partitioning information. The transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (255) to reconstruct the image blocks. Prediction blocks (270) can be obtained from intra-frame prediction (260) or motion-compensated prediction (i.e., inter-frame prediction) (275). A loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference image buffer (280).

[0035] The decoded image can also undergo post-decoding processing (285), such as inverse color transformation (e.g., a transformation from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping of the remapping process performed in the pre-encoding process (101). Post-decoding processing can utilize metadata derived in the pre-encoding process and signaled in the bitstream.

[0036] Figure 3 illustrates a block diagram of an example system implementing various aspects and examples therein. System 300 may be embodied as a device including the various components described below and configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 300 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing and encoder / decoder elements of system 300 are distributed across multiple ICs and / or discrete components. In various examples, system 300 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, system 300 is configured to implement one or more aspects of the aspects described in this document.

[0037] System 300 includes at least one processor 310 configured to execute instructions loaded thereon for implementing various aspects, such as those described in this document. Processor 310 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 300 includes at least one memory 320 (e.g., a volatile memory device and / or a non-volatile memory device). System 300 includes a storage device 340 that may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 340 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0038] System 300 includes an encoder / decoder module 350 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 350 may include its own processor and memory. The encoder / decoder module 350 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both of an encoding module and a decoding module. Furthermore, the encoder / decoder module 350 may be implemented as a standalone element of system 300, or may be incorporated within processor 310 as a combination of hardware and software known to those skilled in the art.

[0039] Program code to be loaded onto processor 310 or encoder / decoder 350 to execute the various aspects described in this document may be stored in storage device 340 and subsequently loaded onto memory 320 for execution by processor 310. According to various examples, one or more of processor 310, memory 320, storage device 340, and encoder / decoder module 350 may store one or more items from various categories during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded or partially decoded video, bitstreams, matrices, variables, and intermediate or final results of processing equations, formulas, operations, and operational logic.

[0040] In some examples, the memory within processor 310 and / or encoder / decoder module 350 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, external memory (e.g., the processing device may be processor 310 or encoder / decoder module 350) is used for one or more of these functions. External memory may be memory 320 and / or storage device 340, such as volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external volatile memory such as RAM is used as working memory for video encoding and decoding operations, such as MPEG-2 (MPEG stands for Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Various Video Coding).

[0041] Inputs to the components of system 300 may be provided via various input devices as shown in box 360. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Other examples not shown in Figure 3 include composite video.

[0042] In various examples, the input device of box 360 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some examples), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF section of various examples includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various functions among these functions, including, for example, down-converting received signals to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various examples rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.

[0043] Furthermore, the USB and / or HDMI terminals may include corresponding interface processors for connecting system 300 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within processor 310. Similarly, aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface IC or within processor 310. Demodulated streams, error-corrected streams, and demultiplexed streams are provided to various processing elements, including, for example, processor 310 and encoder / decoder 350, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on the output device.

[0044] Various components of system 300 can be housed within an integrated housing. Within the integrated housing, various components can be interconnected using a suitable connection arrangement 370 (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards) and data can be transferred between these components.

[0045] System 300 includes a communication interface 380 capable of communicating with other devices via a communication channel 382. The communication interface 380 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 382. The communication interface 380 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 382 may be implemented, for example, in a wired and / or wireless medium.

[0046] In various examples, wireless networks (such as Wi-Fi networks), for example IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers), are used to stream or otherwise provide data to system 300. Wi-Fi signals in these examples are received via a communication channel 382 and a communication interface 350 suitable for Wi-Fi communication. The communication channel 382 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other cross-platform communications. Other examples use a set-top box to provide streaming data to system 300, delivering data via an HDMI connection to input box 360. Still other examples use an RF connection to input box 360 to provide streaming data to system 300. As mentioned above, various examples provide data in a non-streaming manner. Additionally, various examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0047] System 300 can provide output signals to various output devices, including a display 392, a speaker 394, and other peripheral devices 396. Various examples of the display 392 include one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 392 can be used in televisions, tablets, laptops, cellular phones (mobile phones), or other devices. The display 392 can also be integrated with other components (e.g., as in a smartphone) or standalone (e.g., an external monitor for a laptop). In various examples, other peripheral devices 396 include one or more of a standalone digital video disc (or digital universal disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various examples use one or more peripheral devices 396 that provide functionality based on the output of system 300. For example, a disc player performs the function of playing the output of system 300.

[0048] In various examples, signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols capable of device-to-device control with or without user intervention is used to transmit control signals between system 300 and display 392, speaker 394, or other peripheral devices 396. Output devices can be communicatively coupled to system 300 via dedicated connections through corresponding interfaces 330, 332, and 334. Alternatively, output devices can be connected to system 300 via communication interface 380 using communication channel 382. Display 392 and speaker 394 can be integrated into a single unit with other components of system 300 in electronic devices such as televisions. In various examples, display interface 330 includes display drivers, such as timing controller (TCon) chips.

[0049] Alternatively, for example, if the RF section of input 370 is part of a separate set-top box, the display 392 and speaker 394 may be separate from one or more other components. In various examples where the display 392 and speaker 394 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0050] These examples can be executed by processor 310 or by computer software implemented by hardware or a combination of hardware and software. As a non-limiting example, these examples can be implemented by one or more integrated circuits. As a non-limiting example, memory 320 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 310 can be of any type suitable for the technical environment and can encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0051] Various specific implementations participate in decoding. As used in this application, "decoding" may encompass all or part of a process performed, for example, on a received encoded sequence, to produce a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such a process also includes, or alternatively includes, processes performed by a decoder of various embodiments described in this application, such as receiving instructions for sub-picture level surround motion compensation, performing luminance sample bilinear interpolation, etc.

[0052] As another example, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" specifically refers to a subset of operations or broadly refers to a wider decoding process will be clear based on the specific context of the description and is believed to be well understood by those skilled in the art.

[0053] Various specific implementations participate in encoding. In a manner similar to the discussion above regarding “decoding,” the term “encoding,” as used herein, can encompass, for example, all or part of the processes performed on an input video sequence to produce an encoded bitstream. In various examples, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various examples, such processes also include, or alternatively include, processes performed by an encoder of the various embodiments described herein, such as determining whether wraparound motion compensation (e.g., geometric filling) should be enabled or disabled for individual sub-pictures.

[0054] As another example, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of differential decoding and entropy decoding. Whether the phrase "encoding process" specifically refers to a subset of operations or broadly refers to a wider encoding process will be clear based on the specific context of the description and is believed to be well understood by those skilled in the art.

[0055] Note that the syntax elements used in this article (e.g., subpic_wraparound_enabled_flag, subpic_ref_wraparound_offset_minus1, etc.) are descriptive terms. Therefore, they do not preclude the use of other syntax element names.

[0056] When the accompanying drawings are presented as flowcharts, it should be understood that block diagrams of the corresponding devices are also provided. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that flowcharts of the corresponding methods / processes are also provided.

[0057] Various examples involve rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often taking into account computational complexity constraints. Rate-distortion optimization is generally formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist to solve the rate-distortion optimization problem. For example, these approaches may be based on extensive testing of all encoding options, including all considered modes or values ​​of encoding parameters, and a complete evaluation of their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to reduce encoding complexity, particularly for the computation of approximate distortion based on prediction or prediction of the residual signal rather than the reconstructed residual signal. A hybrid of these two approaches can also be used, such as by using approximate distortion for only some of the possible encoding options and full distortion for others. Other methods evaluate only a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding costs and associated distortion.

[0058] The specific embodiments and aspects described herein may be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single form of specific embodiment (e.g., discussed only as a method), specific embodiments of the discussed features may be implemented in other forms (e.g., apparatus or program). Apparatus may be implemented, for example, in suitable hardware, software, and firmware. Methods may be implemented, for example, in a processor that generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0059] References to “an example” or “an example” or “an embodiment” or “an embodiment” and their other variations mean that a particular feature, structure, characteristic, etc., described in connection with that example is included in at least one example. Therefore, the phrases “in an example” or “in the example” or “in an embodiment” or “in the embodiment” and any other variations appearing in various places throughout this application do not necessarily refer to the same example.

[0060] Additionally, this application may involve "determining" various types of information. Determining information may include, for example, one or more of the following: estimation information, calculation information, prediction information, or information retrieved from memory.

[0061] Furthermore, this application may relate to "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or more of these.

[0062] Furthermore, this application may relate to "receiving" various types of information. Like "access," "receiving" is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) or more. Moreover, "receiving" typically involves one or more of the following during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0063] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such phrases are intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many items as possible listed.

[0064] Moreover, as used herein, the term "signaling" refers to (among other things) instructing the corresponding decoder to do something. For example, in some examples, the encoder signals a specific parameter among several parameters used for sub-picture encoding. Thus, in one example, the same parameter is used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit (explicit signaling) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter and others, signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various examples by avoiding the transmission of any actual function. It should be understood that signaling can be implemented in various ways. For example, in various examples, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the verb form of the term "signal" was mentioned above, the term "signal" can also be used as a noun in this article.

[0065] It will be apparent to those skilled in the art that specific implementations can generate various signals formatted to carry, for example, storable or transmissible information. The information may include, for example, instructions for performing a method or data generated by one of the specific implementations. For example, a signal may be formatted to carry a bitstream of the example. Such signals may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of a spectrum) or baseband signals. Formatting may include, for example, encoding a data stream and modulating the carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is known that signals can be transmitted via various wired or wireless links. The signal may be stored on a processor-readable medium.

[0066] The video processing apparatus described herein can be configured to divide an image into one or more rows and / or columns of tiles and / or one or more sub-images. A tile may include a sequence of coding tree units (CTUs) that can cover a rectangular area of ​​the image. In the example, a tile may be further divided into one or more bricks, and each brick may include one or more rows of CTUs for the tile. A sub-image may include one or more slices that can collectively cover an area of ​​the image (e.g., a rectangular area). Slices may be rectangular slices, raster scan slices, etc. Raster scan slices (e.g., generated and / or used in raster scan slice mode) may include one or more tiles (e.g., a sequence of tiles) that can be derived via a tile raster scan of the image. A rectangular slice (e.g., generated and / or used in rectangular slice mode) may include one or more bricks that can collectively form an area of ​​the image (e.g., a rectangular area). Bricks within a rectangular slice may be arranged based on the order of the brick raster scans of the corresponding slice. Figure 4 illustrates an example of an image divided into sub-images, slices (e.g., rectangular slices), tiles, and coding units (e.g., CTUs).

[0067] The video processing apparatus described herein can be configured to transmit (e.g., if the video processing apparatus includes a video encoder) or receive (e.g., if the video processing apparatus includes a video decoder) a sequence parameter set (PPS) and / or a picture parameter set (PPS). The SPS may include syntax elements (e.g., parameters) defining subpicture grids of pictures, syntax elements (e.g., subpic_treated_as_pic_flag) indicating whether subpictures of encoded pictures (e.g., in an encoded video sequence (CVS)) can be considered as pictures in the decoding process (e.g., in addition to loop filtering operations), etc. The PPS may include syntax elements (e.g., parameters) defining tile and / or brick grids, syntax elements indicating whether wrap-around motion compensation (e.g., horizontal wrap-around motion compensation) is enabled, etc.

[0068] SPS can, for example, use a grid to specify the size and / or position of a subpicture. Figure 5 shows an example of a subpicture grid that can be used to indicate a subpicture identifier (ID) (e.g., using numerical values ​​such as 0, 1, ... 5). As shown in Figure 5, an coded image can be divided (e.g., segmented) into a grid. The number of rows and columns of the grid can be determined based on the size of the grid elements and / or the size of the coded image. Table 1 below includes exemplary syntax for signaling the subpicture ID (e.g., sub_pic_id[i][j]) at the (i, j)th grid position.

[0069] Table 1 - Exemplary SPS Syntax .

[0070] The video processing apparatus described herein can be configured to perform wraparound motion compensation when processing images. Such wraparound motion compensation can be performed, for example, in the horizontal direction. The SPS and / or PPS can include elements indicating whether wraparound motion compensation is enabled. For example, the SPS can include a first parameter (e.g., sps_ref_wraparound_enabled_flag) indicating whether horizontal wraparound motion compensation is enabled or disabled for inter-frame prediction (e.g., by setting sps_ref_wraparound_enabled_flag to 1 or 0, respectively). The SPS can also include a second parameter (e.g., sps_ref_wraparound_offset_minus1 plus 1) that can specify the offset that can be used to calculate the horizontal wraparound position.

[0071] The video processing apparatus described herein can be configured to perform geometric infilling (e.g., horizontal geometric infilling) when surround motion compensation (e.g., horizontal surround motion compensation) is enabled. Figure 6 illustrates an exemplary geometric infilling process for a 360° video in an equal rectangular projection format (ERP). As shown, the video processing apparatus can be configured to fill samples at positions A, B, C, D, E, and / or F (e.g., along the left and / or right boundaries of the image) with samples at positions D', E', F', A', B', and / or C'. Along the top boundary, the video processing apparatus can be configured to fill samples at positions G, H, I, and / or J with samples at positions I', J', G', and / or H'. Along the bottom boundary, the video processing apparatus can be configured to fill samples at positions K, L, M, and / or N with samples at positions M', N', K', and / or L'.

[0072] The video processing apparatus described herein can be configured to support Access Unit Delimiters (AUDs), which can be signaled in the video bitstream and / or AUD Network Abstraction Layer (NAL) units (e.g., for standards-compliant video). Table 2 below illustrates exemplary AUD syntax. The syntax may contain elements (e.g., pic_type) indicating slice_type values ​​that can exist in the encoded picture. Furthermore, access units (e.g., each access unit) may begin with an AUD NAL unit in the syntax and / or one (e.g., at most one) AUD NAL unit may exist in the layer access units according to the syntax.

[0073] Table 2 - Exemplary AUD Syntax .

[0074] In the example (e.g., when an AUD is enforced on (e.g., each) access unit or AU), one or more syntax elements can be signaled in the AUD (e.g., instead of in the slice header), such as signaling in the slice header by constraint to encode those syntax elements with the same value within the picture, to reduce signaling overhead. The AUD and parameter set can be interdependent, and this interdependence can be explored to improve coding efficiency. In the example (e.g., in subpicture coding, when the left and / or right boundaries of the ERP picture are not connected), it might be desirable to enable or disable wrap-around motion compensation (e.g., geometric filling) for individual subpictures. Mechanisms that do not allow enabling or disabling wrap-around motion compensation for individual subpictures (e.g., by signaling only a wrap-around enable flag in the SPS) may be insufficient.

[0075] The video processing apparatus described herein can be configured to process video content using various coding tools and / or High-Level Syntax (HLS). Coding tools can facilitate intra-frame prediction, inter-frame prediction, transform, quantization, entropy coding, cyclic filtering, and more. HLS can support the partitioning of images, sub-images, slices, tiles, and bricks (e.g., for parallelization), and applications such as 360-degree video view-related processing. HLS can also support features such as scalable video coding, reference image resampling (RPR), and progressive decode refresh (GDR).

[0076] The slice header associated with an encoded image may contain one or more syntax elements such as slice_pic_parameter_set_id, non_reference_picture_flag, colour_plane_id, slice_pic_order_cnt_lsb, recovery_poc_cnt, no_output_of_prior_pics_flag, pic_output_flag, and / or slice_temporal_mvp_enabled_flag. The corresponding values ​​of these syntax elements may be the same in multiple (e.g., all) slice headers associated with the encoded image. (For example, slice_pic_parameter_set_id, non_reference_picture_flag, colour_plane_id, slice_pic_order_cnt_lsb, recovery_poc_cnt, no_output_of_prior_pics_flag, pic_output_flag, slice_temporal_mvp_enabled_flag, etc.) can be signaled to one or more of these syntax elements in the picture header (e.g., instead of in multiple slice headers) or in the layer access unit delimiter (layer AUD). Doing so can reduce the cost associated with slicing overhead. The layer AUD can correspond to the NAL unit type and can be used to indicate the boundaries of the layer-encoded picture.

[0077] Access units can include pictures from different layers. One or more (e.g., all) pictures within an access unit can share the same output time instance and / or the same picture order count (POC) value. One or more syntax elements (e.g., slice_pic_order_cnt_lsb) can be signaled in the AUD (e.g., instead of in the slice header). When the one or more syntax elements (e.g., slice_pic_order_cnt_lsb) are signaled in the AUD, a dependency between the AUD and one or more slices can be introduced. Table 3 below shows exemplary syntax elements that can be included in the AUD. As shown, the syntax can include the aud_pic_order_cnt_isb element, which can be otherwise included in multiple slice headers (e.g., as slice_pic_order_cnt_lsb).

[0078] Table 3 - Exemplary Syntax Elements Placed in AUD .

[0079] Non-reference image characteristics can be signaled in the slice header, for example, to indicate sublayer reference and / or non-reference characteristics of the image. A video processing apparatus (e.g., a decoder) receiving the signaling information can determine, based on the signaling information, that one or more images can be discarded under certain circumstances (e.g., when playback is lagging). In the example (e.g., when using a multi-layer coding structure), a non-reference layer not referenced by other layers can be indicated at the Video Parameter Set (VPS) or SPS level, allowing one or more images of the non-reference layer to be discarded. Table 4 below shows an exemplary VPS syntax that includes elements (e.g., vps_non_reference_layer_flag) indicating that a layer (e.g., a non-reference layer) may not be referenced by other layers (e.g., it may not be a direct reference layer for those other layers).

[0080] Table 4 - Exemplary VPS Syntax, indicating that non-referenced layers are not referenced by other layers. .

[0081] In the exemplary syntax shown in Table 4, setting the value of the parameter vps_non_reference_layer_flag[i] to 1 (or another suitable value) indicates that layer i may not be used as a reference layer for inter-layer prediction (e.g., through another layer, such as layer j). Conversely, setting the value of vps_non_reference_layer_flag[i] to 0 (or another suitable value) indicates that layer i may or may not be used as a reference layer for inter-layer prediction (e.g., through another layer, such as layer j).

[0082] The video processing apparatus described herein can be configured to send or receive syntax elements (e.g., via a video bitstream) indicating whether wraparound motion compensation is enabled or disabled for a sub-picture. Table 5 below shows example SPS syntax where elements associated with wraparound motion compensation (e.g., sps_ref_wraparound_enabled_flag) can be signaled (e.g., after one or more elements associated with sub-picture segmentation).

[0083] Table 5 - Exemplary SPS Syntax Including Wrap Indicators .

[0084] In the example (e.g., in a view-dependent flow), wrap-around motion compensation may or may not be applied to specific sub-images, such as sub-images within a merged image. Figure 7 illustrates an example of wrap-around motion compensation for the original image and the merged image. The original 360-degree image may be encoded as a set of one or more high-resolution images (e.g., represented by 1-6 out of 702) and / or a set of one or more low-resolution images (e.g., represented by 1-6 out of 704). Wrap-around motion compensation may be applied to these images (e.g., the high-resolution image of 702 and the low-resolution image of 704) with corresponding (e.g., different) wrap-around offsets. A new image 706 including two high-resolution sub-images (e.g., represented by 1 and 3 in Figure 7) and four low-resolution sub-images (e.g., represented by 4, 5, 2, and 6 in Figure 7) may be derived, for example, via extraction and / or merging. Sub-images 1 and 3 may be grouped into a first sub-image, and sub-images 4, 5, 2, and 6 may be grouped into a second sub-image. In cases like this, it might be desirable to apply different surround motion compensations (e.g., different surround offsets) to the first and second sub-images.

[0085] The video processing apparatus described herein can be configured to send or receive instructions on whether to enable or disable wrap-around motion compensation for a sub-picture (e.g., via a video bitstream) (e.g., for each sub-picture) and / or instructions on wrap-around offsets to be applied (e.g., via a video bitstream) (e.g., when wrap-around motion compensation is enabled). For example, the video encoding apparatus described herein can be configured to encode a picture including a first sub-picture and / or a second sub-picture. The video encoding apparatus can obtain information indicating whether wrap-around motion compensation is enabled for the first sub-picture and / or the second sub-picture and the corresponding wrap-around offsets associated with the first and second sub-pictures of the encoded picture. The video encoding apparatus can then form a set of encoded data including the encoded picture and the obtained information. In an example, the obtained information included in the set of encoded data can indicate that wrap-around motion compensation is enabled for the first sub-picture and disabled for the second sub-picture. In an example, the obtained information included in the set of encoded data can include a Picture Parameter Set (PPS) syntax element indicating that wrap-around motion compensation is enabled and a Sequence Parameter Set (SPS) syntax element indicating that the first sub-picture will be treated as a picture. In the example, the video encoding device can be configured to send a set of encoded data to a receiving device such as a video decoding device.

[0086] For example, a subpico-level wraparound indicator as described herein can be provided when an indicator for treating a subpico as a picture (e.g., `subpic_treated_as_pic_flag`) is set to true or 1, and when an SPS indicator regarding wraparound motion compensation (e.g., `sps_ref_wraparound_enabled_flag`) is also set to true or 1. Table 6 below shows example PPS syntax structures for signaling the number of subpicos (e.g., maximum number), wraparound motion compensation enabled / disabled indicators (e.g., for subpicos), wraparound motion compensation offsets (e.g., for subpicos), etc. Although signaling occurs in the PPS, one or more of the syntax elements in Table 6 can also be signaled in the SPS.

[0087] Table 6 - Exemplary PPS syntax, including wrapping tags and offsets .

[0088] The exemplary syntax shown in Table 6 may include elements such as `max_subpics_minus2 plus 2`, which specify the maximum number of subpicks that can exist in a Code Video Sequence (CVS). The value of this element can range from 0 to 254 (e.g., 255 can be reserved for future use). The exemplary syntax may include elements such as `all_subpic_wraparound_enabled_flag`, which indicate whether wraparound motion compensation is enabled (e.g., for all subpicks, when the element has a value of 1) or disabled / skipped (e.g., not applying wraparound to all subpicks, when the element has a value of zero) for one or more subpicks. When this element (e.g., `all_subpic_wraparound_enabled_flag`) is absent, its value can be inferred to be equal to 0.

[0089] The exemplary syntax shown in Table 6 may include elements such as subpic_wraparound_offset_sps_flag, which indicates whether to infer that the subpicture wraparound motion compensation offset is equal to the value sps_ref_wraparound_offset_minus1 plus 1 (e.g., when subpic_wraparound_offset_sps_flag is set to 1) or is specified by another element, such as subpic_wraparound_offset_minus1 (e.g., when subpic_wraparound_offset_sps_flag is set to zero).

[0090] The exemplary syntax shown in Table 6 may include elements such as `subpic_ref_wraparound_enabled_flag[i]`, which indicates whether wraparound motion compensation (e.g., horizontal wraparound motion compensation) (e.g., for inter-frame prediction of the i-th subpic) is enabled or disabled for the i-th subpic. When this element is set to 1, it indicates that horizontal wraparound motion compensation is enabled (e.g., applied) for the i-th subpic. When this element is set to zero, it indicates that horizontal wraparound motion compensation is disabled (e.g., not applied) for the i-th subpic. When this element is not present in the signaling syntax, its value can be inferred to be equal to the value of the `all_subpic_wraparound_enabled_flag` element described herein.

[0091] The exemplary syntax shown in Table 6 may include an element, such as `subpic_ref_wraparound_offset_minus1 plus 1`, which specifies the offset associated with wraparound motion compensation (e.g., for calculating the horizontal wraparound position of the i-th subpic). The offset value can be specified in units of `MinCbSizeY` luminance samples. For example, the value of `subpic_ref_wraparound_offset_minus1` can be set in the range of (CtbSizeY / MinCbSizeY)+1 to (subpic_width_in_luma_samples[i] / MinCbSizeY)−1 (inclusive), where `subpic_width_in_luma_sample[i]` can represent the width of the i-th subpic in the luminance samples. When this element (e.g., `subpic_ref_wraparound_offset_minus1`) is absent, its value can be inferred to be equal to the value of the `sps_ref_wraparound_offset_minus1` element described herein.

[0092] The video processing apparatus described herein can apply surround motion compensation to sub-images (e.g., sub-image boundaries) based on the syntax elements described herein. For example, when performing luminance sample bilinear interpolation (e.g., when determining the luminance position in a full sample cell (xInt)). i yInt i The video processing device may consider image wraparound motion compensation indicators (e.g., subpic_ref_wraparound_enabled_flag) for sub-i = 0..1, as well as other syntax elements (e.g., subpic_treated_as_pic_flag), as shown below.

[0093] If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies: xInt i = Clip3( SubPicLeftBoundaryPos, SubPicRightBoundaryPos, subPic_ref_wraparound_enabled_flag ClipH( ( subpic_ref_wraparound_offset_minus1 + 1 ) *MinCbSizeY, SubPicWidth, ( xInt L + i ) ) : xInt L + i )yInt i = Clip3( SubPicTopBoundaryPos, SubPicBotBoundaryPos, yInt L + i).

[0094] When performing luminance sample interpolation filtering (e.g., when determining the luminance position in a full sample unit (xInt)). i yInt i The video processing device may consider subpicture wraparound motion compensation indicators (e.g., subpic_ref_wraparound_enabLED_flag) for i = 0..7, as well as other syntax elements (e.g., subpic_treated_ed_as_pic_flag), as shown below.

[0095] If subpic_treated_as_pic_flag[SubPicIdx] equals 1, then the following applies: xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, subPic_ref_wraparound_enabled_flag) ClipH((subpic_ref_wraparound_offset_minus1 + 1) * MinCbSizeY,SubPicWidth, xIntL + i − 3) : xIntL + i − 3)yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i −3).

[0096] When performing chromaticity sample interpolation filtering (e.g., when determining the chromaticity position in a full sample unit (xInt)). i yInt i The video processing device may consider subpicture wraparound motion compensation indicators (e.g., subpic_ref_wraparound_enabled_flag) for i = 0..3, as well as other syntax elements (e.g., subpic_treated_as_pic_flag), as shown below.

[0097] If `subpic_treated_as_pic_flag[SubPicIdx]` equals 1, then the following applies: `xInti = Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC, subpic_ref_wraparound_enabled_flag)` ClipH(xOffset, SubPicWidthC, xIntC + i) : xIntL + i)yIntti = Clip3(SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yIntL + i) In the example (e.g., when the sub-picture boundary is not aligned with the corresponding picture boundary), the hardware (HW) decoder may not perform wrap-around filling. Figure 8 shows an example of sub-picture wrap-around filling. In the circled area, wrap-around filling can be performed inside the picture. In some implementations (e.g., when using multiple decoders), each sub-picture boundary can be the same as the corresponding picture boundary. In some implementations (e.g., when using a single decoder), the picture boundary may not be applied to the sub-picture boundaries inside the picture.

[0098] Horizontal motion prediction around images can improve the encoding efficiency of certain types of content, such as 360-degree video that can be delivered via view-dependent streaming. For example, if sub-picture level motion prediction around images is not supported, it can be disabled for sub-picture view-dependent streaming. Conversely, if it is supported, it can be enabled for sub-picture view-dependent streaming.

[0099] Sub-image level wrap-around motion prediction can be enabled if the sub-image boundaries are aligned with the image boundaries. Figure 9 shows when sub-images can be wrap-around filled and when they may not. As shown, the top two sub-images on the left side of Figure 9 can be wrap-around filled because their left and right boundaries are aligned with the boundaries of the constituent image on the right side of Figure 9. In contrast, the bottom two sub-images on the left side of Figure 9 may not be wrap-around filled because one or more vertical boundaries of those sub-images (e.g., the left and right vertical boundaries) are not aligned with the boundaries of the constituent image. Table 7 below shows exemplary syntax associated with wrap-around motion prediction that can support the examples in Figure 9.

[0100] Table 7 - Exemplary syntax elements associated with wrapping padding .

[0101] As shown in the figure, the exemplary syntax in Table 7 may include elements such as the indicator `sps_subpic_wraparound_enabled_flag`. When set to 1 or true, this element can indicate the presence of one or more subpic wraparound syntax elements in SPS, such as `sps_subpic_wraparound_boundaries_pos_y0[i]`, `sps_subpic_wraparound_boundaries_pos_y1[i]`, and / or `sps_subpic_wraparound_offset_minus1[i]`. When the indicator element `sps_subpic_wraparound_enabled_flag` is set to false or zero, it can indicate the presence of a picture wraparound offset indicator in SPS, such as `sps_ref_wraparound_offset_minus1`.

[0102] The exemplary syntax in Table 7 may include the element `num_wraparound_boundaries_minus1`, which specifies the total number of boundary segments for which wraparound filling can be performed. In the exemplary syntax, one or more elements (e.g., a combination of `sps_subpic_wraparound_boundaries_pos_y0[i]` and `sps_subpic_wraparound_boundaries_pos_y1[i]`) may specify the location of the i-th boundary segment (e.g., in brightness samples or CTs). The exemplary syntax may also include the element `sps_subpic_wraparound_offset_minus1[i]`, which specifies the offset value to be applied to the i-th boundary segment.

[0103] The video processing apparatus described herein can be configured to send or receive (e.g., via a video bitstream) sub-picture locations, sub-picture sizes, and / or sub-picture IDs. Sub-picture locations, sizes, and / or IDs can be signaled based on a sub-picture grid, which may have a 4 × 4 size (e.g., the smallest grid element size). The bit count associated with signaling can depend on the number of sub-pictures in the sub-picture grid. For example, for a 4Kx2K picture, the signaling bit count could be 47 bits for 6 sub-pictures, 149 bits for 24 sub-pictures, and 701 bits for 96 sub-pictures. For example, the bit count can be reduced without explicitly signaling the ID of each sub-picture when the sub-pictures in the sub-picture grid share the same size (e.g., in a cube map projection (CMP)) and the sub-picture ID is derived from the sub-picture grid.

[0104] Figure 10 illustrates three exemplary subpicture grids. The first grid contains 6 subpictures, the second grid contains 24 subpictures, and the third grid contains 10 subpictures. Each grid in the first and second grids may contain subpictures of the same size, and the third grid may contain 10 subpictures of different sizes (e.g., indicated by the different shading shown in Figure 10). Syntax elements such as single_subpic_per_grid_flag (e.g., SPS syntax elements) can be used to adjust the signaling of subsubpic_grid_idx[i][j] (e.g., to indicate whether to skip the signaling of subpic_grid_idx[i][j]), as shown in Table 8 below. Setting the value of `subpic_per_grid_flag` to 1 (or another suitable value) indicates that `subpic_grid_idx[i][j]` does not exist in the SPS RBSP syntax, and setting the value of `single_subpic_per_grid_flag` to 0 (or another suitable value) indicates that `subpic_grid_idx[i][j]` exists in the SPS RBSP syntax. When the element `single_subpic_per_grid_flag` does not exist, its value (e.g., the value of `single_subpic_per_grid_flag`) can be inferred to be equal to 1 (or another suitable value indicating that `subpic_grid_idx[i][j]` does not exist in the SPS RBSP syntax).

[0105] Table 8 - Exemplary SPS Syntax Including Signaling Indicators .

[0106] Using the exemplary syntax shown in Table 8, when single_subpic_per_grid_flag is equal to 1 (or another suitable value indicating that subpic_grid_idx[ i ][ j ] is not signaled), subpic_grid_idx[ i ][ j ] can be derived as follows: for(i = 0; i<NumSubPicGridRows; i++) for(j = 0; j<NumSubPicGridCols; j++) subpic_grid_idx[ i ][ j ]= i * NumSubPicGridRows +NumSubPicGridColsThe value of subpic_grid_idx can be in the range from 0 to max_subpics_minus1, including the end values. By adjusting the subpicture signaling in the manner described herein, the bit count associated with the signaling can be reduced to 30 bits for, e.g., 6, 24, and 96 subpictures.

[0107] Syntax elements such as subpic_grid idx as described herein can be used to signal the subpicture ID. The minimum subpicture grid size can be 4×4. A subpicture can correspond to a rectangular region of one or more slices within a picture, and a slice can include a sequence of multiple complete tiles or a complete brick of a tile (e.g., a consecutive sequence). The slice position and / or size can be signaled, e.g., in the PPS. The slice ID can be used to indicate the subpicture position and / or subpicture size to improve signaling efficiency.

[0108] The slice ID and / or slice_address can be signaled in the slice header. slice_address can be equal to the set of slice IDs in the PPS (e.g., set explicitly) or the slice index of a rectangular slice or the brick ID of a raster-scanned slice. In an example (e.g., for a rectangular slice), the slice position can be determined by syntax elements (e.g., parameters) such as bottom_right_brick_idx (e.g., bottom_right_brick_idx can be used to derive TopLeftBrickIdx and / or BottomRightBrickIdx). Raster-scanned slices can be used at least in low-latency scenarios. Subpictures can be restricted to include one or more rectangular slices (e.g., only rectangular slices), and such constrained subpictures can be signaled based on the slices included in the subpicture to facilitate slice extraction. Table 9 below shows an exemplary syntax for subpicture signaling.

[0109] Table 9 - Exemplary Sub-image Syntax 。

[0110] As shown in the table, the exemplary subpicture signaling syntax may include a first element, `subpic_per_slice_flag`, which can be used to indicate, for example, whether each slice is a subpicture (e.g., when `subpic_per_slice_flag` is set to 1) or whether a slice comprises one or more slices (e.g., when `subpic_per_slice_flag` is set to 0). The exemplary syntax may include a second element, `num_slices_minus1[i] plus 1`, specifying the total number of slices within the i-th subpicture, and a third element, `slice_address[i][j]`, specifying the address of the j-th slice of the i-th subpicture. In the example (e.g., when elements such as `signaled_slice_id_flag` are signaled in the PPS and set to 1), the value of `slice_address[i][j]` can be equal to the slice ID of the slice, and the value of `slice_address[i][j]` can range from 0 to 2. (signalled_slice_id_length_minus1 + 1) -1, including end values. In the example (e.g., when signaled_slice_id_flag is set to 0), the value of slice_address[i][j] can be in the range of 0 to nnum_slices_in_pic_minus1, including end values.

[0111] Syntax elements such as `recovery_poc_cnt` can be signaled in the slice header of a Gradual Decode Refresh (GDR) NAL unit. These elements can specify the recovery point of the decoded image in the output order. An image can be called a recovery point image when it follows the current GDR image in decoding order and its POC value is equal to the GDR's POC value plus the value of `recovery_poc_cnt`. Figure 11 shows an example of a layer-based coding structure. A layer image can refer to a GDR image from its dependent layers. Constraints can be imposed requiring that the NAL unit type (NUT) of the layer image for which the inter-layer reference image is a GDR image is set to GDR_NUT (e.g., the NAL unit type for GDR images) and / or that the value of `recovery_poc_cnt` in the current slice is the same as the value of `recovery_poc_cnt` in the corresponding inter-layer reference image.

[0112] Figure 12A is a schematic diagram illustrating an exemplary communication system 1200 that can be implemented in one or more of the disclosed embodiments. Communication system 1200 can be a multiple access system providing content such as voice, data, video, messaging, and broadcasting to multiple wireless users. Communication system 1200 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, communication system 1200 can employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.

[0113] As shown in Figure 12A, the communication system 1200 may include wireless transmit / receive units (WTRUs) 1202a, 1202b, 1202c, 1202d, RAN 1204 / 1213, CN 1206 / 1215, public switched telephone network (PSTN) 1208, Internet 1210, and other networks 1212. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 1202a, 1202b, 1202c, and 1202d can be any type of device configured to operate and / or communicate in a wireless environment. As an example, WTRUs 1202a, 1202b, 1202c, and 1202d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 1202a, 1202b, 1202c, and 1202d may be interchangeably referred to as a UE.

[0114] The communication system 1200 may also include base stations 1214a and / or 1214b. Each of the base stations 1214a and 1214b may be any type of device configured to wirelessly interface with at least one of the WTRUs 1202a, 1202b, 1202c, and 1202d to facilitate access to one or more communication networks, such as CN 1206 / 1215, Internet 1210, and / or other networks 1212. As an example, base stations 1214a and 1214b may be base transceiver stations (BTS), Node Bs, evolved Node Bs, home Node Bs, home evolved Node Bs, gNBs, NR Node Bs, site controllers, access points (APs), wireless routers, etc. Although base stations 1214a and 1214b are each depicted as a single element, it should be understood that base stations 1214a and 1214b may include any number of interconnected base stations and / or network elements.

[0115] Base station 1214a may be part of RAN 1204 / 1213, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 1214a and / or base station 1214b may be configured to transmit and / or receive radio signals on one or more carrier frequencies (which may be referred to as cells (not shown)). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage of radio services to a specific geographic area, which may be relatively fixed or changeable over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 1214a may be divided into three sectors. Therefore, in one embodiment, base station 1214a may include three transceivers, i.e., one transceiver per sector of the cell. In one embodiment, base station 1214a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in the desired spatial direction.

[0116] Base stations 1214a and 1214b can communicate with one or more of WTRUs 1202a, 1202b, 1202c, and 1202d via air interface 1216, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish air interface 1216.

[0117] More specifically, as noted above, the communication system 1200 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 1214a and WTRUs 1202a, 1202b, and 1202c in RAN 1204 / 1213 can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which can use Wideband CDMA (WCDMA) to establish air interfaces 1215 / 1216 / 1217. WCDMA may include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or evolved HSPA (HSPA+). HSPA may include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSDPA).

[0118] In one implementation, base station 1214a and WTRUs 1202a, 1202b, 1202c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-A Pro) to establish air interface 1216.

[0119] In one implementation, base station 1214a and WTRUs 1202a, 1202b, 1202c can implement radio technologies such as NR radio access, which can use New Radio (NR) to establish air interface 1216.

[0120] In one implementation, base station 1214a and WTRUs 1202a, 1202b, and 1202c can implement multiple radio access technologies. For example, base station 1214a and WTRUs 1202a, 1202b, and 1202c can, for example, use a dual connectivity (DC) principle to implement both LTE and NR radio access together. Therefore, the air interface used by WTRUs 1202a, 1202b, and 1202c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).

[0121] In other implementations, base station 1214a and WTRUs 1202a, 1202b, and 1202c can implement radio technologies such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), GSM Enhanced Data Rate Evolution (EDGE), and GSMEDGE (GERAN).

[0122] Base station 1214b in Figure 12A can be, for example, a wireless router, a home node B, a home evolution node B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in local areas such as commercial locations, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for use by drones), roads, etc. In one embodiment, base station 1214b and WTRUs 1202c, 1202d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, base station 1214b and WTRUs 1202c, 1202d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 1214b and WTRUs 1202c, 1202d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. As shown in Figure 12A, base station 1214b may have a direct connection to Internet 1210. Therefore, base station 1214b may not need to access Internet 1210 via CN 1206 / 1215.

[0123] RAN 1204 / 1213 can communicate with CN 1206 / 1215, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU 1202a, 1202b, 1202c, and 1202d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 1206 / 1215 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, etc., and / or perform advanced security functions such as user authentication. Although not shown in Figure 12A, it should be understood that RAN 1204 / 1213 and / or CN 1206 / 1215 can communicate directly or indirectly with other RANs using the same RAT as RAN 1204 / 1213 or a different RAT. For example, in addition to connecting to RAN 1204 / 1213 which utilizes NR radio technology, CN1206 / 1215 can also communicate with another RAN (not shown) that uses GSM, UMTS, CDMA 2000, WiMAX, E-UTRA or WiFi radio technology.

[0124] CN 1206 / 1215 may also act as a gateway for WTRU 1202a, 1202b, 1202c, 1202d to access PSTN 1208, Internet 1210, and / or other networks 1212. PSTN 1208 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). Internet 1210 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 1212 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 1212 may include another CN connected to one or more RANs, which may use the same RAT as RAN 1204 / 1213 or a different RAT.

[0125] Some or all of the WTRUs 1202a, 1202b, 1202c, and 1202d in the communication system 100 may include multi-mode capabilities (e.g., WTRUs 1202a, 1202b, 1202c, and 1202d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, the WTRU 1202c shown in Figure 12A may be configured to communicate with a base station 1214a that may employ cellular-based radio technology and with a base station 1214b that may employ IEEE 802 radio technology.

[0126] Figure 12B is a system diagram illustrating an exemplary WTRU 1202. As shown in Figure 12B, the WTRU 1202 may include a processor 1218, a transceiver 1220, a transmit / receive element 1222, a speaker / microphone 1224, a keypad 1226, a display / touchpad 1228, non-removable memory 1230, removable memory 1232, a power supply 1234, a Global Positioning System (GPS) chipset 1236, and / or other peripheral devices 1238, etc. It should be understood that the WTRU 1202 may include any sub-combination of the foregoing elements while remaining consistent with the implementation.

[0127] Processor 1218 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 1218 may perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 1202 to operate in a wireless environment. Processor 1218 may be coupled to transceiver 1220, which may be coupled to transmitting / receiving element 1222. Although Figure 12B depicts processor 1218 and transceiver 1220 as separate components, it should be understood that processor 1218 and transceiver 1220 may be integrated together in an electronic package or on a chip.

[0128] Transmitting / receiving element 1222 can be configured to transmit signals to or receive signals from a base station (e.g., base station 1214a) via air interface 1216. For example, in one embodiment, transmitting / receiving element 1222 can be an antenna configured to transmit and / or receive RF signals. In one embodiment, transmitting / receiving element 1222 can be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 1222 can be configured to transmit and / or receive RF and optical signals. It should be understood that transmitting / receiving element 1222 can be configured to transmit and / or receive any combination of wireless signals.

[0129] Although the transmitting / receiving element 1222 is depicted as a single element in Figure 12B, the WTRU 1202 may include any number of transmitting / receiving elements 1222. More specifically, the WTRU 1202 may employ MIMO technology. Therefore, in one embodiment, the WTRU 1202 may include two or more transmitting / receiving elements 1222 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 1216.

[0130] Transceiver 1220 can be configured to modulate signals transmitted by transmitting / receiving element 1222 and demodulate signals received by transmitting / receiving element 1222. As noted above, WTRU 1202 may have multi-mode capability. Therefore, transceiver 1220 may include multiple transceivers to enable WTRU 1202 to communicate via various RATs such as NR and IEEE 802.11.

[0131] The processor 1218 of WTRU 1202 can be coupled to a speaker / microphone 1224, a keypad 1226, and / or a display / touchpad 1228 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit) and can receive user input data therefrom. The processor 1218 can also output user data to the speaker / microphone 1224, the keypad 1226, and / or the display / touchpad 1228. Furthermore, the processor 1218 can access information from any type of suitable memory (such as non-removable memory 1230 and / or removable memory 1232) and store data in any type of suitable memory. Non-removable memory 1230 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 1232 may include a Subscriber Identity Module (SIM) card, a Memory Stick, a Secure Digital (SD) memory card, etc. In other embodiments, the processor 1218 may access memory information that is never physically located on the WTRU 1202 (such as on a server or home computer (not shown)) and store the data in that memory.

[0132] The processor 1218 may receive power from the power supply 1234 and may be configured to distribute and / or control power to other components in the WTRU 1202. The power supply 1234 may be any suitable device for powering the WTRU 1202. For example, the power supply 1234 may include one or more dry cell battery packs (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0133] The processor 1218 may also be coupled to a GPS chipset 1236, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 1202. In addition to or instead of the information from the GPS chipset 1236, the WTRU 1202 may receive location information from base stations (e.g., base stations 1214a, 1214b) via air interface 1216 and / or determine its location based on the time of signal reception from two or more nearby base stations. It should be understood that, while remaining consistent with the implementation, the WTRU 1202 may acquire location information using any suitable location determination method.

[0134] The processor 1218 may also be coupled to other peripheral devices 1238, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripheral device 1238 may include an accelerometer, electronic compass, satellite transceiver, digital camera (for photos and / or video), Universal Serial Bus (USB) port, vibration device, television transceiver, hands-free headset, Bluetooth. ® Modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 1238 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.

[0135] WTRU 1202 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference through signal processing via hardware (e.g., a choke) or via a processor (e.g., a separate processor (not shown) or via processor 1218). In one embodiment, WTRU 1202 may include a full-duplex radio for which the transmission and reception of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception) may be concurrent and / or simultaneous.

[0136] Figure 12C is a system diagram illustrating RAN 1204 and CN 1206 according to one embodiment. As described above, RAN 1204 can communicate with WTRUs 1202a, 1202b, and 1202c via air interface 1216 using E-UTRA radio technology. RAN 1204 can also communicate with CN 1206.

[0137] RAN 1204 may include evolved Node Bs 1260a, 1260b, and 1260c; however, it should be understood that RAN 1204 may include any number of evolved Node Bs while remaining consistent with the implementation scheme. Evolved Node Bs 1260a, 1260b, and 1260c may each include one or more transceivers to communicate with WTRUs 1202a, 1202b, and 1202c via air interface 1216. In one implementation, evolved Node Bs 1260a, 1260b, and 1260c may implement MIMO technology. Therefore, evolved Node B 1260a may, for example, use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 1202a.

[0138] Each of the evolved Nodes B 1260a, 1260b, and 1260c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in the UL and / or DL, etc. As shown in Figure 12C, the evolved Nodes B 1260a, 1260b, and 1260c can communicate with each other via the X2 interface.

[0139] The CN 1206 shown in Figure 12C may include a Mobility Management Entity (MME) 1262, a Serving Gateway (SGW) 1264, and a Packet Data Network (PDN) Gateway (or PGW) 1266. Although each of the foregoing elements is depicted as part of the CN 1206, it should be understood that any of these elements may be owned and / or operated by an entity other than the CN operator.

[0140] The MME 1262 can connect to each of the evolved nodes B 1260a, 1260b, and 1260c in RAN 1204 via the S1 interface and can be used as a control node. For example, the MME 1262 can be responsible for authenticating users of WTRUs 1202a, 1202b, and 1202c, bearer activation / deactivation, selecting a specific serving gateway during the initial attachment of WTRUs 1202a, 1202b, and 1202c, etc. The MME 1262 can provide control plane functions for handover between RAN 1204 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.

[0141] The SGW 1264 can connect to each of the evolved Node Bs 1260a, 1260b, and 1260c in RAN 1204 via the S1 interface. The SGW 1264 typically routes and forwards user data packets to / from WTRUs 1202a, 1202b, and 1202c. The SGW 1264 can perform other functions such as anchoring the user plane during inter-evolved Node B handovers, triggering paging when DL data is available for WTRUs 1202a, 1202b, and 1202c, and managing and storing the context of WTRUs 1202a, 1202b, and 1202c.

[0142] The SGW 1264 can be connected to the PGW 1266, which provides WTRUs 1202a, 1202b, and 1202c with access to packet-switched networks (such as Internet 1210) to facilitate communication between WTRUs 1202a, 1202b, 1202c and IP-enabled devices.

[0143] CN 1206 can facilitate communication with other networks. For example, CN 1206 can provide WTRU 1202a, 1202b, and 1202c with access to a circuit-switched network (such as PSTN 1208) to facilitate communication between WTRU 1202a, 1202b, and 1202c and conventional landline communication equipment. For example, CN 1206 may include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN 1206 and PSTN 1208, or be able to communicate with such an IP gateway. Additionally, CN 1206 can provide WTRU 1202a, 1202b, and 1202c with access to other networks 1212, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0144] Although the WTRU is described as a wireless terminal in Figures 12A to 12D, it is conceivable that in some representative embodiments, such a terminal may (e.g., temporarily or permanently) use a wired communication interface with a communication network.

[0145] In a representative implementation, the other network 1212 may be a WLAN.

[0146] A WLAN in Infrastructure Basic Services Set (BSS) mode may have an access point (AP) for the BSS and one or more sites (STAs) associated with the AP. The AP may have access or an interface to a distribution system (DS) or another type of wired / wireless network that carries traffic to and / or carries traffic out of the BSS. Traffic originating outside the BSS and destined for a STA can reach and be delivered to the STA via the AP. Traffic originating from a STA and destined for a destination outside the BSS can be sent to the AP for delivery to the appropriate destination. Traffic between STAs within the BSS can be sent via the AP, for example, where a source STA can send traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as point-to-point traffic. Point-to-point traffic can be sent between source and destination STAs (e.g., directly between them) using Direct Link Establishment (DLS). In some representative implementations, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). WLANs using Standalone BSS (IBSS) mode may not have access points (APs), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. IBSS communication mode may sometimes be referred to as "ad-hoc" communication mode in this document.

[0147] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel can be of fixed width (e.g., a 20 MHz wide bandwidth) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative implementations, Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) can be implemented, for example, in an 802.11 system. For CSMA / CA, each STA (including the AP) can sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, that STA can back off. A single STA (e.g., only one station) can transmit in a given BSS at any given time.

[0148] High-throughput (HT) STAs can communicate using a 40MHz wide channel, for example, by combining a primary 20MHz channel with adjacent or non-adjacent 20MHz channels to form a 40MHz wide channel.

[0149] Very High Throughput (VHT) STAs support channels with widths of 20MHz, 40MHz, 80MHz, and / or 160MHz. 40MHz and / or 80MHz channels can be formed by combining consecutive 20MHz channels. A 160MHz channel can be formed by combining eight consecutive 20MHz channels, or by combining two non-consecutive 80MHz channels (this can be referred to as an 80+80 configuration). For the 80+80 configuration, after channel coding, data can be split into two streams by a segment parser. Each stream can be processed individually using Inverse Fast Fourier Transform (IFFT) and time-domain processing. These streams can be mapped to two 80MHz channels, and data can be transmitted via a transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration can be reversed, and the combined data can be sent to Media Access Control (MAC).

[0150] 802.11af and 802.11ah support operating modes below 1 GHz. Compared to those used in 802.11n and 802.11ac, 802.11af and 802.11ah reduce channel operating bandwidth and carrier. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative implementations, 802.11ah may support instrument-type control / machine-type communications, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support (e.g., only support) certain bandwidths and / or limited bandwidths. MTC devices may include batteries with battery life above a threshold (e.g., to maintain a very long battery life).

[0151] WLAN systems supporting multiple channels, as well as channel bandwidths such as 802.11n, 802.11ac, 802.11af, and 802.11ah, include channels that can be designated as primary channels. A primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by STAs operating in the BSS (each supporting a minimum bandwidth operating mode). In the 802.11ah example, for STAs supporting (e.g., only supporting) a 1MHz mode (e.g., MTC-type devices), the primary channel can be 1MHz wide, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes. Carrier Sense and / or Network Allocation Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy, for example, because an STA (supporting only the 1MHz operating mode) is transmitting to the AP, the entire available band can be considered busy even if most of the band remains idle and potentially available.

[0152] In the United States, the available frequency bands for 802.11ah are 902MHz to 928MHz. In South Korea, the available frequency bands are 917.5MHz to 923.5MHz. In Japan, the available frequency bands are 916.5MHz to 927.5MHz. The total available bandwidth for 802.11ah ranges from 6MHz to 26MHz, depending on the country code.

[0153] Figure 12D is a system diagram illustrating RAN 1213 and CN 1215 according to one embodiment. As noted above, RAN 1213 can communicate with WTRUs 1202a, 1202b, and 1202c via air interface 1216 using NR radio technology. RAN 1213 can also communicate with CN 1215.

[0154] RAN 1213 may include gNBs 1280a, 1280b, and 1280c; however, it should be understood that RAN 1213 may include any number of gNBs, while remaining consistent with the implementation scheme. gNBs 1280a, 1280b, and 1280c may each include one or more transceivers for communication with WTRUs 1202a, 1202b, and 1202c via air interface 1216. In one implementation, gNBs 1280a, 1280b, and 1280c may implement MIMO technology. For example, gNBs 1280a and 1280b may utilize beamforming to transmit signals to and / or receive signals from gNBs 1280a, 1280b, and 1280c. Therefore, gNB 1280a can, for example, use multiple antennas to transmit and / or receive radio signals from WTRU 1202a. In one embodiment, gNBs 1280a, 1280b, and 1280c can implement carrier aggregation technology. For example, gNB 1280a can transmit multiple component carriers to WTRU 1202a (not shown). A subset of these component carriers may be on unlicensed spectrum, while the remaining component carriers may be on licensed spectrum. In one embodiment, gNBs 1280a, 1280b, and 1280c can implement Cooperative Multipoint (CoMP) technology. For example, WTRU 1202a can receive cooperative transmissions from gNB 1280a and gNB 1280b (and / or gNB 1280c).

[0155] WTRU 1202a, 1202b, and 1202c can communicate with gNB 1280a, 1280b, and 1280c using transmissions associated with scalable parameter sets. For example, OFDM symbol spacing and / or OFDM subcarrier spacing can vary depending on different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRU 1202a, 1202b, and 1202c can communicate with gNB 1280a, 1280b, and 1280c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing different numbers of OFDM symbols and / or continuously varying absolute time lengths).

[0156] gNBs 1280a, 1280b, and 1280c can be configured to communicate with WTRUs 1202a, 1202b, and 1202c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 1202a, 1202b, and 1202c can communicate with gNBs 1280a, 1280b, and 1280c without accessing other RANs (e.g., evolved Node Bs 1260a, 1260b, and 1260c). In standalone configuration, WTRUs 1202a, 1202b, and 1202c can use one or more of gNBs 1280a, 1280b, and 1280c as mobility anchors. In a standalone configuration, WTRUs 1202a, 1202b, and 1202c can communicate with gNBs 1280a, 1280b, and 1280c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 1202a, 1202b, and 1202c can communicate or connect to gNBs 1280a, 1280b, and 1280c, while also communicating or connecting to another RAN (such as Evolved Node B 1260a, 1260b, and 1260c). For example, WTRUs 1202a, 1202b, and 1202c can implement DC principles to communicate substantially simultaneously with one or more gNBs 1280a, 1280b, and 1280c and one or more Evolved Node B 1260a, 1260b, and 1260c. In a non-standalone configuration, evolved nodes B 1260a, 1260b, and 1260c can be used as mobility anchors for WTRU 1202a, 1202b, and 1202c, and gNB 1280a, 1280b, and 1280c can provide additional coverage and / or throughput for serving WTRU 1202a, 1202b, and 1202c.

[0157] Each of the gNBs 1280a, 1280b, and 1280c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in the UL and / or DL, network slicing support, dual connectivity, interoperability between NR and E-UTRA, routing of user plane data to User Plane Functions (UPF) 1284a and 1284b, routing of control plane information to Access and Mobility Management Functions (AMF) 1282a and 1282b, etc. As shown in Figure 12D, the gNBs 1280a, 1280b, and 1280c can communicate with each other via the Xn interface.

[0158] The CN 1215 shown in Figure 12D may include at least one AMF 1282a, 1282b, at least one UPF 1284a, 1284b, at least one Session Management Function (SMF) 1283a, 1283b, and possibly a Data Network (DN) 1285a, 1285b. Although each of the foregoing elements is depicted as part of the CN 1215, it should be understood that any one of these elements may be owned and / or operated by an entity other than the CN operator.

[0159] AMF 1282a and 1282b can connect to one or more of gNB 1280a, 1280b, and 1280c via the N2 interface in RAN 1213 and can be used as control nodes. For example, AMF 1282a and 1282b can be responsible for authenticating users of WTRU 1202a, 1202b, and 1202c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMF 1283a and 1283b, managing registration areas, terminating NAS signaling, mobility management, etc. AMF 1282a and 1282b can use network slicing to customize CN support for WTRU 1202a, 1202b, and 1202c based on the type of service used by WTRU 1202a, 1202b, and 1202c. For example, different network slices can be established for different use cases, such as services that rely on Ultra-Reliable Low Latency (URLLC) access, services that rely on Enhanced Mobile Broadband (eMBB) access, services for Machine Type Communication (MTC) access, etc. The AMF 1282 can provide control plane functions for handover between RAN 1213 and other RANs (not shown) that employ other radio technologies (such as LTE, LTE-A, LTE-A Pro) and / or non-3GPP access technologies (such as WiFi).

[0160] SMFs 1283a and 1283b can connect to AMFs 1282a and 1282b in CN 1215 via the N11 interface. SMFs 1283a and 1283b can also connect to UPFs 1284a and 1284b in CN 1215 via the N4 interface. SMFs 1283a and 1283b can select and control UPFs 1284a and 1284b, and configure traffic routing through UPFs 1284a and 1284b. SMFs 1283a and 1283b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.

[0161] UPF 1284a and 1284b can connect via the N3 interface to one or more of the gNBs 1280a, 1280b, and 1280c in RAN 1213. These gNBs can provide WTRU 1202a, 1202b, and 1202c with access to packet-switched networks (such as Internet 1210) to facilitate communication between WTRU 1202a, 1202b, 1202c and IP-enabled devices. UPF 1284 and 1284b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multihomed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0162] CN 1215 may facilitate communication with other networks. For example, CN 1215 may include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN 1215 and PSTN 1208, or may communicate with such an IP gateway. Additionally, CN 1215 may provide WTRU 1202a, 1202b, and 1202c with access to other networks 1212, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU 1202a, 1202b, and 1202c may be connected to DN 1285a and 1285b via UPF 1284a and 1284b through the N3 interface to UPF 1284a and 1284b and the N6 interface between UPF 1284a and 1284b and local data networks (DNs) 1285a and 1285b.

[0163] Given the corresponding descriptions in Figures 12A to 12D, one or more of the functions described herein with reference to one or more of the following can be performed by one or more emulation devices (not shown): WTRU1202a-d, base station 1214a-b, evolved Node B 1260a-c, MME 1262, SGW 1264, PGW 1266, gNB 1280a-c, AMF 1282a-b, UPF 1284a-b, SMF 1283a-b, DN 1285a-b, and / or any other devices described herein. An emulation device can be one or more devices configured to mimic one or more of the functions described herein. For example, an emulation device can be used to test other devices and / or simulate network and / or WTRU functions.

[0164] Simulation devices can be designed to perform one or more tests on other devices in laboratory and / or carrier network environments. For example, the one or more simulation devices may perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. The one or more simulation devices may perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices may be directly coupled to another device for testing purposes and / or may use over-the-air wireless communication to perform tests.

[0165] The one or more emulation devices may perform one or more (including all) functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be used in test scenarios within a test laboratory and / or non-deployed (e.g., testing) wired and / or wireless communication networks to perform testing of one or more components. The one or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas) may be used by the emulation devices to transmit and / or receive data.

[0166] Although features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with other features and elements. Furthermore, the methods described herein may be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted over wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks and digital versatile optical discs (DVDs)). A processor associated with the software may be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A video decoding device, comprising: A processor configured to: receive video data associated with a current image; extract an indication of an image order count associated with the current image from an image header included in the video data; and decode the current image based at least on the indication of the image order count extracted from the image header.

2. The video decoding apparatus of claim 1, wherein the processor is further configured to extract an indication from the image header as to whether a temporal motion vector predictor (MVP) is enabled for the current image, and further decode the current image based on the indication as to whether a temporal MVP is enabled for the current image.

3. The video decoding device according to claim 1, wherein the processor is further configured to: extract an indication from the image header whether the current image is used as a reference image for at least one other image, and decode at least one other image based on the indication whether the current image is used as a reference image for at least one other image.

4. The video decoding device of claim 3, wherein the processor is further configured to: extract from the video parameter set associated with the current image an indication of whether the first layer associated with the current image is used as a reference layer for the second layer associated with the current image.

5. The video decoding device of claim 4, wherein the processor is further configured to: discard the image associated with the first layer if an indication extracted from the video parameter set indicates that the first layer is not used as a reference layer for the second layer.

6. The video decoding device according to claim 1, wherein the processor is further configured to: extract a flag from the image header and determine whether to output the decoded current image based on the flag.

7. The video decoding device according to claim 1, wherein the processor is further configured to: extract an identifier of an image parameter set associated with the current image from the image header.

8. The video decoding device of claim 1, wherein the indication of the picture order count is not included in any slice header associated with the current picture.

9. The video decoding apparatus of claim 8, wherein the processor is further configured to: determine that an indication of the image order count extracted from the image header applies to a plurality of slices associated with the current image.

10. A method for video decoding, comprising: Receive video data associated with the current image; extract an indication of the image order count associated with the current image from the image header included in the video data; And decode the current image based at least on an indication of the image order count extracted from the image header.

11. The method of claim 10, further comprising: Extract an indication from the image header as to whether the Temporal Motion Vector Predictor (MVP) is enabled for the current image, and further decode the current image based on the indication as to whether the Temporal MVP is enabled for the current image.

12. The method of claim 10, further comprising: Extract an indication from the image header whether the current image is used as a reference image for at least one other image, and decode at least one other image based on the indication that the current image is used as a reference image for at least one other image.

13. The method of claim 12, further comprising: Extract an indication from the video parameter set associated with the current image whether the first layer associated with the current image is used as a reference layer for the second layer associated with the current image.

14. The method of claim 13, further comprising: If the indication extracted from the video parameter set indicates that the first layer is not used as a reference layer for the second layer, the image associated with the first layer is discarded.

15. The method of claim 10, further comprising: Extract the flags from the image header and determine whether to output the decoded current image based on the flags.

16. The method of claim 10, further comprising: Extract the identifier of the set of image parameters associated with the current image from the image header.

17. The method of claim 10, wherein the indication of the image order count is not included in any slice header associated with the current image.

18. The method of claim 17, further comprising: The indicator that determines the image order count extracted from the image header applies to multiple slices associated with the current image.

19. A video encoding device, comprising: The processor is configured to encode the current image into video data; Add an image order count indicator to the image header associated with the current image; And the image order count indicator is encoded into the video data.

20. The video encoding apparatus of claim 19, wherein the processor is further configured to: encode at least one of the following into video data: an indication of whether a temporal motion vector predictor (MVP) is enabled for the current image, an indication of whether the current image is used as a reference image for at least one other image, an indication of whether the current image is output after it is decoded, or an identifier of an image parameter set associated with the current image.

21. The video encoding apparatus of claim 19, wherein the processor is further configured to encode an indication into video data whether a first layer associated with the current image is used as a reference layer for a second layer associated with the current image.

22. A method for video encoding, comprising: Encode the current image into video data; Add an image order count indicator to the image header associated with the current image; And the image order count indicator is encoded into the video data.

23. The method of claim 22, further comprising: Encode at least one of the following into the video data: an indication of whether the Temporal Motion Vector Predictor (MVP) is enabled for the current image, an indication of whether the current image is used as a reference image for at least one other image, an indication of whether the current image is output after it is decoded, or an identifier for a set of image parameters associated with the current image.

24. The method of claim 22, further comprising: The indication of whether the first layer associated with the current image is used as a reference layer for the second layer associated with the current image is encoded into the video data.

25. A video decoding apparatus, comprising: The processor is configured to: determine whether to enable surround motion compensation for an image; based on the determination that surround motion compensation is enabled for the image, obtain parameters for calculating the surround position; based on the obtained parameters, obtain a reference sample position, wherein the reference sample position is associated with a reference sample; and decode the image based on the reference sample.

26. The video decoding apparatus of claim 25, wherein the parameters include an offset associated with the current sample position.

27. A video decoding method, comprising: Determine whether to enable surround motion compensation for the image; Based on the determination that surround motion compensation is enabled for the image, parameters for calculating the surround position are obtained; based on the obtained parameters, a reference sample position is obtained, wherein the reference sample position is associated with a reference sample; and the image is decoded based on the reference sample.

28. The video decoding method of claim 27, wherein the parameters include an offset associated with the current sample position.