Video encoder, video decoder, and corresponding method

By allowing sub-pictures to have non-CTU size dimensions at picture boundaries, the method addresses decoding errors and resource inefficiencies in video coding, achieving improved compression ratios and efficiency in video encoding and decoding.

JP2025111629AActive Publication Date: 2025-07-30HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025071560
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-01-09
Filing Date
2025-04-23
Publication Date
2025-07-30
Estimated Expiration
2040-01-09

AI Technical Summary

Technical Problem

Existing video coding systems face challenges in efficiently compressing video data without sacrificing image quality due to constraints on sub-picture sizes being multiples of the coding tree unit (CTU) size, which can lead to decoding errors and inefficient resource usage.

Method used

The solution involves allowing sub-pictures to have heights and widths that are not multiples of the CTU size, particularly at the right or bottom boundaries of the picture, to accommodate various picture layouts without causing decoding errors, thereby improving encoder and decoder efficiency.

Benefits of technology

This approach enhances the compression ratio by reducing resource usage in encoders and decoders, allowing for more efficient coding and decoding of video data without compromising image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111629000001_ABST
    Figure 2025111629000001_ABST
Patent Text Reader

Abstract

To provide a decoder, an encoder, a method, and a non-temporary computer readable medium which improve a compression ratio with little or no sacrifice in image quality.SOLUTION: A method of decoding a bit stream includes: receiving 1101 a bitstream comprising one or more sub-pictures partitioned from a picture, each sub-picture including a sub-picture that is an integer multiple of a coding tree unit (CTU) size when each sub-picture includes a right border that does not coincide with picture's right border; analyzing the bitstream to obtain 1103 the one or more sub-pictures; decoding the one or more sub-pictures to create a video sequence; and forwarding 1105 the video sequence for display.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 790,207, filed Jan. 9, 2019, by Ye - Kui Wang et al., entitled “Sub - Pictures in Video Coding,” which is incorporated herein by reference.

[0002] This disclosure generally relates to video coding, and more particularly, to sub - picture management in video coding.

Background Art

[0003] Even relatively short videos require a fairly large amount of video data to depict them, which can pose difficulties when the data is to be streamed or otherwise communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. Since memory resources can be limited, the size of the video can also be an issue when the video is stored on a storage device. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Due to limited network resources and the continuing increase in the demand for higher video quality, improving the compression ratio without sacrificing little or no image quality. improvements that improve the compression ratio without sacrificing little or no image quality.​​​​​​​​​​​ Desired compression and decompression techniques are provided. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0004] In one embodiment, the present disclosure includes a method implemented in a decoder, the method comprising: receiving, by a receiver of the decoder, a bitstream comprising one or more sub-pictures segmented from a picture such that when a first sub-picture includes a right boundary that coincides with the right boundary of the picture, the first sub-picture has a sub-picture width that includes an incomplete coding tree unit (CTU); analyzing, by a processor of the decoder, the bitstream to obtain the one or more sub-pictures; decoding, by the processor, the one or more sub-pictures to create a video sequence; and transferring, by the processor, the video sequence for display. Some video systems may limit the sub-pictures to have a height and width that are multiples of the CTU size. However, pictures may include a height and width that are not multiples of the CTU size. Thus, the sub-picture size constraints prevent the sub-pictures from operating correctly with many picture layouts. In the disclosed examples, the sub-picture width and sub-picture height are constrained to be multiples of the CTU size. However, these constraints are removed when the sub-picture is located at the right boundary or the bottom boundary of the picture, respectively. By allowing the bottom and right sub-pictures to include a height and width that are not multiples of the CTU size, any picture can be processed without causing decoding errors. obtaining, by a processor of the decoder, one or more sub-pictures by analyzing the bitstream; decoding, by the processor, the one or more sub-pictures to create a video sequence; and transferring, by the processor, the video sequence for display. Some video systems may limit the sub-pictures to have a height and width that are multiples of the CTU size. However, pictures may include a height and width that are not multiples of the CTU size. Thus, the sub-picture size constraints prevent the sub-pictures from operating correctly with many picture layouts. In the disclosed examples, the sub-picture width and sub-picture height are constrained to be multiples of the CTU size. However, these constraints are removed when the sub-picture is located at the right boundary or the bottom boundary of the picture, respectively. By allowing the bottom and right sub-pictures to include a height and width that are not multiples of the CTU size, any picture can be processed without causing decoding errors. Some video systems may limit the sub-pictures to have a height and width that are multiples of the CTU size. However, pictures may include a height and width that are not multiples of the CTU size. Thus, the sub-picture size constraints prevent the sub-pictures from operating correctly with many picture layouts. In the disclosed examples, the sub-picture width and sub-picture height are constrained to be multiples of the CTU size. However, these constraints are removed when the sub-picture is located at the right boundary or the bottom boundary of the picture, respectively. By allowing the bottom and right sub-pictures to include a height and width that are not multiples of the CTU size, any picture can be processed without causing decoding errors. When the bottom and right sub-pictures include a height and width that are not multiples of the CTU size, respectively, decoding errors can be avoided for any picture. By allowing the bottom and right sub-pictures to include a height and width that are not multiples of the CTU size, respectively, any picture can be processed without causing decoding errors. Subpictures may also be used for both the video and audio streams. This allows for improved encoder and decoder capabilities. Additionally, the improvements allow the encoder to code pictures more efficiently. This allows for network resources, memory resources, and / or Reduces the use of processing resources in the encoder and decoder.

[0005] Optionally, in any of the preceding aspects, another implementation of the aspect is When a picture contains a bottom border that does not coincide with the bottom border of the picture, the second subpicture is aligned. It specifies having a sub-picture height that contains several complete CTUs.

[0006] Optionally, in any of the preceding aspects, another implementation of the aspect further comprises: When a picture contains a right border that does not coincide with the right border of the picture, the third subpicture is aligned. It specifies having a sub-picture width that contains several complete CTUs.

[0007] Optionally, in any of the preceding aspects, another implementation of the aspect further comprises: The fourth subpicture is incomplete when the picture contains a bottom border that coincides with the bottom border of the picture. It specifies that the subpicture height includes the entire CTU.

[0008] In an embodiment, the present disclosure includes a method implemented in a decoder, the method comprising: The decoder receiver may detect that each subpicture has a right boundary that does not coincide with the right boundary of the picture. When the subpicture size is an integer multiple of the coding tree unit (CTU) size, With one or more sub-pictures segmented from the picture, such that the sub-picture width is included receiving a bitstream; and decoding the bitstream by a processor of the decoder. analyzing the frame to obtain one or more sub-pictures; decoding the sub-pictures to create a video sequence; and transferring the video sequence for display. The system may restrict subpictures to contain heights and widths that are multiples of the CTU size. However, a picture may contain heights and widths that are not multiples of the CTU size. Therefore, the subpicture size constraints are applicable to many picture layouts. This prevents the CTU from working properly. In the example shown, , the subpicture width and subpicture height are constrained. These constraints are If the bottom and right subpictures have heights and widths that are not multiples of the CTU size, they will be removed. By allowing the inclusion of any picture, Subpictures may also be used for both the video and audio streams. This allows for improved encoder and decoder capabilities. Additionally, the improvements allow the encoder to code pictures more efficiently. This allows for network resources, memory resources, and / or Reduces the use of processing resources in the encoder and decoder.

[0009] Optionally, in any of the preceding aspects, another implementation of the aspect is When a subpicture contains a bottom border that does not coincide with the bottom border of the picture, each subpicture is It specifies that the subpicture height must be an integer multiple of

[0010] Optionally, in any of the preceding aspects, another implementation of the aspect is such that when each sub-picture includes a right boundary that coincides with the right boundary of the picture, at least one of the sub-pictures includes a sub-picture width that is not an integer multiple of the CTU size.

[0011] Optionally, in any of the preceding aspects, another implementation of the aspect is such that when each sub-picture includes a bottom boundary that coincides with the bottom boundary of the picture, at least one of the sub-pictures includes a sub-picture height that is not an integer multiple of the CTU size.

[0012] Optionally, in any of the preceding aspects, another implementation of the aspect is such that the picture includes a picture width that is not an integer multiple of the CT U size.

[0013] Optionally, in any of the preceding aspects, another implementation of the aspect is such that the picture includes a picture height that is not an integer multiple of the CT U size.

[0014] Optionally, in any of the preceding aspects, another implementation of the aspect is such that the CTU size is specified to be measured in units of luma samples.

[0015] In one embodiment, the present disclosure includes a method implemented in an encoder, the method comprising, by a processor of the encoder, dividing a picture into a plurality of sub-pictures such that each sub-picture has a sub-picture width that is an integer multiple of the CTU size when each sub-picture does not include a right boundary that coincides with the right boundary of the picture, and, by the processor, encoding one or more of the sub-pictures into a bitstream, and encoding one or more of the sub-pictures into a bitstream, and encoding and storing the bitstream in a memory of the decoder for communication to the decoder. Some video systems are sub-pistoned to include heights and widths that are multiples of the CTU size. However, pictures may be limited in size and height that are not multiples of the CTU size. Therefore, the subpicture size constraints may include the number of picture layers. This prevents the subpicture from working properly with the CTU size. The subpicture width and height are constrained to be multiples of the size. while the subpicture is located at the right border of the picture or the bottom border of the picture, respectively. These constraints are removed when the bottom and right subpictures are not multiples of the CTU size. By allowing the height and width to be included separately, it is possible to Subpictures may be used with any picture. Furthermore, the improved functionality allows the encoder to process the picture more efficiently. This allows for efficient coding, saving network and memory resources. Reduce the use of bandwidth and / or processing resources in the encoder and decoder.

[0016] Optionally, in any of the preceding aspects, another implementation of the aspect is When a subpicture contains a bottom border that does not coincide with the bottom border of the picture, each subpicture is It specifies that the subpicture height must be an integer multiple of

[0017] Optionally, in any of the preceding aspects, another implementation of the aspect is When a subpicture contains a right border that coincides with the right border of the picture, at least one of the subpictures It is specified to include a sub-picture width that is not an integer multiple of the CTU size.

[0018] Optionally, in any of the preceding aspects, another implementation form of the aspect is that when each sub-picture includes a lower boundary that coincides with the lower boundary of the picture, at least one of the sub-pictures is specified to include a sub-picture height that is not an integer multiple of the CTU size.

[0019] Optionally, in any of the preceding aspects, another implementation form of the aspect is that the picture includes a picture width that is not an integer multiple of the CT U size.

[0020] Optionally, in any of the preceding aspects, another implementation form of the aspect is that the picture includes a picture height that is not an integer multiple of the CT U size.

[0021] Optionally, in any of the preceding aspects, another implementation form of the aspect is that the CTU size is specified to be measured in units of luma samples.

[0022] In one embodiment, the present disclosure includes a video coding device comprising a processor, a memory, a receiver coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the memory, the receiver, and the transmitter are configured to execute any of the methods of the preceding aspects.

[0023] In one embodiment, the present disclosure includes a non-transitory computer-readable medium comprising a computer program product for use by a video coding device, wherein the computer program product, when executed by a processor, causes the video coding device to perform any of the methods preceding. Comprising computer-executable instructions stored on a non-transitory computer-readable medium that cause a method of any of the aspects to be executed.

[0024] In one embodiment, the present disclosure provides receiving means for receiving a bitstream comprising one or more sub-pictures partitioned from a picture, wherein each sub-picture comprises a sub-picture width that is an integer multiple of the CTU size when each sub-picture includes a right boundary that does not coincide with the right boundary of the picture; analyzing means for analyzing the bitstream to obtain the one or more sub-pictures; decoding means for decoding the one or more sub-pictures to create a video sequence; and transferring means for transferring the video sequence for display, including a decoder.

[0025] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the decoder is further configured to execute a method of any of the preceding aspects.

[0026] In one embodiment, the present disclosure provides partitioning means for partitioning a picture into a plurality of sub-pictures such that each sub-picture includes a sub-picture width that is an integer multiple of the CTU size when each sub-picture includes a right boundary that does not coincide with the right boundary of the picture; encoding means for encoding one or more of the sub-pictures into a bitstream; and storage means for storing the bitstream for communication to a decoder, including an encoder.

[0027] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the encoder is further configured to execute a method of any of the preceding aspects.

[0028] ​​​​​​​​​​​​​ For clarity, any one of the foregoing embodiments may be combined with any one or more of any other of the foregoing embodiments to create new embodiments within the scope of the present disclosure.

[0029] These and other features will become more clearly understood from the following detailed description in conjunction with the accompanying drawings and claims.

[0030] For a more complete understanding of the present disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.

Brief Description of the Drawings

[0031]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

[0032] Illustrative implementations of one or more embodiments are provided below, but the disclosed system Any system and / or method, whether currently known or existing, It should be understood at the outset that this disclosure may be implemented using any number of techniques. including the example designs and implementations illustrated and described herein. The illustrative implementations, diagrams, and techniques illustrated below should not be considered limiting. Modifications may be made within the scope of the appended claims, along with their full range of equivalents.

[0033] Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coding Video Sequence (CVS), Joint Video Expert Team JVET, Motion Constrained Tile Set (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Level Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sequence Sequence Parameter Set (SPS), Versatile Video Coding (VVC), and Working Various acronyms are utilized herein, such as Draft (WD).

[0034] To reduce the size of video files while minimizing data loss, many video compression techniques can be utilized. For example, video compression techniques can include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video picture or a part of a video picture) may be partitioned into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTB), coding tree units (CTU), coding units (CU), and / or coding nodes. Video blocks within an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture. Video blocks within an inter-coded unidirectional prediction (P) or bidirectional prediction (B) slice of a picture may be coded by utilizing spatial prediction with respect to reference samples in neighboring blocks within the same picture, or temporal prediction with respect to reference samples in other reference pictures. A picture may also be referred to as a frame and / or an image, and a reference picture may also be referred with the vector and residual data indicating a difference between the coded block and the prediction block Thus, it is encoded. An intra-coded block is encoded according to the intra-coding mode and the residual data. For further compression, the residual data may be transferred from the pixel cell region to the transform region. These result in residual transform coefficients that may be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to produce a one-dimensional vector of transform coefficients. Entropy coding may be applied to achieve further compression. Such video compression techniques are discussed in more detail below.

[0035] To ensure that the encoded video can be decoded correctly, the video is encoded and decoded in accordance with the corresponding video coding standard. Video coding standards include the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) H.261, the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC Advanced Video Coding (AVC), also known as ISO / IEC MPEG-4 Part 10, and ITU-T H.265 or High Efficiency Video Coding (HEVC), also known as MPEG-H Part 2. AVC includes Scalable Video Coding (SVC), Multi-View Video Coding (MVC) and Multi-View Video Coding Plus Depth (MVC+D), as well as three-dimensional (3D) AV including extensions such as C(3D-AVC). HEVC includes extensions such as scalable HEVC (SHVC), multi-view HEVC (MV-HEV C), and 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVT) of ITU-T and ISO / IEC has started the development of a video coding standard called Versatile Video Coding (VVC). VVC is included in the Working Draft (WD), which includes JVET-L1001-v9.

[0036] To code a video image, the image is first segmented, and the segments are coded into a bitstream. Various picture segmentation methods are available. For example, an image can be segmented into ordinary slices, dependent slices, tiles, and / or according to wavefront parallel processing (WPP). For simplicity, when slicing into groups of CTBs for video coding, HEVC constrains the encoder so that ordinary slices, dependent slices, tiles, WPP, and combinations thereof can be used. Such segmentation can be applied to support adaptation of the maximum transmission unit (MTU) size, parallel processing, and reduction of end-to-end delay. The MTU represents the maximum amount of data that can be transmitted in a single packet. If the packet payload exceeds the MTU, the payload is split into two packets through a process called fragmentation. Ordinary slices, also simply called slices, can be reconstructed independently of other ordinary slices in the same picture, despite a certain degree of interdependence due to loop filtering operations.

[0037] It is a segmented part of a possible image. Each normal slice is encapsulated into a unique network abstraction layer (NAL) unit for transmission. Furthermore, intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Furthermore, intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture. It is encapsulated into a network abstraction layer (NAL) unit. Intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies along slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization utilizes minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Additionally, normal slices may be utilized to support compliance with MTU size requirements. Specifically, since normal slices are encapsulated into separate NAL units and can be coded independently, each normal slice must be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Therefore, the goals of parallelization and MTU size compliance can impose conflicting requirements on the slice layout within a picture.

[0038] Dependent slices are similar to normal slices but have a shortened slice header and enable the segmentation of image tree block boundaries without breaking intra-picture prediction. Therefore, dependent slices allow normal slices to be fragmented into multiple NAL units. Dependent slices are similar to normal slices but have a shortened slice header and enable the segmentation of image tree block boundaries without breaking intra-picture prediction. Therefore, dependent slices allow normal slices to be fragmented into multiple NAL units. Dependent slices are similar to normal slices but have a shortened slice header and enable the segmentation of image tree block boundaries without breaking intra-picture prediction. Therefore, dependent slices allow normal slices to be fragmented into multiple NAL units. This enables a part of a normal slice to be sent out before the encoding of the entire normal slice is completed, resulting in a reduction of the end-to-end delay.

[0039] A tile is a segmented part of an image created by horizontal and vertical boundaries that create tile columns and rows. Tiles may be coded in raster scan order (from right to left and from top to bottom). The scan order of CTBs is local within a tile. Thus, the CTBs of the first tile are coded in raster scan order before proceeding to the CTBs of the next tile. Similar to normal slices, tiles break the dependencies of intra-picture prediction and the dependencies of entropy decoding. However, since tiles may not be contained in individual NAL units, tiles may not be used for MTU size adaptation. Each tile can be processed by one processor / core, and the inter-processor / core communication utilized for intra-picture prediction between processing units decoding neighboring tiles is limited to carrying the shared slice header (when adjacent tiles are in the same slice) and performing sharing related to loop filtering of the reconstructed samples and metadata. When more than one tile is included in a slice, the entry point byte offset for each tile other than the first entry point offset in the slice may be signaled in the slice header. For each slice and tile, the conditions that 1) all coded tree blocks in the slice belong to the same slice and 2) all coded blocks in the tile belong to the same slice At least one of them should be satisfied.

[0040] In WPP, the picture is partitioned into a single row of CTBs. The entropy decoding and prediction mechanism may use data from CTBs in other rows. Parallel processing is made possible through parallel decoding of CTB rows. For example, the current row may be decoded in parallel with the preceding row. However, the decoding of the current row lags behind the decoding process of the preceding row by two CTBs. This delay ensures that data regarding the CTBs above and to the upper right of the current CTB in the current row is available before the current CTB is coded. This method appears as a wavefront when graphically represented. This shifted start enables parallelization using up to the same number of processors / cores as the number of CTB rows the picture contains. Since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication for enabling intra-picture prediction can be quite extensive. The WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, ordinary slices can be used with WPP, with some coding overhead, to perform MTU size adaptation upon request. from CTBs in other rows. Parallel processing is made possible through parallel decoding of CTB rows. For example, the current row may be decoded in parallel with the preceding row. However, the decoding of the current row lags behind the decoding process of the preceding row by two CTBs. This delay ensures that data regarding the CTBs above and to the upper right of the current CTB in the current row is available before the current CTB is coded. This method appears as a wavefront when graphically represented. This shifted start enables parallelization using up to the same number of processors / cores as the number of CTB rows the picture contains. Since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication for enabling intra-picture prediction can be quite extensive. The WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, the decoding of the current row lags behind the decoding process of the preceding row by two CTBs. This delay ensures that data regarding the CTBs above and to the upper right of the current CTB in the current row is available before the current CTB is coded. This method appears as a wavefront when graphically represented. This shifted start enables parallelization using up to the same number of processors / cores as the number of CTB rows the picture contains. Since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication for enabling intra-picture prediction can be quite extensive. The WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, this delay ensures that data regarding the CTBs above and to the upper right of the current CTB in the current row is available before the current CTB is coded. This method appears as a wavefront when graphically represented. This shifted start enables parallelization using up to the same number of processors / cores as the number of CTB rows the picture contains. Since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication for enabling intra-picture prediction can be quite extensive. The WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, this delay ensures that data regarding the CTBs above and to the upper right of the current CTB in the current row is available before the current CTB is coded. This method appears as a wavefront when graphically represented. This shifted start enables parallelization using up to the same number of processors / cores as the number of CTB rows the picture contains. Since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication for enabling intra-picture prediction can be quite extensive. The WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, this delay ensures that data regarding the CTBs above and to the upper right of the current CTB in the current row is available before the current CTB is coded. This method appears as a wavefront when graphically represented. This shifted start enables parallelization using up to the same number of processors / cores as the number of CTB rows the picture contains. Since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication for enabling intra-picture prediction can be quite extensive. The WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, this shifted start enables parallelization using up to the same number of processors / cores as the number of CTB rows the picture contains. Since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication for enabling intra-picture prediction can be quite extensive. The WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication for enabling intra-picture prediction can be quite extensive. The WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication for enabling intra-picture prediction can be quite extensive. The WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, the WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, the WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support adaptation of the MTU size. However, ordinary slices can be used with WPP, with some coding overhead, to perform MTU size adaptation upon request.

[0041] A tile may also include a motion constraint tile set. A motion constraint tile set (MCTS) is a tile set designed such that the associated motion vectors are constrained to indicate only integer sample positions within the MCTS and fractional sample positions that require only integer sample positions within the MCTS for interpolation. Furthermore, the temporal motion vectors derived from blocks outside the MCTS are constrained to indicate only integer sample positions within the MCTS and fractional sample positions that require only integer sample positions within the MCTS for interpolation. are constrained to indicate only integer sample positions within the MCTS and fractional sample positions that require only integer sample positions within the MCTS for interpolation. are constrained to indicate only integer sample positions within the MCTS and fractional sample positions that require only integer sample positions within the MCTS for interpolation. The use of motion vector candidates for Ctu prediction is not allowed. In this way, each MCTS may be decoded independently without the presence of tiles not included in the MCTS. Temporal MCTS supplement The enhancement information (SEI) message indicates the presence of MCTS in the bitstream and may be used to signal the MCTS. The MCTS SEI message provides supplementary information that can be used in MCTS sub-bitstream extraction (defined as part of the semantics of the SEI message) to generate a bitstream conforming to the MCTS set. The information includes several sets of extracted information, each of which defines the number of MCTS sets and contains the replacement video parameter sets (VPS), sequence parameter sets (SPS), and picture parameter sets (PPS) low-order bitstream payload (RBSP) bytes to be used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) may be rewritten or replaced, and the slice header may be updated, because one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) may use different values in the extracted sub-bitstream. A picture may also be divided into one or more sub-pictures. A sub-picture starts with a tile group having a tile_group_address equal to 0, and the length of the tile group / slice The use of motion vector candidates for Ctu prediction is not allowed. In this way, each MCTS may be decoded independently without the presence of tiles not included in the MCTS. Temporal MCTS supplement The enhancement information (SEI) message indicates the presence of MCTS in the bitstream and may be used to signal the MCTS. The MCTS SEI message provides supplementary information that can be used in MCTS sub-bitstream extraction (defined as part of the semantics of the SEI message) to generate a bitstream conforming to the MCTS set. The information includes several sets of extracted information, each of which defines the number of MCTS sets and contains the replacement video parameter sets (VPS), sequence parameter sets (SPS), and picture parameter sets (PPS) low-order bitstream payload (RBSP) bytes to be used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) may be rewritten or replaced, and the slice header may be updated, because one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) may use different values in the extracted sub-bitstream. A picture may also be divided into one or more sub-pictures. A sub-picture starts with a tile group having a tile_group_address equal to 0, and the length of the tile group / slice ) The enhancement information (SEI) message indicates the presence of MCTS in the bitstream and may be used to signal the MCTS. The MCTS SEI message provides supplementary information that can be used in MCTS sub-bitstream extraction (defined as part of the semantics of the SEI message) to generate a bitstream conforming to the MCTS set. The information includes several sets of extracted information, each of which defines the number of MCTS sets and contains the replacement video parameter sets (VPS), sequence parameter sets (SPS), and picture parameter sets (PPS) low-order bitstream payload (RBSP) bytes to be used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) may be rewritten or replaced, and the slice header may be updated, because one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) may use different values in the extracted sub-bitstream. A picture may also be divided into one or more sub-pictures. A sub-picture starts with a tile group having a tile_group_address equal to 0, and the length of the tile group / slice ) The enhancement information (SEI) message indicates the presence of MCTS in the bitstream and may be used to signal the MCTS. The MCTS SEI message provides supplementary information that can be used in MCTS sub-bitstream extraction (defined as part of the semantics of the SEI message) to generate a bitstream conforming to the MCTS set. The information includes several sets of extracted information, each of which defines the number of MCTS sets and contains the replacement video parameter sets (VPS), sequence parameter sets (SPS), and picture parameter sets (PPS) low-order bitstream payload (RBSP) bytes to be used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) may be rewritten or replaced, and the slice header may be updated, because one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) may use different values in the extracted sub-bitstream. A picture may also be divided into one or more sub-pictures. A sub-picture starts with a tile group having a tile_group_address equal to 0, and the length of the tile group / slice The use of motion vector candidates for Ctu prediction is not allowed. In this way, each MCTS may be decoded independently without the presence of tiles not included in the MCTS. Temporal MCTS supplement

[0042] A picture may also be divided into one or more sub-pictures. A sub-picture starts with a tile group having a tile_group_address equal to 0, and the length of the tile group / slice A picture may also be divided into one or more sub-pictures. A sub-picture starts with a tile group having a tile_group_address equal to 0, and the length of the tile group / slice It is a square set. Since each sub-picture may refer to a separate PPS, it may have separate tile divisions. Sub-pictures may be treated like pictures in the decoding process. The reference sub-picture for decoding the current sub-picture extracts an area at the same position as the current sub-picture from the reference picture in the decoded picture buffer and is generated thereby. The extracted area is treated as the decoded sub-picture. Inter-prediction may be performed between sub-pictures of the same size and at the same position within a picture. A tile group, also known as a slice, is a sequence of related tiles within a picture or sub-picture. To determine the position of sub-pictures within a picture, several items can be derived. For example, each current sub-picture may be placed at the next unoccupied position in the CTU raster scan order within a picture large enough to contain the current sub-picture within the picture boundary. Furthermore, picture division may be based on picture-level tiles and sequence-level tiles. The sequence-level tile may include the function of MCTS and may be implemented as a sub-picture. For example, the picture-level tile may be defined as a rectangular area of coding tree blocks within a specific tile column and a specific tile row within a picture. The sequence-level tile may be defined as a set of rectangular areas of coding tree blocks included in different frames, and each rectangular area further includes one or more

[0043] picture-level tiles, and the set of rectangular areas of coding tree blocks is similar. picture-level tiles, and the set of rectangular areas of coding tree blocks is similar. ​​​​​​It is decodable independently from any other set of rectangular regions. Sequence level tiling A group set (STGPS) is a group of such sequence level tilings. STGPS may be signaled in non-video coding layer (VCL) NAL units along with the relevant identifier (ID) in the NAL unit header.

[0044] The previous sub-picture based partitioning scheme may be associated with a problem. For example when the sub-picture is valid, tiling (partitioning of the sub-picture into tiles) within the sub-picture may be used to support parallel processing. The tile partitioning of the sub-picture for which parallel processing is the goal may vary from picture to picture (e.g., for the purpose of balancing the parallel processing load), and thus may be managed at the picture level (e.g., in the PPS). However, the sub-picture partitioning (partitioning of the picture into sub-pictures) may be utilized to support regions of interest (ROI) and sub-picture based picture access. In such a case the signaling of the sub-picture or MCTS in the PPS is not efficient.

[0045] In another example, when any sub-picture within a picture is coded as a temporally motion-constrained sub-picture all sub-pictures within the picture may be coded as temporally motion-constrained sub-pictures. Such picture partitioning may be limited . For example, coding a sub-picture as a temporally motion-constrained sub-picture may reduce the coding efficiency in exchange for additional functionality. However, regions of interest ​​​​In the case of base application, usually only one or several of the sub-pictures are subject to the temporal motion constraint Use the sub-picture-based function. Therefore, the remaining sub-pictures suffer from a decrease in coding efficiency without bringing any practical benefits

[0046] In another example, the syntax element for specifying the size of the sub-picture may be specified in units of the luma CTU size Therefore, both the width and height of the sub-picture should be an integer multiple of CtbSizeY. This mechanism for specifying the width and height of the sub-picture may bring various problems For example, the sub-picture division is only applicable to pictures with a picture width and / or picture height that is an integer multiple of CtbSizeY. This makes the sub-picture division unavailable for pictures containing dimensions that are not integer multiples of CTbSizeY When the picture dimensions are not integer multiples of CtbSizeY, if the sub-picture division is applied to the width and / or height of the picture, the derivation of the sub-picture width and / or sub-picture height in terms of luma sample units for the rightmost sub-picture and the bottommost sub-picture will be inaccurate And for some coding tools, such inaccurate derivations will lead to incorrect results

[0047] In another example, the position of the sub-picture in the picture may not be signaled Instead, the position is derived using the following rule. The current sub-picture is placed at the next such unoccupied position in the CTU raster scan order within a picture large enough to contain the sub-picture within the picture boundaries In some cases, the sub-picture is arranged in such a way ​​​​​​​​Deriving the sub-picture position may cause errors. For example, if a sub-picture is lost during transmission, then the positions of other sub-pictures are derived incorrectly, and the decoded samples are placed in the wrong positions. When sub- pictures arrive in the wrong order, the same problem applies.

[0048] In another example, decoding a sub-picture may require extracting the sub- picture at the same position in the reference picture. This may impose additional complexity and resulting burden in terms of processor and memory resource usage.

[0049] In another example, when a sub-picture is designed as a time motion constraint sub-picture, the loop filter that scans the sub-picture boundaries is disabled. This occurs regardless of whether the loop filter that scans the tile boundaries is enabled. Such constraints can be too strict and may result in visual artifacts for video pictures that utilize multiple sub-pictures.

[0050] In another example, the relationship between SPS, STGPS, PPS, and tile group headers is as follows. STGPS refers to SPS, PPS refers to STGPS, and the tile group header / slice header refers to PPS. However, STGPS and PPS must be orthogonal, rather than PPS referring to STGPS. The aforementioned configuration may also not allow all tile groups of the same picture to refer to

[0051] the same PPS. In another example, each STGPS may include IDs for the four sides Sub - pictures sharing the same boundary are identified using an ID, and their relative spatial relationships can be defined. However, in some cases, such information may not be sufficient to derive the position and size information for a sequence - level tile group set. In other cases, signaling the position and size information may be redundant.

[0052] In another example, the STGPS ID may be signaled in the NAL unit header of the VCL NAL unit using 8 bits. This may assist in sub - picture extraction. Such signaling may unnecessarily increase the length of the NAL unit header. Another problem is that, unless the sequence - level tile group set is constrained to prevent duplicates, one tile group may be associated with multiple sequence - level tile group sets.

[0053] To address one or more of the above - mentioned problems, various mechanisms are disclosed herein. In a first example, the sub - picture layout information is included in the SPS instead of the PPS. The sub - picture layout information includes sub - picture position and sub - picture size. The sub - picture position is the offset between the top - left sample of the sub - picture and the top - left sample of the picture. The sub - picture size is the height and width of the sub - picture as measured in luma samples. As mentioned above, since the tile may vary from picture to picture, some streams include tiling information in the PPS. However, for ROI application and sub - picture - based ​​​​​Sub - pictures may be used to support access to these. These features do not vary picture - by - picture. Furthermore, a video sequence may contain a single SPS (or one per video segment), and may contain one PPS per picture. Placing sub - picture layout information in the SPS ensures that the layout is signaled only once for the sequence / segment rather than redundantly signaled for each PPS. Thus, signaling the sub - picture layout in the SPS increases coding efficiency and reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. Also, some systems have sub - picture information derived by the decoder. Signaling the sub - picture information reduces the probability of errors in case of packet loss and supports additional functions related to extracting sub - pictures. Thus, signaling the sub - pic ture layout in the SPS enhances the functionality of the encoder and / or

[0054] In a second example, the width of the sub - picture and the height of the sub - picture are constrained to be multiples of the CTU size. However, these constraints are removed when the sub - picture is placed at the right border of the picture or at the bottom border of the picture, respectively. As described above, some video systems may constrain sub - pictures to have heights and widths that are multiples of the CTU size. This prevents By allowing the inclusion of sub-pictures, the sub-pictures can be decoded without causing decoding errors. It may be used with any picture. This is a function of the encoder and decoder functions. Additionally, the improved functionality allows the encoder to code pictures more efficiently. This allows for network linking at the encoder and decoder. Reduce the use of sources, memory resources, and / or processing resources.

[0055] In the third example, the subpicture is constrained to encompass the picture without gaps or overlaps. As mentioned above, some video coding systems use the Allows for gaps and overlaps, which means that a tile group / slice can have multiple sub-tiles. This creates the possibility of associating a picture with the image, if this is allowed by the encoder. , the decoder will be able to use such a codec even when the decoding scheme is rarely used. Subpicture gaps and overlaps must be supported. By not allowing multiple subpictures, the decoder avoids latency when determining the size and position of the subpicture. Reduces decoder complexity since potential gaps and overlaps do not need to be considered. Furthermore, by not allowing gaps and overlaps between subpictures, the video sequence can be This allows the encoder to avoid considering gap and overlap cases when selecting the encoding for a sequence. This reduces the complexity of the rate-distortion optimization (RDO) process in the encoder. Therefore, avoiding gaps and overlaps reduces memory usage in the encoder and decoder. May reduce resource and / or processing resource usage.

[0056] In the fourth example, a flag for indicating when a sub-picture is a temporally motion-constrained sub-picture may be signaled in the SPS. As described above, some systems either set all sub-pictures as temporally motion-constrained sub-pictures, or may not fully permit the use of temporally motion-constrained sub-pictures. Such temporally motion constrained sub-pictures provide an independent extraction function at the expense of a reduction in coding efficiency. However, in the case of region-of-interest-based applications, the region of interest should be coded for independent extraction, while the regions outside the region of interest do not require such a function. Thus, the remaining sub-pictures result in a reduction in coding efficiency without bringing any real benefit. Therefore, this flag enables a mixture of temporally motion-constrained sub-pictures that provide an independent extraction function and non-motion-constrained sub-pictures in order to increase coding efficiency when independent extraction is not desired. Therefore, this flag enables an improvement in functionality and / or coding efficiency, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.

[0057] In the fifth example, a complete set of sub-picture IDs is signaled in the SPS, and the slice header includes the sub-picture ID indicating the sub-picture containing the corresponding slice. As described above, some systems signal the relative picture position with respect to other sub-pictures. This causes problems when a sub-picture is lost or when it is extracted separately. By specifying each sub-picture with an ID, the sub-picture It can be positioned and sized without referring to other sub - pictures. And this supports applications such as error correction and extracting only a part of the sub - picture to avoid transmission of other sub - pictures A complete list of all sub - picture IDs can be transmitted in the SPS along with the relevant slice information. Each slice header may include a sub - picture ID indicating the sub - picture containing the corresponding slice. In this way, the sub - picture and the corresponding slice can be extracted and positioned without referring to other sub - pictures. Thus, the sub - picture ID helps to improve the function and / or the coding efficiency, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.

[0058] In the sixth example, a level is signaled for each sub - picture. In some video coding systems, a level is signaled for a picture. The level indicates the hardware resources required to decode the picture. As described above, in some cases, different sub - pictures may have different functions and thus may be treated differently during the coding process. Therefore, the picture - based level may not be useful for decoding some pictures. Thus, the present disclosure includes a level for each sub - picture. In this way, each sub - picture is decoded without imposing an overly high decoding requirement on sub - pictures coded according to a more complex mechanism, without unduly burdening the decoder and independently of other sub - pictures. ​​​​​​​​​​​​can be coded. The sub-picture level information signaled helps improve functionality and / or coding efficiency, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. and / or improve coding efficiency, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. Figure 1 is a flowchart of an exemplary operating method 100 for coding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size reduces the associated bandwidth overhead while enabling the compressed video file to be transmitted to the user. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to stably reconstruct the video signal.

[0059] Figure 1 is a flowchart of an exemplary operating method 100 for coding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size reduces the associated bandwidth overhead while enabling the compressed video file to be transmitted to the user. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to stably reconstruct the video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size reduces the associated bandwidth overhead while enabling the compressed video file to be transmitted to the user. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to stably reconstruct the video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size reduces the associated bandwidth overhead while enabling the compressed video file to be transmitted to the user. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to stably reconstruct the video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size reduces the associated bandwidth overhead while enabling the compressed video file to be transmitted to the user. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to stably reconstruct the video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size reduces the associated bandwidth overhead while enabling the compressed video file to be transmitted to the user. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to stably reconstruct the video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size reduces the associated bandwidth overhead while enabling the compressed video file to be transmitted to the user. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to stably reconstruct the video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size reduces the associated bandwidth overhead while enabling the compressed video file to be transmitted to the user. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to stably reconstruct the video signal.

[0060] In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live video streaming. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, give the visual effect of movement. The frames include luma components (or luma samples), pixels represented with respect to light as referred to herein, and chroma components. In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live video streaming. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, give the visual effect of movement. The frames include luma components (or luma samples), pixels represented with respect to light as referred to herein, and chroma components. In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live video streaming. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, give the visual effect of movement. The frames include luma components (or luma samples), pixels represented with respect to light as referred to herein, and chroma components. In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live video streaming. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, give the visual effect of movement. The frames include luma components (or luma samples), pixels represented with respect to light as referred to herein, and chroma components. In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live video streaming. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, give the visual effect of movement. The frames include luma components (or luma samples), pixels represented with respect to light as referred to herein, and chroma components. In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live video streaming. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, give the visual effect of movement. The frames include luma components (or luma samples), pixels represented with respect to light as referred to herein, and chroma components. In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live video streaming. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, give the visual effect of movement. The frames include luma components (or luma samples), pixels represented with respect to light as referred to herein, and chroma components. It includes pixels that are expressed in terms of a color called a component (or color sample). In some examples, the frame may also include depth values to support three-dimensional viewing.

[0061] In step 103, the video is segmented into blocks. The segmentation involves re-dividing the pixels of each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), the frame is first divided into Coding Tree Units (CTUs), where a CTU is a block of a predetermined size (e.g., 64 pixels by 64 pixels). A CTU contains both luma samples and chroma samples. The coding tree may be used to divide the CTU into blocks and then recursively re-divide the blocks until a configuration that supports further encoding is achieved. For example, the luma component of the frame may be re-divided until the individual blocks contain relatively uniform illumination values. Further, the chroma component of the frame may be re-divided until the individual blocks contain relatively uniform color values. Therefore, the segmentation mechanism varies depending on the content of the video frame.

[0062] In step 105, various compression mechanisms are utilized to compress the image blocks segmented in step 103. For example, inter prediction and / or intra prediction may be utilized. Inter prediction is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Therefore, among the reference frames The block that depicts an object does not need to be repeatedly described in adjacent frames. Specifically, an object such as a table can remain in a fixed position over multiple frames. Therefore, the table is described once, and adjacent frames can return to the reference frame and refer to it. To match objects across multiple frames, a pattern matching mechanism may be used. Furthermore, a moving object may be represented across multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving around on the screen over multiple frames. Motion vectors may be used to describe such motion. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in the reference frame. Therefore, inter prediction can encode the image blocks in the current frame as a set of motion vectors that indicate the offsets from the corresponding blocks in the reference frame. Specifically, an object such as a table can remain in a fixed position over multiple frames. Therefore, the table is described once, and adjacent frames can return to the reference frame and refer to it. To match objects across multiple frames, a pattern matching mechanism may be used. Furthermore, a moving object may be represented across multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving around on the screen over multiple frames. Motion vectors may be used to describe such motion. A motion vector is a two-dimensional vector that gives the offset from the coordinates of an object in a frame to the coordinates of the object in the reference frame. Therefore, inter prediction can encode the image blocks in the current frame as a set of motion vectors that indicate the offsets from the corresponding blocks in the reference frame.

[0063] Intra prediction encodes blocks in a common frame. Intra prediction utilizes the fact that the luma and chroma components tend to cluster in a frame. For example, green spots in a part of a tree tend to be positioned next to similar green spots. Intra prediction uses multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. The directional mode indicates that the current block is similar to / the same as the samples of neighboring blocks in the corresponding direction. The planar mode interpolates a series of blocks along a row / column (e.g., a plane) based on neighboring blocks at the end of the row. Intra prediction encodes blocks in a common frame. Intra prediction utilizes the fact that the luma and chroma components tend to cluster in a frame. For example, green spots in a part of a tree tend to be positioned next to similar green spots. Intra prediction uses multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. The directional mode indicates that the current block is similar to / the same as the samples of neighboring blocks in the corresponding direction. The planar mode interpolates a series of blocks along a row / column (e.g., a plane) based on neighboring blocks at the end of the row. ​​​​ which can be shown. The planar mode substantially shows a relatively constant gradient by changing the value and utilizes a smooth transition of light / color across rows / columns. The DC mode is utilized for boundary smoothing and indicates that all neighboring blocks associated with the angular direction of the directional prediction mode have the same average value as the block. For boundary smoothing, it is associated with the average value of all neighboring blocks related to the angular direction of the directional prediction mode and indicates that the block is similar / the same. Therefore, the intra prediction block can represent the image block as various relationship prediction mode values instead of the actual value. Furthermore, the inter prediction block can represent the image block as a motion vector value instead of the actual value. In either case, the prediction block may not exactly represent the image block in some cases. Any difference is accumulated in the residual block. To further compress the file, a transformation may be applied to the residual block. In step 107, various filtering techniques may be applied. In HEVC, the filter is applied according to the in-loop filtering method. The block-based prediction discussed above may result in the creation of a block-based image at the decoder. Furthermore, the block-based prediction method may encode the block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering method repeatedly applies noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to the block / frame. These filters are applied such that the encoded file can be accurately reconstructed, such that such blocking artifacts are removed. are removed so that the encoded file can be accurately reconstructed. Therefore, the prediction block may not exactly represent the image block in some cases. Any difference is accumulated in the residual block. To further compress the file, a transformation may be applied to the residual block. are removed so that the encoded file can be accurately reconstructed. are removed so that the encoded file can be accurately reconstructed.

[0064] In step 107, various filtering techniques may be applied. In HEVC, the filter is applied according to the in-loop filtering method. The block-based prediction discussed above may result in the creation of a block-based image at the decoder. Furthermore, the block-based prediction method may encode the block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering method repeatedly applies noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to the block / frame. These filters are applied such that the encoded file can be accurately reconstructed, such that such blocking artifacts are removed. The block-based prediction discussed above may result in the creation of a block-based image at the decoder. Furthermore, the block-based prediction method may encode the block and then reconstruct the encoded block for later use as a reference block. The block-based prediction method may encode the block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering method repeatedly applies noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to the block / frame. These filters are applied such that the encoded file can be accurately reconstructed, such that such blocking artifacts are removed. are removed so that the encoded file can be accurately reconstructed. are removed so that the encoded file can be accurately reconstructed. Furthermore, these filters reduce the artifacts in the reconstructed reference blocks. Reduce artifacts so that the artifacts are based on the reconstructed reference block. This can lead to additional artifacts in subsequent blocks that are coded based on It will be lower.

[0065] Once the video signal has been segmented, compressed and filtered, in step 109: The resulting data is encoded in a bitstream, which is the same as the bitstream discussed above. to support the encoded data and reconstruction of the appropriate video signal at the decoder. For example, such data may be data, prediction data, residual blocks, and coding instructions to the decoder. The bitstream may contain flags such as: The bitstream may also be broadcast to multiple decoders. The creation of the bitstream may be iterative and / or multicast. Therefore, steps 101, 103, 105, 107, and 109 are performed to The time series may occur sequentially and / or simultaneously across multiple time series and blocks. The order in which the video is recorded is presented for clarity and ease of discussion. It is not intended to limit the coding process to any particular order.

[0066] In step 111, the decoder receives the bitstream and begins the decoding process. Specifically, the decoder uses an entropy decoding method to decode the bitstream. Convert it into the corresponding syntax and video data. In step 111, the decoder however, uses the syntax data from the bitstream to determine the partitioning for the frame This partitioning must match the result of the block partitioning in step 103 Entropy coding / decoding as used in step 111 is described here The encoder makes many choices during the compression process, such as selecting a block partitioning method from several possible options based on the spatial arrangement of the values in the input image Many choices are made during the compression process, such as selecting a block partitioning method from several possible options based on the spatial arrangement of the values in the input image Precise selection signaling may use a large number of bins. In this specification, a bin is a binary value treated as a variable (e.g., a bit value that may change depending on the situation) Entropy coding allows the encoder to discard any option that is clearly not executable for a particular case, leaving a set of acceptable options. Then, each acceptable option is assigned a codeword. The length of the codeword is based on the number of acceptable options (e.g., 1 bin for 2 options, 2 bins for 3 to 4 options, etc.) The encoder then encodes the codeword for the selected option. This method reduces the size of the codeword because the codeword is of the desired size to uniquely indicate a selection from a small subset of acceptable options rather than from a large possible set of all options The decoder then decodes the selection by determining the set of acceptable options in the same way as the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder This partitioning must match the result of the block partitioning in step 103 Entropy coding allows the encoder to discard any option that is clearly not executable for a particular case, leaving a set of acceptable options. Then, each acceptable option is assigned a codeword. The length of the codeword is based on the number of acceptable options (e.g., 1 bin for 2 options, 2 bins for 3 to 4 options, etc.) The encoder then encodes the codeword for the selected option. This method reduces the size of the codeword because the codeword is of the desired size to uniquely indicate a selection from a small subset of acceptable options rather than from a large possible set of all options The decoder then decodes the selection by determining the set of acceptable options in the same way as the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder The decoder then decodes the selection by determining the set of acceptable options in the same way as the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder The decoder then decodes the selection by determining the set of acceptable options in the same way as the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder The decoder then decodes the selection by determining the set of acceptable options in the same way as the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder

[0067] In step 113, the decoder performs block decoding. Specifically, the decoder uses inverse transformation to generate a residual block. Next, the decoder uses the residual block and the corresponding prediction block to reconstruct an image block according to the partition. The prediction block may include both an intra prediction block and an inter prediction block as generated by the encoder in step 105. The reconstructed image block is then positioned into the frame of the reconstructed video signal according to the partition data determined in step 111. The syntax for step 113 may also be signaled in the bitstream via entropy

[0068] In step 115, in the encoder, filtering is performed on the frame of the reconstructed video signal in a manner similar to step 107. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove blocking artifacts. Once the frame is filtered, the video signal may be output to a display in step 117 for viewing by an end user.

[0069] FIG. 2 is a schematic diagram of an exemplary encoding and decoding (codec) system 2 00 for video coding. Specifically, the codec system 200 provides functions to support the implementation of operation method 100. The codec system 200 includes both an encoder and a decoder ​is generalized to depict components used therein. The codec sys tem 200 receives and separates a video signal as discussed with respect to steps 101 and 103 in operation method 100, which results in a separated video signal 201. The codec sys tem 200 then compresses the separated video signal 201 into a coded bit stream when operating as an encoder, as discussed with respect to steps 105, 107, and 109 of method 100. When operating as a decoder, the codec system 200 generates an output video signal from the bit stream as discussed with respect to steps 111, 113, 115, and 117 of operation method 100. The codec system 200 includes a general-purpose coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219 , a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, as well as a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are combined as shown. In FIG. 2, the black lines indicate the movement of data to be encoded / decoded, and the dashed lines indicate the movement of control data that controls the operation of other components. The components of the codec sys tem 200 may all be present within the encoder. The decoder may include a subset of the components of the codec sys tem 200. For example, the decoder may include a subset of the components of the codec system 200. For example, the decoder from the bit stream as discussed with respect to steps 111, 113, 115, and 117 of operation method 100. The codec system 200 includes a general-purpose coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219 , a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, as well as a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are combined as shown. In FIG. 2, the black lines indicate the movement of data to be encoded / decoded, and the dashed lines indicate the movement of control data that controls the operation of other components. The components of the codec sys tem 200 may all be present within the encoder. The decoder may include a subset of the components of the codec sys tem 200. For example, the decoder may include a subset of the components of the codec system 200. For example, the decoder from the bit stream as discussed with respect to steps 111, 113, 115, and 117 of operation method 100. The codec system 200 includes a general-purpose coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219 , a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, as well as a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are combined as shown. In FIG. 2, the black lines indicate the movement of data to be encoded / decoded, and the dashed lines indicate the movement of control data that controls the operation of other components. The components of the codec sys tem 200 may all be present within the encoder. The decoder may include a subset of the components of the codec sys tem 200. For example, the decoder may include a subset of the components of the codec system 200. For example, the decoder may include a subset of the components of the codec system 200. For example, the decoder Da may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a reference picture buffer component 223. These components are described here.

[0070] The segmented video signal 201 is a captured video sequence that is segmented into pixel blocks by a coding tree. The coding tree uses various partitioning modes to further divide pixel blocks into smaller pixel blocks. These blocks can then be further divided into even smaller blocks. The blocks may be referred to as nodes on the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is divided is called the depth of the node / coding tree. In some cases, the divided blocks may be included in a coding unit (CU). For example, a CU may be the lower part of a CTU that includes a luma block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block, along with corresponding syntax instructions for the CU. The partitioning modes may include a binary tree (BT), a ternary tree (TT), and a quadtree (QT) that are used to divide each node into two, three, or four child nodes, respectively, depending on the partitioning 7. and is transferred to the motion estimation component 221.

[0071] The general-purpose coder control component 211 is configured to make decisions regarding the coding of the video sequence into the bit stream of the video image according to the constraints of the application example. For example, the general-purpose coder control component 211 manages the optimization of the bit rate / bit stream size versus the reconstructed quality. Such decisions may be made based on the availability of memory space / bandwidth and the requirements of the image resolution. The general-purpose coder control component 211 also considers the transmission speed in order to reduce the problems of buffer underrun and overrun, and manages the buffer utilization. To manage these problems, the general-purpose coder control component 211 manages the partitioning, prediction, and filtering by other components. For example, the general-purpose coder control component 211 may increase the resolution by dynamically increasing the complexity of compression to improve the bandwidth utilization, or may decrease the resolution and bandwidth utilization by reducing the complexity of compression. Therefore, the general-purpose coder control component 211 controls other components of the codec system 200 to balance the quality of the video signal reconstruction and the bit rate problem. The general-purpose coder control component 211 creates control data which controls the operation of other components. The control data is also transferred to the header formatting and CABAC component 231 and is encoded in the bit stream to signal the parameters for decoding at the decoder. The partitioned video signal 201 is also input to the motion estimation component 221 for inter-prediction and is transferred to the motion estimation component 221. The control data is formatted by the header formatter and is also transferred to the CABAC component 231 and is encoded in the bit stream to signal the parameters for decoding at the decoder.

[0072] The partitioned video signal 201 is also input to the motion estimation component 221 for inter-prediction It is sent to the call and motion compensation component 219. The frame of the segmented video signal 201 or the slice may be divided into a plurality of video blocks. The motion estimation component 22 1 and the motion compensation component 219 perform inter-prediction coding of the received video block relative to one or more blocks in one or more reference frames to perform temporal prediction. The codec system 200 executes a plurality of coding paths to, for example, select an appropriate coding mode for each block of video data.

[0073] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process of generating a motion vector that estimates the motion of a video block. The motion vector may indicate, for example, the displacement of a coded object relative to a prediction block. The prediction block is a block that is found to closely match the block to be coded with respect to pixel differences. The prediction block may also be referred to as a reference block. Such pixel differences may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. HEVC utilizes several coded objects, including coding tree units (CTUs), coding tree blocks (CTBs), and coding units (CUs). For example, a CTU can be divided into CTBs, and then the CTB can be divided into CUs for inclusion. A CU is a prediction unit (PU) that contains prediction data and / or ​​​​​​​​It can be encoded as a transform unit (TU) containing transformed residual data for the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate distortion analysis as part of the rate distortion optimization process. For example, the motion estimation component 221 may determine a plurality of reference blocks, a plurality of motion vectors, etc. for the current block / frame, and select a reference block, a motion vector, etc. having the best rate distortion characteristics. The best rate distortion characteristics balance the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final encoding). In some examples, the codec system 200 may calculate values for sub - integer pixel positions of reference pictures stored in the decoded picture buffer component 2 23. For example, the video codec system 200 may interpolate values at quarter - pixel positions, eighth - pixel positions, or other fractional pixel positions of the reference picture. Thus, the motion estimation component 221 may perform motion search for integer pixel positions and fractional pixel positions

[0074] and output motion vectors with fractional pixel accuracy. The motion estimation component 22 1 calculates the motion vector for the PU of a video block in an inter - coded slice by comparing the position of the PU with the position of the predicted block of the reference picture. The motion estimation component 221 outputs the calculated motion vector as motion data for encoding to the header formatting and CABAC component 231, and outputs the motion to the motion compensation component 219. component 219. and output motion vectors with fractional pixel accuracy. The motion estimation component 22 1 calculates the motion vector for the PU of a video block in an inter - coded slice by comparing the position of the PU with the position of the predicted block of the reference picture. The motion estimation component 221 calculates the motion vector for the PU of a video block in an inter - coded slice by comparing the position of the PU with the position of the predicted block of the reference picture. The motion estimation component 221 outputs the calculated motion vector as motion data for encoding to the header formatting and CABAC component 231, and outputs the motion to the motion compensation component 219. component 219.

[0075] The motion compensation performed by the motion compensation component 219 fetches or generates a prediction block based on the motion vector determined by the motion estimation component 2 21. This may be accompanied by fetching or generating a prediction block based on the motion vector determined by the motion estimation component 221. Again, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated in some examples. When receiving a motion vector for the current video block's PU, the motion compensation component 219 may locate the prediction block indicated by the motion vector. The residual video block is then formed by subtracting the pixel values of the prediction block from the pixel values of the currently coded current video block to form pixel difference values. Generally, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The prediction block and the residual block are transferred to the transform scaling and quantization component 213. The segmented video signal 201 is also sent to the intra-picture estimation component 215 and the intra-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated, but are shown separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217, as described above, perform motion estimation between frames in the same way as the motion estimation component 221 and the motion compensation component 219. The intra-picture estimation component 215 and the intra-picture prediction component 217 may be functionally integrated in some examples. When receiving a motion vector for the current video block's PU, the motion compensation component 219 may locate the prediction block indicated by the motion vector. The residual video block is then formed by subtracting the pixel values of the prediction block from the pixel values of the currently coded current video block to form pixel difference values. Generally, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The prediction block and the residual block are transferred to the transform scaling and quantization component 213.

[0076] The segmented video signal 201 is also sent to the intra-picture estimation component 215 and the intra-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated, but are shown separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217, as described above, perform motion estimation between frames in the same way as the motion estimation component 221 and the motion compensation component 219. The intra-picture estimation component 215 and the intra-picture prediction component 217 may be functionally integrated in some examples. When receiving a motion vector for the current video block's PU, the motion compensation component 219 may locate the prediction block indicated by the motion vector. The residual video block is then formed by subtracting the pixel values of the prediction block from the pixel values of the currently coded current video block to form pixel difference values. Generally, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The prediction block and the residual block are transferred to the transform scaling and quantization component 213. As an alternative to inter prediction performed by the motion compensation component 219, the current block in the current frame is intra predicted with respect to the current block. Specifically, the intra picture estimation component 215 determines the intra prediction mode to be used to encode the current block. In some examples, the intra picture estimation component 215 selects an appropriate intra prediction mode from a plurality of tested intra prediction modes to encode the current block. The selected intra prediction mode is then transferred to the header formatting and CABAC component 231 for encoding.

[0077] For example, the intra picture estimation component 215 calculates rate-distortion values using rate-distortion analysis for various tested intra prediction modes and selects the intra prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original uncoded block that was encoded to produce the encoded block, as well as the bit rate (e.g., the number of bits) used to produce the encoded block. The intra picture estimation component 215 calculates ratios from the distortion and rate for various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block. Additionally, the intra picture estimation component 215 may be configured to code depth blocks of depth maps using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).

[0078] When the intra-picture prediction component 217 is implemented on the encoder, it generates a residual block from a prediction block based on a selected intra prediction mode determined by the intra-picture estimation component 215, or when implemented on the decoder, it may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma and chroma components. When the intra-picture prediction component 217 is implemented on the encoder, it generates a residual block from a prediction block based on a selected intra prediction mode determined by the intra-picture estimation component 215, or when implemented on the decoder, it may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma and chroma components. When the intra-picture prediction component 217 is implemented on the encoder, it generates a residual block from a prediction block based on a selected intra prediction mode determined by the intra-picture estimation component 215, or when implemented on the decoder, it may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma and chroma components. When the intra-picture prediction component 217 is implemented on the encoder, it generates a residual block from a prediction block based on a selected intra prediction mode determined by the intra-picture estimation component 215, or when implemented on the decoder, it may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma and chroma components. When the intra-picture prediction component 217 is implemented on the encoder, it generates a residual block from a prediction block based on a selected intra prediction mode determined by the intra-picture estimation component 215, or when implemented on the decoder, it may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma and chroma components. When the intra-picture prediction component 217 is implemented on the encoder, it generates a residual block from a prediction block based on a selected intra prediction mode determined by the intra-picture estimation component 215, or when implemented on the decoder, it may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma and chroma components. When the intra-picture prediction component 217 is implemented on the encoder, it generates a residual block from a prediction block based on a selected intra prediction mode determined by the intra-picture estimation component 215, or when implemented on the decoder, it may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma and chroma components. When the intra-picture prediction component 217 is implemented on the encoder, it generates a residual block from a prediction block based on a selected intra prediction mode determined by the intra-picture estimation component 215, or when implemented on the decoder, it may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma and chroma components.

[0079] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block with residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying scale factors to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 may also perform quantization on the scaled residual information. The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block with residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying scale factors to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 may also perform quantization on the scaled residual information. The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block with residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying scale factors to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 may also perform quantization on the scaled residual information. The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block with residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying scale factors to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 may also perform quantization on the scaled residual information. The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block with residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying scale factors to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 may also perform quantization on the scaled residual information. The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block with residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying scale factors to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 may also perform quantization on the scaled residual information. The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block with residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying scale factors to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 may also perform quantization on the scaled residual information. The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block with residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying scale factors to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 may also perform quantization on the scaled residual information. The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block with residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying scale factors to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 may also perform quantization on the scaled residual information. The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform to the residual block to produce a video block with residual transform coefficient values. Wavelet transform, integer transform, subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying scale factors to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 may also perform quantization on the scaled residual information. The quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be adjusted by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of the matrix containing the quantized transform coefficients. The quantized transform coefficients are transferred to the header formatting and CABAC component 231 and encoded in the bitstream. The scaling and inverse transform component 229 applies the inverse operation of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transformation, and / or quantization to reconstruct the residual block in the pixel domain, for example, as a reference block that may later be used as a predicted block for another current block. The motion estimation component 221 and / or the motion compensation component 219 may calculate the reference block by adding the residual block back to the corresponding predicted block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transformation. Such artifacts could otherwise cause inaccurate predictions (and generate further artifacts) when subsequent blocks are predicted. The degree of quantization may be adjusted by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of the matrix containing the quantized transform coefficients. The quantized transform coefficients are transferred to the header formatting and CABAC component 231 and encoded in the bitstream. The scaling and inverse transform component 229 applies the inverse operation of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transformation, and / or quantization to reconstruct the residual block in the pixel domain, for example, as a reference block that may later be used as a predicted block for another current block.

[0080] The motion estimation component 221 and / or the motion compensation component 219 may calculate the reference block by adding the residual block back to the corresponding predicted block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transformation. Such artifacts could otherwise cause inaccurate predictions (and generate further artifacts) when subsequent blocks are predicted. The scaling and inverse transform component 229 applies inverse scaling, transformation, and / or quantization to reconstruct the residual block in the pixel domain, for example, as a reference block that may later be used as a predicted block for another current block. The motion estimation component 221 and / or the motion compensation component 219 may calculate the reference block by adding the residual block back to the corresponding predicted block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transformation. Such artifacts could otherwise cause inaccurate predictions (and generate further artifacts) when subsequent blocks are predicted. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transformation. Such artifacts could otherwise cause inaccurate predictions (and generate further artifacts) when subsequent blocks are predicted. Such artifacts could otherwise cause inaccurate predictions (and generate further artifacts) when subsequent blocks are predicted. Such artifacts could otherwise cause inaccurate predictions (and generate further artifacts) when subsequent blocks are predicted.

[0081] ​ The filter control analysis component 227 and the in-loop filter component 225 apply the filter to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter may then be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Like the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is header formatted as filter control data and transferred to the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, a SAO filter, and an adaptive loop filter. Such a filter, depending on the example, is in the spatial / pixel domain (e.g., the reconstructed pi xel), to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is header formatted as filter control data and transferred to the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, a SAO filter, and an adaptive loop filter. Such a filter, depending on the example, is in the spatial / pixel domain (e.g., the reconstructed pi xel), to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is header formatted as filter control data and transferred to the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, a SAO filter, and an adaptive loop filter. Such a filter, depending on the example, is in the spatial / pixel domain (e.g., the reconstructed pi It may be applied on the pixel block or in the frequency domain.

[0082] When operating as an encoder, the filtered and reconstructed image blocks, the residual blocks, and / or the prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as discussed above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks as part of the output video signal and transfers them to the display. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.

[0083] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission to the decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data, as well as residual data in the form of quantized transform coefficient data, are all encoded in the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original segmented video signal 201. Such information includes the intra prediction mode index table (codeword ​​​​​​​​​​​​​Definition of coding context for various blocks (also referred to as a mapping table) It may also include an indication of the most probable intra prediction mode, an indication of segmentation information, etc. Such data may be encoded by using entropy coding . For example, the information may be encoded by context adaptive variable length coding (CAVLC), CABAC, syntax - based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or by using another entropy coding technique . Following entropy coding, the coded bitstream may be transmitted to another device (e.g., a video decoder), or may be archived for later transmission or retrieval.

[0084] FIG. 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of operation method 100. The encoder 300 partitions an input video signal and provides a partitioned video signal 301 that is substantially the same as the partitioned video signal 201. The partitioned video signal 301 is then compressed and encoded into a bitstream by components of the encoder 300. Specifically, the partitioned video signal 301 is transferred to an intra picture prediction component 317 for intra

[0085] prediction. The intra picture prediction component 317 performs intra prediction on the intra picture prediction component 317 for intra prediction. The intra picture prediction component 317 performs intra The picture estimation component 215 and the intra-picture prediction component 217 may be substantially the same. The segmented video signal 301 is also transferred to the motion compensation component 321 for inter-prediction based on the reference block in the decoded picture buffer component 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction block and the residual block from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to the conversion and quantization component 313 for the conversion and quantization of the residual block . The conversion and quantization component 313 may be substantially the same as the conversion scaling and quantization component 213. The converted and quantized residual block and the corresponding prediction block are transferred to the entropy coding component 331 for coding into the bitstream (along with the relevant control data). The entropy coding component 331 may be substantially the same as the header formatting and CABAC component 231.

[0086] The converted and quantized residual block and / or the corresponding prediction block are also transferred from the conversion and quantization component 313 to the inverse conversion and quantization component 329 for the reconstruction of the reference block used by the motion compensation component 321. The inverse conversion and quantization component 329 may be substantially the same as the scaling and inverse conversion component 229 . The in-loop filter in the in-loop filter component 325 ​​also applies to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component 325 is connected to the filter control analysis component 227 and the loop The in-loop filter component may be substantially similar to the in-loop filter component 225. The in-loop filter component 325 may be any of a number of filters such as those discussed with respect to the in-loop filter component 225. The filtered block may then be filtered by a motion compensation component. Component 321 of the decoded picture buffer for use as a reference block. The decoded picture buffer component 323 stores the decoded picture The component 223 may be substantially similar to the component 223.

[0087] 4 is a block diagram illustrating an example video decoder 400. To implement the decoding functionality of the codec system 200 and / or the steps of the operating method 100 It may be utilized to perform steps 111, 113, 115, and / or 117. The encoder 400 receives the bitstream, for example from the encoder 300, and displays it to the end user. To do this, a reconstructed output video signal is generated based on the bitstream.

[0088] The bitstream is received by the entropy decoding component 433. The Tropey Decoding Component 433 supports CAVLC, CABAC, SBAC, PIPE coding, or other configured to implement an entropy decoding scheme, such as the entropy coding technique of For example, the entropy decoding component 433 may decode the encoded data in the bitstream. To provide context for interpreting additional data encoded as language, the he ader information may be utilized. The decoded information includes any desired information for decoding a video signal, such as general control data, filter control data , section information, motion information, prediction data, and quantized transform coefficients from residual blocks . The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into the residual block . The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329 .

[0089] The reconstructed residual block and / or prediction block are transferred to the intra-picture prediction component 417 for reconstruction into the picture block based on the intra prediction operation . The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically , the intra-picture prediction component 417 utilizes a prediction mode to locate a reference block within the frame and applies the residual block to the result to reconstruct the intra-predicted image block. The reconstructed intra-predicted image block and / or residual block and the corresponding inter-prediction data are transferred through the loop filter component 42 5 to the decoded picture buffer component 423, which may be substantially similar to the decoded picture buffer component 223 and the loop filter component 225, respectively. The loop filter component 425 reconstructs the image block . The decoded picture buffer component 423 and the loop filter component 425 may be substantially similar to the decoded picture buffer component 223 and the loop filter component 225, respectively. The loop filter component 425 reconstructs the image block Filter the hook, residual block, and / or prediction block, and such information is stored in the decoded picture buffer component 423. The reconstructed image block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses the motion vector from the reference block to generate a prediction block and applies the residual block to the result to reconstruct the image block. The resulting reconstructed block may also be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks, which can be reconstructed into frames via segmentation information and such frames may also be placed in the sequence. The sequence is output to the display as the reconstructed output video signal. and can be reconstructed into frames and such frames may also be placed in the sequence. Such frames may also be placed in the sequence. The sequence is output to the display as the reconstructed output video signal.

[0090] FIG. 5 is a schematic diagram showing an exemplary bitstream 500 and a sub-bitstream 501 extracted from the bitstream 500. For example, the bitstream 500 can be generated by the codec system 200 and / or the encoder 300 for decoding by the codec system 200 and / or the decoder 400. As another example, the bitstream 500 can be generated by the encoder in step 109 of method 100 for use by the decoder in step 111. and / or can be generated by the encoder 300. As another example, the bitstream 500 can be generated by the encoder in step 109 of method 100 for use by the decoder in step 111. [[ID=Z36]]and can be generated by the encoder in step 109 of method 100 for use by the decoder in step 111.

[0091] The bitstream 500 includes a sequence parameter set (SPS) 510, a plurality of picture para meter sets (PPSs) 512, a plurality of slice headers 514, picture data 520, and one or more SEI messages 515. The SPS 510 includes sequence data common to all pictures in the video sequence included in the bitstream 500. Such data may include picture size, bit depth, coding tool parameters, bitrate limits, and the like. The PPS 512 includes parameters specific to one or more corresponding pictures. Thus, each picture in the video sequence may refer to one PPS 512. The PPS 512 may indicate coding tools available for tiles in the corresponding picture, quantization parameters, offsets, picture-specific coding tool parameters (e.g., filter control), and so on. The slice header 514 includes parameters specific to one or more corresponding slices 524 in the picture. Thus, each slice 524 in the video sequence may refer to a slice header 514. The slice header 514 may include slice type information, picture order count (POC), reference picture list, prediction weights, tile entry point, deblocking parameters, and the like. In some examples, the slice 524 may be referred to as a tile group. In such a case, the slice header 514 may be referred to as a tile group header. The SEI message 515 is an optional message that includes metadata not required for block decoding, but may be used for related purposes such as indicating the timing of picture output, display settings, loss detection, loss concealment, and the like.

[0092] The picture data 520 includes video data encoded according to inter prediction and / or intra prediction, as well as corresponding transformed and quantized residual data. Such picture data 520 is classified according to the partition used to partition the picture before encoding. For example, the video sequence is split into pictures 521. The picture 521 may be further split into sub-pictures 522, and the sub-picture 522 is split into slices 524. The slice 524 may be further split into tiles and / or CTUs. The CTU is further split into coding blocks based on the coding tree. The coding block can then be encoded / decoded according to the prediction mechanism. For example, the picture 521 may include one or more sub-pictures 522. The sub-picture 522 may include one or more slices 524. The picture 521 refers to the PPS 512, and the slice 524 refers to the slice header 514. The sub-picture 522 may refer to the SPS 510 since it may be stably partitioned across the entire video sequence (also known as a segment). Each slice 524 may include one or more tiles. Each slice 524, and thus the picture 521 and the sub-picture 522, may also include a plurality of CTUs. and corresponding transformed and quantized residual data. Such picture data 520 is classified according to the partition used to partition the picture before encoding. For example, the video sequence is split into pictures 521. The picture 521 may be further split into sub-pictures 522, and the sub-picture 522 is split into slices 524. The slice 524 may be further split into tiles and / or CTUs. The CTU is further split into coding blocks based on the coding tree. The coding block can then be encoded / decoded according to the prediction mechanism. For example, the picture 521 may include one or more sub-pictures 522. The sub-picture 522 may include one or more slices 524. The pic ture 521 refers to the PPS 512, and the slice 524 refers to the slice header 514. The sub-pic ture 522 may refer to the SPS 510 since it may be stably partitioned across the entire video sequence (also known as a segment), so it may refer to the SPS 510. Each slice 524 may include one or more tiles. Each slice 524, and thus the picture 521 and the sub-picture 522, may also include a plurality of CTUs.

[0093] Each picture 521 may include a complete set of visual data associated with the video sequence for the corresponding moment. However, in some applications, it may be desirable to display only a part of the picture 521 in some cases. For example, in a virtual may display a user-selected area of picture 521, which may It creates the feeling of being in the scene depicted in the The region is unknown at the time the bitstream 500 is encoded. Picture 521 is a sub-picture 522, each of which may be viewed by the user. These regions may be separately decoded and displayed based on user input. In other applications, the region of interest may be displayed separately. A television with a picture function can capture a specific area from a video sequence, and therefore A user may wish to display a picture 522 over a picture 521 from an unrelated video sequence. In yet another example, the teleconferencing system may display a general picture of the user currently speaking. A sub-picture 521 of the user not currently speaking and a sub-picture 522 of the user not currently speaking may be displayed. A sub-picture 522 may contain a defined region of picture 521. The sub-picture 522 may be separately decodable from the rest of the picture 521. The temporal motion constrained subpicture does not refer to samples outside the temporal motion constrained subpicture. Since it is coded without reference to the rest of picture 521, there is enough information to perform a complete decoding. Includes information.

[0094] Each slice 524 has a length defined by a CTU in the upper left corner and a CTU in the lower right corner. In some examples, the slices 524 are arranged in a left-to-right and top-to-bottom order. In another example, slice 52 includes a series of tiles and / or CTUs in a raster scan order including: 4 is a rectangular slice. A rectangular slice is a rectangle that is the width of the picture in raster scan order. Instead, rectangular slices are scanned over CTUs and / or tile rows. , and picture 521 and / or The sub-picture 522 may include rectangular and / or square regions. is the smallest unit that can be individually displayed by a decoder. These slices 524 are divided into different sub-pictures to separately depict desired regions of the picture 521. may be assigned to channel 522.

[0095] The decoder may display one or more sub-pictures 523 of a picture 521. The sub-pictures 523 are user-selected subgroups of the sub-pictures 522 or predefined sub-groups. For example, picture 521 is divided into nine sub-pictures 522. However, the decoder may select only a single sub-picture 523 from a group of sub-pictures 522. Subpicture 523 includes slice 525, which is a selection of slice 524. Subpicture 52 is a selected or predefined subgroup. Sub-bitstream 501 is derived from bitstream 500 to allow for three separate displays. The extraction 529 may be performed if the decoder receives only the sub-bitstream 501. In other cases, the entire bitstream 500 may be is sent to the decoder, which extracts the sub-bitstreams 501 for separate decoding. The sub-bitstream 501 is sometimes also referred to generically as the bitstream (529). Note that it may be exposed. The sub-bitstream 501 includes the SPS 510, PPS 512, selected sub-picture 523, as well as the slice header 514 and the SEI message 515 related to the sub-picture 523 and / or slice 525.

[0096] This disclosure signals various data to support efficient coding of sub-picture 522 for selection and display of sub-picture 5 in a decoder. The SPS 510 includes the sub-picture size 531, sub-picture position 532, and sub-picture ID 533 related to the complete set of sub-picture 522. The sub-picture size 531 includes the sub-picture height in luma sample units and the sub- picture width in luma sample units for the corresponding sub-picture 522. The sub-picture position 532 includes the offset distance between the top-left sample of the corresponding sub-picture 522 and the top-left sample of picture 521. The sub-picture position 532 and sub-picture size 531 define the layout of the corresponding sub-picture 522. The sub-picture ID 533 includes data that uniquely identifies the corresponding sub-picture 522. The sub-picture ID 533 may be the raster scan index of sub-picture 522 or other defined values. Thus, the decoder can read the SPS 510 and determine the size, position, and ID of each sub-picture 522. In some video coding systems, data related to sub-picture 522 may be included in the PPS 512 because the sub-picture 522 is partitioned from picture 521. However, to create sub-picture 522 The ID 533 may be the raster scan index of sub-picture 522 or other defined values. Therefore, the decoder can read the SPS 510 and determine the size, position, and ID of each sub-picture 522. In some video coding systems, data related to sub-picture 522 may be included in the PPS 512 because the sub-picture 522 is partitioned from picture 521. However, to create sub-picture 522 for some video coding systems, data related to sub-picture 52 is included in PPS 512 because sub-picture 52 is partitioned from picture 521. However, to create sub-picture 522 for some video coding systems, data related to sub-picture 52 is included in PPS 512 because sub-picture 52 is ​The divisions used include application examples based on ROI, VR application examples, etc., and are used by application examples that depend on the division of the consistent sub-picture 522 throughout the video sequence / segment It may also be used by an application example that depends on the division of the consistent sub-picture 522 throughout the entire video sequence / segment Therefore, the division of the sub-picture 522 generally does not change from picture to picture. Placing the layout information of the sub-picture 522 in the SPS510 ensures that it is not redundantly signaled for each PPS512 (which may be signaled for each picture 521 depending on the case), but is signaled only once for the sequence / segment . Also, instead of depending on the decoder to derive such information, signaling the sub-picture 522 information reduces the probability of errors in case of packet loss and supports additional functions related to extracting the sub-picture 523. Therefore, signaling the layout of the sub-picture 522 in the SPS510 improves the functions of the encoder and / or decoder . Signaling the sub-picture 522 information reduces the probability of errors in case of packet loss and supports additional functions related to extracting the sub-picture 523. Therefore, signaling the layout of the sub-picture 522 in the SPS510 improves the functions of the encoder and / or decoder

[0097] The SPS510 also includes a motion constraint sub-picture flag 534 related to the complete set of sub-pictures 522 . The motion constraint sub-picture flag 534 indicates whether each sub-picture 522 is a temporal motion constraint sub-picture . Therefore, the decoder can read the motion constraint sub-picture flag 534 and determine which of the sub-pictures 522 can be separately extracted and displayed without decoding other sub-pictures 522 . This enables the selected sub-picture 522 to be coded as a temporal motion constraint sub-picture while allowing other sub-pictures 522 to be coded without such constraints for improved coding efficiency ​​​Enable it to be dinged.

[0098] Sub-picture ID 533 is also included in slice header 514. Each slice header 514 contains data related to the corresponding set of slices 524. Therefore, slice header 514 contains only the sub-picture ID 533 corresponding to the slice 524 associated with slice header 514. Therefore, the decoder can receive slice 524, obtain the sub-picture ID 533 from slice header 514, and determine which sub-picture 522 contains slice 524 . The decoder can also use the sub-picture ID 533 from slice header 514 to correlate with the relevant data in SPS 510. Therefore, the decoder can determine how to position sub-pictures 522 / 523 and slices 524 / 525 by reading SPS 510 and the relevant slice header 514. This enables sub-pictures 523 and slices 525 to be decoded even if some sub-pictures 522 are lost during transmission or are deliberately omitted for coding efficiency. The SEI message 515 may also include a sub-picture level 535. The sub-picture level 535 indicates the hardware resources required to decode the corresponding sub-picture 522. In this way, each sub-picture 522 can be coded independently of other sub-pictures 522. This ensures that each sub-picture 522 can be allocated the correct amount of hardware resources at the decoder. Without such a sub-picture level 535 it would not be possible. even if some sub-pictures 522 are lost during transmission or are deliberately omitted for coding efficiency. .

[0099] The SEI message 515 may also include a sub-picture level 535. The sub-picture level 535 indicates the hardware resources required to decode the corresponding sub-picture 522. In this way, each sub-picture 522 can be coded independently of other sub-pictures 522. This ensures that each sub-picture 522 can be allocated the correct amount of hardware resources at the decoder. Without such a sub-picture level 535 it would not be possible. to ensure that each sub-picture 522 can be allocated the correct amount of hardware resources at the decoder. Without such a sub-picture level 535 it would not be possible. ​Each sub-picture 522 is allocated enough resources to decode the most complex sub-picture 522. Therefore, the subpicture level 535 is the hard level at which the subpicture 522 changes. Decoders may over-utilize hardware resources when combined with hardware resource requirements. Prevent allocation.

[0100] FIG. 6 is a schematic diagram illustrating an example picture 600 partitioned into sub-pictures 622. For example, picture 600 may be encoded by, for example, codec system 200, encoder 300, and / or is encoded in a bitstream 500 by the decoder 400, 0. Furthermore, picture 600 supports encoding and decoding according to method 100. The sub-bitstreams 501 may be partitioned and / or included to accommodate the above-mentioned needs.

[0101] Picture 600 may be substantially similar to picture 521. Furthermore, picture 600 may be The image may be partitioned into sub-pictures 622, which are substantially the same as sub-picture 522. Similarly, each subpicture 622 includes a subpicture size 631. 631 may be included in the bitstream 500 as the subpicture size 531. The subpicture size 631 includes a subpicture width 631a and a subpicture height 631b. Width 631a is the width of the corresponding subpicture 622 in units of luma samples. The height 631b is the height of the corresponding subpicture 622 in units of luma samples. 22 each includes a subpicture ID 633, and the subpicture ID 633 is a bit It may be included in the bitstream 500. The sub-picture ID 633 may be any value that uniquely identifies each sub-picture 622. In the example shown, the sub-picture ID 633 is the index of sub-picture 6 22. Each sub-picture 622 includes a position 632, and the position 632 may be included in the bitstream 500 as the sub-picture position 532. The position 632 is represented as the offset between the top-left sample of the corresponding sub-picture 622 and the top-left sample 642 of the picture 600.

[0102] Also as shown, some sub-pictures 622 may be time motion constraint sub-pictures 634, while other sub-pictures 622 may not. In the example shown, the sub-picture 622 with the sub-picture ID 633 of 5 is a time motion constraint sub-picture 634. This indicates that the sub-picture 622 identified as 5 can be decoded without referring to any other sub-picture 622, and thus can be extracted and decoded separately without considering the data from other sub-pictures 622. The indication of which sub-picture 622 is the time motion constraint sub-picture 634 can be signaled in the motion constraint sub-picture flag 534 in the bitstream 500.

[0103] As shown, the sub-pictures 622 may be restricted to enclose the picture 600 without gaps or overlaps. A gap is an area of the picture 600 that is not included in any sub-picture 622. An overlap is an area of the picture 600 that is included in more than one sub-picture 622. In the example shown in FIG. 6, the sub-pictures 622 enclose the picture 600 so as to prevent both gaps and overlaps. ​​​​It is divided from. Due to the gap, samples of the picture 600 remain outside the sub - picture 622. Due to duplication, related slices are included in multiple sub - pictures 622. Thus, the gap and duplication cause the samples to be affected by different handling when the sub - pictures 622 are coded differently. This may be the case. If this is allowed in the encoder, the decoder must support such a coding method even when the decoding method is rarely used. By not allowing gaps and duplications in the sub - pictures 622, the complexity of the decoder can be reduced because it is not necessary for the decoder to consider potential gaps and duplications when determining the sub - picture size 631 and position 632. Furthermore, not allowing gaps and duplications in the sub - pictures 622 reduces the complexity of the RDO process in the encoder. This is because the encoder can omit considering cases of gaps and duplications when selecting the coding for the video sequence. Thus, avoiding gaps and duplications may reduce the use of memory resources and / or processing resources in the encoder and decoder.

[0104] FIG. 7 is a schematic diagram showing an exemplary mechanism 700 for associating a slice 724 with the layout of sub - pictures 722. For example, the mechanism 700 may be applied to the picture 600. Further, the mechanism 700 may be applied based on data in the bitstream 500, for example, by the codec system 200, the encoder 300, and / or the decoder 400. Further, the mechanism 700 may be utilized to support the encoding and decoding according to the method 100.

[0105] ​​​​ The mechanism 700 can be applied to slices 724 in sub-pictures 722, such as slices 524 / 525 and sub-pictures 522 / 523 respectively. In the example shown, the sub-picture 722 includes a first slice 724a, a second slice 724b, and a third slice 724c. Each of the slice headers of slice 724 includes the sub-picture ID 733 of the sub-picture 722. The decoder can match the sub-picture ID 733 from the slice header with the sub-picture ID 733 in the SPS. The decoder can then determine the position 732 and size of the sub-picture 722 from the SPS based on the sub-picture ID 733. Using the position 732, the sub-picture 722 can be arranged relative to the top-left sample at the top-left corner 742 of the picture. The size can be used to set the height and width of the sub-picture 722 relative to the position 732. Next, the slice 724 can be included in the sub-picture 722. Thus, the slice 724 can be positioned correctly within the correct sub-picture 722 based on the sub-picture ID 733 without referring to other sub-pictures. This helps error correction because other lost sub-pictures do not change the decoding of the sub-picture 722. This also helps application examples that extract only the sub-picture 722 and avoids the transmission of other sub-pictures. Thus, the sub-picture ID 733 helps improve functionality and / or coding efficiency, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.

[0106] FIG. 8 is a schematic diagram showing another exemplary picture 800 that is partitioned into sub-pictures 822 . Picture 800 may be substantially similar to picture 600. Additionally, picture 800 may be encoded in bitstream 500 and decoded from bitstream 500 by, for example, codec system 200, encoder 300, and / or decoder 400 . Furthermore, picture 800 may be partitioned in and / or included in sub-bitstream 501 to support encoding and decoding according to method 100 and / or mechanism 700 .

[0107] Picture 800 includes sub-picture 822, and sub-picture 822 may be substantially similar to sub-pictures 522, 523, 622 , and / or 722. Sub-picture 822 is divided into a plurality of CTUs 825 . CTU 825 is a basic coding unit in a standardized video coding system. CTU 825 is further divided into coding blocks by a coding tree, and the coding blocks are coded according to inter prediction or intra prediction . As shown, some sub-pictures 822a are constrained to include a sub-picture width and a sub-picture height that are multiples of the size of CTU 825. In the example shown, sub-picture 822a has a height of 6 CTUs 825 and a width of 5 CTUs 825. This constraint is removed for sub-picture 822b located at the right boundary 801 of the picture and sub-picture 822c located at the lower boundary 802 of the picture . In the example shown, sub-picture 822b has a width of between 5 and 6 CTUs 825. Thus, sub-picture 822b has 5 complete CTUs 825 and one partial CTU 825. Sub-picture 822c has a height of between 5 and 6 CTUs 825 and a width of 5 CTUs 825​​​​​​​ has a width of CTU825 and one incomplete CTU825. However, the sub-picture 822b that is not located at the bottom boundary 802 of the picture is still restricted to maintain a sub-picture height that is a multiple of the size of CTU825. In the example shown, the sub-picture 822c has a height between 6 and 7 CTU825s. Therefore, the sub-picture 822c has a height of 6 complete CTUs 825s and 1 incomplete CTU825. However, the sub-picture 822c that is not located at the right boundary 801 of the picture is still restricted to maintain a sub-picture width that is a multiple of the size of CTU825. Note that the right boundary 801 of the picture and the bottom boundary 802 of the picture may also be referred to as the picture right boundary and the picture bottom boundary, respectively. It should also be noted that the size of CTU825 is a value determined by the user. The size of CTU825 can be any value between the size of the smallest CTU825 and the size of the largest CTU825. For example, the size of the smallest CTU825 may have a height of 16 luma samples and a width of 16 luma samples. Furthermore, the size of the largest CTU825 may have a height of 128 luma samples and a width of 128 luma samples. As described above, some video systems may restrict the sub-picture 822 to include a height and width that are multiples of the size of CTU825. This may prevent the sub-picture 822 from operating correctly with a number of picture layouts, such as the picture 800 that includes an overall width or height that is not a multiple of the size of CTU825. The lower sub-picture 822c

[0108] As mentioned above, some video systems may limit the sub-picture 822 to include a height and width that are multiples of the size of CTU825. This may prevent the sub-picture 822 from operating correctly with a number of picture layouts, such as the picture 800 that includes an overall width or height that is not a multiple of the size of CTU825. The lower sub-picture 822c ​​​​​​​and the right sub-picture 822b each include a height and a width that are not multiples of the size of the CTU 825 By allowing this, the sub-picture 822 may be used with any picture 800 without causing decoding errors. This brings about an improvement in the functions of the encoder and decoder. Furthermore, due to the improved function, the encoder can code pictures more efficiently, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.

[0109] As described herein, the present disclosure describes the design of sub-picture-based picture partitioning in video coding. A sub-picture is a rectangular area within a picture that can be independently decoded using a decoding process similar to that used for the picture. The present disclosure relates to the signaling of sub-pictures in a coded video sequence and / or bitstream, as well as the process for sub-picture extraction. The description of the techniques is based on VVC by ITU-T and ISO / IEC's JVET. However, the techniques are also applicable to other video codec specifications. The following are exemplary embodiments described herein. Such embodiments may be applied individually or in combination. The description of the techniques is based on VVC by ITU-T and ISO / IEC's JVET. However, the techniques are also applicable to other video codec specifications. The following are exemplary embodiments described herein. Such embodiments may be applied individually or in combination.

[0110] Information regarding sub-pictures that may be present in a coding video sequence (CVS) may be signaled in a sequence-level parameter set such as SPS. Such signaling may include the following information. Sub-pictures present in each picture of the CVS ​​​​​​​​​​​​​The number of tiles may be signaled in the SPS. In the context of the SPS or CVS, all sub - pictures at the same position for all access units (AUs) may be collectively called a sub - picture sequence. A loop for further specifying the information describing the nature of each sub - picture may also be included in the SPS. This information may include sub - picture identification information, the position of the sub - picture (e.g., the offset distance between the top - left luma sample of the sub - picture and the top - left luma sample of the picture), and the size of the sub - picture. In addition, the SPS may signal whether each sub - picture is a motion - constrained sub - picture (including the function of MCTS). Profile, layer, and level information for each sub - picture may also be signaled in the decoder or may be derivable. Such information may be used to determine the profile, layer, and level information for the bitstream created by extracting the sub - picture from the original bitstream. The profile and layer of each sub - picture may be derived as being the same as the profile and layer of the entire bitstream. The level of each sub - picture may be explicitly signaled. Such signaling may be present in the loop included in the SPS. Sequence - level Hypothetical Reference Decoder (HRD) parameters may be signaled in the Video Usability Information (VUI) section of the SPS for each sub - picture (or equivalently, for each sub - picture sequence). When a picture is not divided into two or more sub - pictures, the characteristics of the sub - picture (e.g.,

[0111] ​​​​​​​​​​​If so, (such as position and size) except for the sub-picture ID, do not exist in the bitstream / may not be signaled in the bitstream. When the sub-picture in the CVS is extracted, each access unit in the new bitstream may not contain a sub-picture. In this case, the picture in each AU in the new bitstream is not divided into multiple sub-pictures. Therefore, there is no need to signal any sub-picture characteristics such as position and size in the SPS, because such information can be derived from the picture properties. However, even so, the sub-picture identification information may still be signaled, because the ID may be referenced by the VCL NAL unit / tile group included in the extracted sub-picture. This may enable the sub-picture ID to remain the same when extracting the sub-picture. The position of the sub-picture in the picture (x offset and y offset) can be signaled in units of luma samples. The position represents the distance between the top-left luma sample of the sub-picture and the top-left luma sample of the picture. Alternatively, the position of the sub-picture in the picture can be signaled in units of the minimum coding luma block size (MinCbSizeY). Alternatively, the unit of the sub-picture position offset may be explicitly indicated by the syntax element in the parameter set. The unit may be CtbSizeY, MinCbSizeY, luma sample, or other values. When the sub-picture in the CVS is extracted, each access unit in the new bitstream may not contain a sub-picture. In this case, the picture in each AU in the new bitstream is not divided into multiple sub-pictures. Therefore, there is no need to signal any sub-picture characteristics such as position and size in the SPS, because such information can be derived from the picture properties. However, even so, the sub-picture identification information may still be signaled, because the ID may be referenced by the VCL NAL unit / tile group included in the extracted sub-picture. This may enable the sub-picture ID to remain the same when extracting the sub-picture. The size of the sub-picture (sub-picture width and sub-picture height) is in units of luma samples. The position of the sub-picture in the picture (x offset and y offset) can be signaled in units of luma samples. The position represents the distance between the top-left luma sample of the sub-picture and the top-left luma sample of the picture.

[0112] The position of the sub-picture in the picture (x offset and y offset) can be signaled in units of luma samples. The position represents the distance between the top-left luma sample of the sub-picture and the top-left luma sample of the picture. Alternatively, the position of the sub-picture in the picture can be signaled in units of the minimum coding luma block size (MinCbSizeY). Alternatively, the unit of the sub-picture position offset may be explicitly indicated by the syntax element in the parameter set. The unit may be CtbSizeY, MinCbSizeY, luma sample, or other values. The size of the sub-picture (sub-picture width and sub-picture height) is in units of luma samples. The position of the sub-picture in the picture (x offset and y offset) can be signaled in units of luma samples.

[0113] The size of the sub-picture (sub-picture width and sub-picture height) is in units of luma samples. can be signaled in bits. Alternatively, the size of the sub-picture can be signaled in units of the minimum coding luma block size (MinCbSizeY). Alternatively, the unit of the sub-picture size can be explicitly indicated by a syntax element in the parameter set can be. The unit can be CtbSizeY, MinCbSizeY, luma samples, or other values. When the right boundary of the sub-picture does not coincide with the right boundary of the picture, the width of the sub-picture may be required to be an integer multiple of the luma CTU size (CtbSizeY). Similarly, when the lower boundary of the sub-picture does not coincide with the lower boundary of the picture, the height of the sub-picture may be required to be an integer multiple of the luma CTU size (CtbSizeY). When the width of the sub-picture is not an integer multiple of the luma CTU size, the sub-picture may be required to be located at the rightmost position in the picture. Similarly, when the height of the sub-picture is not an integer multiple of the luma CTU size the sub-picture may be required to be located at the lowest position in the picture. In some cases, the width of the sub-picture can be signaled in units of the luma CTU size, but the width of the sub-picture is not an integer multiple of the luma CTU size. In this case, the actual width in luma sample units can be derived based on the offset position of the sub-picture. The width of the sub-picture can be derived based on the luma CTU size, and the height of the picture can be derived based on the luma sample units. Similarly, the height of the sub-picture may be signaled in units of the luma CTU size, but the height of the sub-picture is not an integer multiple of the luma CTU size. In such a case, the actual height in luma sample units is based on the offset position of the sub-picture can be. In some cases, the width of the sub-picture can be signaled in units of the luma CTU size, but the width of the sub-picture is not an integer multiple of the luma CTU size. In this case, the actual width in luma sample units can be derived based on the offset position of the sub-picture. The width of the sub-picture can be derived based on the luma CTU size, and the height of the picture can be derived based on the luma sample units. Similarly, the height of the sub-picture may be signaled in units of the luma CTU size, but the height of the sub-picture is not an integer multiple of the luma CTU size. In such a case, the actual height in luma sample units is based on the offset position of the sub-picture size, but the height of the sub-picture is not an integer multiple of the luma CTU size. In such a case, the actual height in luma sample units is based on the offset position of the sub-picture position. Deriving the sub-picture height based on the luma CTU size and the picture height can be derived based on the luma samples.

[0114] For any given subpicture, the subpicture ID is different from the subpicture index. The subpicture index may be used as a signature in a subpicture loop in an SPS. The subpicture ID may be the index of the subpicture to be nulled. The index of the subpicture in the picture's subpicture raster scan order. The value of the subpicture ID of each subpicture is the same as the subpicture index. When a subpicture is generated, the subpicture ID may be signaled or derived. When the picture ID differs from the subpicture index, the subpicture ID must be explicitly signaled. The number of bits for signaling the subpicture ID depends on the subpicture feature may be signaled in the same parameter set (e.g., in the SPS) that contains Some values for the subpicture ID may be reserved for certain purposes. For example, a tile group header specifies which subpicture contains the tile group. Prevents the accidental inclusion of anti-emulation code when including subpicture IDs for To ensure that the first few bits of the tile group header are not all zeros, Additionally, the value 0 is reserved for subpictures and may not be used. In the optional case where the image does not encompass the entire area of the picture without gaps and overlaps, A value (e.g., value 1) is reserved for tile groups that are not part of any subpicture. This may be done. Alternatively, the sub-picture IDs of the remaining areas may be explicitly signaled The number of bits for signaling the sub-picture ID may be constrained as follows The value range must be sufficient to uniquely identify all sub-pictures in the picture, including the reserved values of the sub-picture ID For example, the minimum number of bits for the sub-picture ID can be the value of Ceil(Log2(number of sub-pictures in the picture + number of reserved sub-picture IDs))

[0115] It may be constrained that the union of sub-pictures must cover the entire picture without gaps and overlaps When this constraint applies, for each sub-picture, there may be a flag to specify whether the sub-picture is a motion-constrained sub-picture which indicates that the sub-picture can be extracted. Alternatively, the union of sub-pictures may not cover the entire picture, but overlaps may not be allowed

[0116] To assist the sub-picture extraction process without requiring the extractor to analyze the rest of the NAL unit bits, the sub-picture ID may be present immediately after the NAL unit header For VCL NAL units, the sub-picture ID may be present in the first bit of the tile group header For non-VCL NAL units, the following may apply. In the SPS, the sub-picture ID does not need to be present immediately after the NAL unit header. In the PPS, if all tile groups of the same picture are constrained to refer to the same PPS, the sub-picture ID does not need to be present immediately after the NAL unit header For all tile groups of the same picture, if they are constrained to refer to the same PPS, the sub-picture ID does not need to be present immediately after the NAL unit header ​​​​​​​When different tile groups are allowed to refer to different PPSs, the sub-picture ID may be present in the first bit (e.g., immediately following the NAL unit header). In this case, it may be allowed for any tile group of a picture to share the same PPS. Alternatively, when tile groups of the same picture are allowed to refer to different PPSs and different tile groups of the same picture are also allowed to share the same PPS, the sub-picture ID may not be present in the PPS syntax. Alternatively, when tile groups of the same picture are allowed to refer to different PPSs and different tile groups of the same picture are also allowed to share the same PPS, a list of sub-picture IDs may be present in the PPS syntax. This list indicates the sub-pictures to which the PPS applies. For other non-VCL NAL units, when non-VCL units (e.g., access unit delimiter, end of sequence, end of bitstream, etc.) are applied at the picture level or above, then, the sub-picture ID may not be present immediately following the NAL unit header. Otherwise, the sub-picture ID may be present immediately following the NAL unit header. For other non-VCL NAL units, when non-VCL units (e.g., access unit delimiter, end of sequence, end of bitstream, etc.) are applied at the picture level or above, then, the sub-picture ID may not be present immediately following the NAL unit header. Otherwise, the sub-picture ID may be present immediately following the NAL unit header. When different tile groups are allowed to refer to different PPSs, the sub-picture ID may be present in the first bit (e.g., immediately following the NAL unit header). In this case, it may be allowed for any tile group of a picture to share the same PPS. Alternatively, when tile groups of the same picture are allowed to refer to different PPSs and different tile groups of the same picture are also allowed to share the same PPS, the sub-picture ID may not be present in the PPS syntax. When different tile groups are allowed to refer to different PPSs, the sub-picture ID may be present in the first bit (e.g., immediately following the NAL unit header). In this case, it may be allowed for any tile group of a picture to share the same PPS. Alternatively, when tile groups of the same picture are allowed to refer to different PPSs and different tile groups of the same picture are also allowed to share the same PPS, the sub-picture ID may not be present in the PPS syntax and different tile groups of the same picture are also allowed to share the same PPS, a list of sub-picture IDs may be present in the PPS syntax. Alternatively, when tile groups of the same picture are allowed to refer to different PPSs and different tile groups of the same picture are also allowed to share the same PPS, a list of sub-picture IDs may be present in the PPS syntax. This list indicates the sub-pictures to which the PPS applies. For other non-VCL NAL units, when non-VCL units (e.g., access unit delimiter, end of sequence, end of bitstream, etc.) are applied at the picture level or above, then, the sub-picture ID may not be present immediately following the NAL unit header. Otherwise, the sub-picture ID may be present immediately following the NAL unit header. This list indicates the sub-pictures to which the PPS applies. For other non-VCL NAL units, when non-VCL units (e.g., access unit delimiter, end of sequence, end of bitstream, etc.) are applied at the picture level or above, then, the sub-picture ID may not be present immediately following the NAL unit header. Otherwise, the sub-picture ID may be present immediately following the NAL unit header. For other non-VCL NAL units, when non-VCL units (e.g., access unit delimiter, end of sequence, end of bitstream, etc.) are applied at the picture level or above, then, the sub-picture ID may not be present immediately following the NAL unit header. Otherwise, the sub-picture ID may be present immediately following the NAL unit header. For other non-VCL NAL units, when non-VCL units (e.g., access unit delimiter, end of sequence, end of bitstream, etc.) are applied at the picture level or above, then, the sub-picture ID may not be present immediately following the NAL unit header. Otherwise, the sub-picture ID may be present immediately following the NAL unit header. For other non-VCL NAL units, when non-VCL units (e.g., access unit delimiter, end of sequence, end of bitstream, etc.) are applied at the picture level or above, then, the sub-picture ID may not be present immediately following the NAL unit header. Otherwise, the sub-picture ID may be present immediately following the NAL unit header. For other non-VCL NAL units, when non-VCL units (e.g., access unit delimiter, end of sequence, end of bitstream, etc.) are applied at the picture level or above, then, the sub-picture ID may not be present immediately following the NAL unit header. Otherwise, the sub-picture ID may be present immediately following the NAL unit header.

[0117] Using the above SPS signaling, the tile partitioning within an individual sub-picture may be signaled in the PPS. Tile groups within the same picture may be allowed to refer to different PPSs. In this case, the tile grouping may be only within each sub-picture. The concept of tile grouping is the partitioning of a sub-picture into tiles. Using the above SPS signaling, the tile partitioning within an individual sub-picture may be signaled in the PPS. Tile groups within the same picture may be allowed to refer to different PPSs. In this case, the tile grouping may be only within each sub-picture. The concept of tile grouping is the partitioning of a sub-picture into tiles. In this case, the tile grouping may be only within each sub-picture. The concept of tile grouping is the partitioning of a sub-picture into tiles.

[0118] Alternatively, a parameter set for describing tile partitioning within each subpicture is defined. Such a parameter set may be referred to as a subpicture parameter set (SPPS). The SPPS refers to the SPS. A syntax element that refers to the SPS ID is present in the SPPS. The SPPS may include a subpicture ID. For the purpose of subpicture extraction, a syntax element that refers to the subpicture ID is the first syntax element in the SPPS. The SPPS includes a tile structure (e.g., number of columns, number of rows, uniform tile separation, etc.). The SPPS may include a flag for indicating whether the loop filter is effective across the subpicture boundary to which it is relevant. Alternatively, the subpicture characteristics for each subpicture may be signaled in the SPPS instead of in the SPS. The tile partitioning within each individual subpicture may itself be signaled in the PPS. Tile groups within the same picture are allowed to refer to different PPSs. When the SPPS is activated, the SPPS continues between consecutive sequences of AUs in decoding order. However, the SPPS may be deactivated / activated in an AU that is not the first in the CVS. At any instant during the decoding process of a single-layer bitstream with multiple subpictures in some AU, multiple SPPSs may be active. The SPPS may be shared by different subpictures of an AU. Alternatively, the SPPS and the PPS may be integrated into one parameter set. In such a case, it is not necessary for all tile groups of the same picture to refer to Constraints may be applied such that the same parameter set may be referenced.

[0119] The number of bits used to signal the subpicture ID may be signaled in the NAL unit header. When present in the NAL unit header, such information may assist the subpicture extraction process when parsing the subpicture ID value that is in the beginning of the NAL unit payload (e.g., the first few bits right after the NAL unit header). For such signaling, a part of the spare bits in the NAL unit header (e.g., 7 spare bits) may be used to avoid increasing the length of the NAL unit header. The number of such signaling bits may include the value of sub-picture-ID-bit-len. For example, 4 out of the 7 spare bits of the VVC NAL unit header may be used for this purpose.

[0120] When decoding a subpicture, the position of each coding tree block (e.g., xCtb and yCtb) may be adjusted to the actual luma sample position in the picture instead of the luma sample position in the subpicture. This way, since the coding tree block is decoded with reference to the picture instead of the subpicture, extraction of the same subpicture at the same position from each reference picture can be avoided. To adjust the position of the coding tree block, the variables SubpictureXOffset and SubpictureYOffset may be derived based on the position of the subpicture (subpic_x_offset and subpic_y_offset). The value of the variables may be ​​​​​​​​​​​​​​The values of the coordinates of the luma sample positions x and y of each coding tree block in the picture may be added thereto respectively.

[0121] The sub-picture extraction process can be defined as follows. The input to the process is the target sub-picture to be extracted. This can be in the form of a sub-picture ID or a sub-picture position. When the input is the position of the sub-picture, the relevant sub-picture ID can be resolved by analyzing the sub-picture information in the SPS. For non-VCL NAL units, the following applies. The syntax elements in the SPS regarding the picture size and level may be updated together with the size and level information of the sub-picture. The following non-VCL NAL units, namely, PPS, Access Unit Delimiter (AUD), End of Sequence (EOS), End of Bitstream (EOB), and any other non-VCL NAL units applicable to the picture level or above, remain unchanged. The remaining non-VCL NAL units with sub-picture IDs equal to the target sub-picture ID may be removed. The VCL NAL units with sub-picture IDs not equal to the target sub-picture ID may also be removed. [[ID=2,6]]The sub-picture extraction process can be defined as follows. The input to the process is the target sub-picture to be extracted. This can be in the form of a sub-picture ID or a sub-picture position. When the input is the position of the sub-picture, the relevant sub-picture ID can be resolved by analyzing the sub-picture information in the SPS. For non-VCL NAL units, the following applies. The syntax elements in the SPS regarding the picture size and level may be updated together with the size and level information of the sub-picture. The following non-VCL NAL units, namely, PPS, Access Unit Delimiter (AUD), End of Sequence (EOS), End of Bitstream (EOB), and any other non-VCL NAL units applicable to the picture level or above, remain unchanged. The remaining non-VCL NAL units with sub-picture IDs equal to the target sub-picture ID may be removed. The VCL NAL units with sub-picture IDs not equal to the target sub-picture ID may also be removed.

[0122] The sequence level sub-picture that nests the SEI message may be used for nesting the AU level SEI message or the sub-picture level SEI message for a set of sub-pictures. This may include the buffering period, picture timing, and non-HRD SEI messages. The syntax of this sub-picture that nests the SEI message ​​​​​​​​​​​​​The syntax and semantics can be as follows. In a system operation in an unoriented media format (OMA F) environment, etc., a set of sub-picture sequences including a viewport may be requested and decoded by an OMAF player. Therefore, the sequence level SEI message is used to carry information about a set of sub-picture sequences collectively encompassing rectangular picture regions. This information can be used by the system and this information indicates the required decoding capabilities as well as the bitrate of the set of sub-picture sequences This information indicates the level of the bitstream containing only the set of sub-picture sequences This information also indicates the bitrate of the bitstream containing only the set of sub-picture sequences Optionally, the sub-bitstream extraction process may be specified for the set of sub-picture sequences. The advantage of doing this is that the bitstream containing only the set of sub-picture sequences can also be made compliant The disadvantage is that considering the possibility of different viewport sizes, in addition to the already numerous individual sub-picture sequences, many such sets can exist.

[0123] In one exemplary embodiment, one or more of the disclosed examples may be implemented as follows. A sub-picture may be defined as a rectangular region of one or more tile groups within a picture The allowed splitting process may be defined as follows. The input to this process is the splitting mode btSplit, the coding block width cbWidth, the coding block height cbHeight, the relative considered coding block for the top-left luma sample of the picture The position (x0, y0) of the top-left luma sample of the coding block, the multi-type tree depth mttDepth, the maximum multi-type tree depth maxMttDepth with offset, the maximum binary tree size maxBtSize, and the partition index partIdx. The output of this process is the variable allowBtSplit.

Table 1

[0124] The variables parallelTtSplit and cbSize are derived as defined above. The variable allowBt Spit is derived as follows. If any of the following conditions, namely, cbSize is less than or equal to MinBtSizeY, cb Width is greater than maxBtSize, cbHeight is greater than maxBtSize, and mttDepth is greater than or equal to maxMt tDepth, is true, then allowBtSplit is set equal to FALSE. Otherwise, if all of the following conditions, namely, btSplit is equal to SPLIT_BT_VER and y0 + cbHeight is greater than SupPicBottomBorderInPic, are true then allowBtSplit is set equal to FALSE. Otherwise, if all of the following conditions, namely btSplit is equal to SPLIT_BT_HOR, x0 + cbWidth is greater than SupPicRightBorderInPic, and y0 + cbHeight is less than or equal to SubPicBottomBorderInPic are true then allowBtSplit is set equal to FALSE. Otherwise, if all of the following conditions, namely ​That is, when all of the conditions that mttDepth is greater than 0, partIdx is equal to 1, and MttSplitMode[x0][y0][mttDepth - 1] is equal to parallelTtSplit are true, allowBtSplit is set equal to FALSE. Otherwise, when all of the following conditions, that is, btSplit is equal to SPLIT_BT_VER, cbWidth is less than or equal to MaxTbSizeY, and cbHeight is greater than MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. Otherwise, when all of the following conditions, that is, btSplit is equal to SPLIT_BT_HOR, cbWidth is greater than MaxTbSizeY, and cbHeight is less than or equal to MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. In other cases, allowBtSplit is set equal to TRUE. If all of the conditions are true, allowBtSplit is set equal to FALSE. Otherwise, when all of the following conditions, that is, btSplit is equal to SPLIT_BT_VER, cbWidth is less than or equal to MaxTbSizeY, and cbHeight is greater than MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. Otherwise, when all of the following conditions, that is, btSplit is equal to SPLIT_BT_HOR, cbWidth is greater than MaxTbSizeY, and cbHeight is less than or equal to MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. In other cases, allowBtSplit is set equal to TRUE. FALSE. Otherwise, when all of the following conditions, that is, btSplit is equal to SPLIT_BT_VER, cbWidth is less than or equal to MaxTbSizeY, and cbHeight is greater than MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. Otherwise, when all of the following conditions, that is, btSplit is equal to SPLIT_BT_HOR, cbWidth is greater than MaxTbSizeY, and cbHeight is less than or equal to MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. In other cases, allowBtSplit is set equal to TRUE. equal, cbWidth is less than or equal to MaxTbSizeY, and cbHeight is greater than MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. Otherwise, when all of the following conditions, that is, btSplit is equal to SPLIT_BT_HOR, cbWidth is greater than MaxTbSizeY, and cbHeight is less than or equal to MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. In other cases, allowBtSplit is set equal to TRUE. If all of the conditions are true, allowBtSplit is set equal to FALSE. Otherwise, when all of the following conditions, that is, btSplit is equal to SPLIT_BT_HOR, cbWidth is greater than MaxTbSizeY, and cbHeight is less than or equal to MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. In other cases, allowBtSplit is set equal to TRUE. Otherwise, when all of the following conditions, that is, btSplit is equal to SPLIT_BT_HOR, cbWidth is greater than MaxTbSizeY, and cbHeight is less than or equal to MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. In other cases, allowBtSplit is set equal to TRUE. greater, and cbHeight is less than or equal to MaxTbSizeY, are true, allowBtSplit is set equal to FALSE. In other cases, allowBtSplit is set equal to TRUE. If all of the conditions are true, allowBtSplit is set equal to FALSE. In other cases, allowBtSplit is set equal to TRUE. set.

[0125] The allowed three - split process may be defined as follows. The input to this process is the three - split mode ttSplit, the coding block width cbWidth, the coding block height cbHeight, the position (x0, y0) of the top - left luma sample of the coding block relative to the top - left luma sample of the picture, the multi - type tree depth mttDepth, the maximum multi - type tree depth maxMttDepth with offset, and the maximum binary tree size maxTtSize. The output of this process is the variable allowTtSplit. That is, the input to this process is the three - split mode ttSplit, the coding block width cbWidth, the coding block height cbHeight, the position (x0, y0) of the top - left luma sample of the coding block relative to the top - left luma sample of the picture, the multi - type tree depth mttDepth, the maximum multi - type tree depth maxMttDepth with offset, and the maximum binary tree size maxTtSize. The output of this process is the variable allowTtSplit. Height, the position (x0, y0) of the top - left luma sample of the coding block considered relative to the top - left luma sample of the picture, the multi - type tree depth mttDepth, the maximum multi - type tree depth maxMttDepth with offset, and the maximum binary tree size maxTtSize. The output of this process is the variable allowTtSplit. of the top - left luma sample of the coding block relative to the top - left luma sample of the picture, the multi - type tree depth mttDepth, the maximum multi - type tree depth maxMttDepth with offset, and the maximum binary tree size maxTtSize. The output of this process is the variable allowTtSplit. depth maxMttDepth, and the maximum binary tree size maxTtSize. The output of this process is the variable allowTtSplit. The output of this process is the variable allowTtSplit. [Table 2]

[0126] The variable cbSize is derived as defined above. The variable allowTtSplit is derived as follows follows. If one or more of the following conditions, namely, cbSize is less than or equal to 2*MinTtSizeY, cbWidth is greater than Min(MaxTbSize Y, maxTtSize), cbHeight is greater than Min(MaxTbSizeY, maxTtSize), mttDepth is greater than or equal to maxMttDepth, x0 + cbWidth is greater than SupPicRightBoderInPic, and y0 + c bHeight is greater than SubPicBottomBorderInPic, is true, then allowTtSplit is set equal to FALSE. Otherwise, allowTtSplit is set equal to TRUE .

[0127] The syntax and semantics of the sequence parameter set RBSP are as follows .

Table 3

[0128] pic_width_in_luma_samples specifies the width of each decoded picture in units of luma samples . pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY . pic_height_in_luma_samples specifies the height of each decoded picture in units of luma samples . pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY . The value obtained by adding 1 to num_subpicture_minus1 is ​, specifying the number of sub - pictures segmented in the coded picture, belongs to the coded video sequence. The value obtained by adding 1 to subpic_id_len_minus1 is used to represent the number of bits for the syntax element subpic_id[i] in SPS, spps_subpic_id in SPPS referring to SPS, and tile_group_subpic_id in the tile group header referring to SPS. The value of subpic_id_len_minus1 shall be in the range from Ceil(Log2(num_subpic_minus1 + 2)) (including both ends) to 8. subpic_id[i] specifies the sub - picture ID of the i - th sub - picture of the picture referring to SPS. The length of subpic_id[i] is subpic_id_len_minus1 + 1 bits. The value of subpic_id[i] shall be greater than 0. subpic_level_idc[i] indicates the level according to which the CVS resulting from the extraction of the i - th sub - picture meets the specified resource requirements. The bit - stream shall not contain values of subpic_level_idc[i] other than those specified. Other values of subpic_level_idc[i] are reserved. If not present, the value of subpic_level_idc[i] is presumed to be equal to the value of general_level_idc.

[0129] subpic_x_offset[i] specifies the horizontal offset of the upper - left corner of the i - th sub - picture relative to the upper - left corner of the picture. If not present, the value of subpic_x_offset[i] is presumed to be equal to 0. The value of the sub - picture x offset is SubpictureXOffset[i]=subpic It is derived like _x_offset[i]. subpic_y_offset[i] is the vertical offset of the upper left corner of the i-th subpicture relative to the upper left corner of the picture. Specifies the vertical offset of the upper left corner of the i-th subpicture relative to the upper left corner of the picture. When it does not exist , the value of subpic_y_offset[i] is presumed to be equal to 0. The value of the subpicture y offset is derived like SubpictureYOffset[i]=subpic_y_offset[i]. subpic_width_in_l uma_samples[i] specifies the width of the i-th decoded subpicture for which this SPS is the active SPS . When the sum of SubpictureXOffset[i] and subpic_width_in_luma_samples[i] is less than pic_wid th_in_luma_samples, the value of subpic_width_in_luma_samples[i] shall be an integer multiple of CtbSizeY . When it does not exist, the value of subpic_width_in_luma_samples[i] is presumed to be equal to the value of pic_width_in_luma_samples. subpic_height_in_luma_sampl es[i] specifies the height of the i-th decoded subpicture for which this SPS is the active SPS . When the sum of SubpictureYOffset[i] and subpic_height_in_luma_samples[i] is less than pic_height_in_ luma_samples, the value of subpic_height_in_luma_samples[i] shall be an integer multiple of CtbSizeY. When it does not exist, the value of subpic_height_in_luma_samples[i] is, pic_ height_in_luma_samples is presumed to be equal to the value.

[0130] The union of sub - pictures should cover the entire picture area without overlap and gaps This is a requirement for bit - stream conformity. subpic_motion_constrain equal to 1 subpic_motion_constrained_flag[i] specifies that the i - th sub - picture is a temporally motion - constrained sub - picture subpic_motion_constrained_flag[i] equal to 0 specifies that the i - th sub - picture may or may not be a temporally motion - constrained sub - picture. When it does not exist, the value of subpic_motion_constrained_flag is assumed to be equal to 0 constrained sub - picture. When it does not exist, the value of subpic_motion_constrained_flag is assumed to be equal to 0 constrained sub - picture. When it does not exist, the value of subpic_motion_constrained_flag is assumed to be equal to 0

[0131] The variables SubpicWidthInCtbsY, SubpicHeightInCtbsY, SubpicSizeInCtbsY, SubpicWidthInM inCbsY, SubpicHeightInMinCbsY, SubpicSizeInMinCbsY, SubpicSizeInSamplesY, Subpic WidthInSamplesC, and SubpicHeightInSamplesC are derived as follows SubpicWidthInLumaSamples[i]=subpic_width_in_luma_samples[i] SubpicHeightInLumaSamples[i]=subpic_height_in_luma_samples[i] SubPicRightBorderInPic[i]=SubpictureXOffset[i]+PicWidthInLumaSamples[i] SubPicBottomBorderInPic[i]=SubpictureYOffset[i]+PicHeightInLumaSamples[i] SubpicWidthInCtbsY[i] = Ceil(SubpicWidthInLumaSamples[i] ÷ CtbSizeY) SubpicHeightInCtbsY[i] = Ceil(SubpicHeightInLumaSamples[i] ÷ CtbSizeY) SubpicSizeInCtbsY[i] = SubpicWidthInCtbsY[i] * SubpicHeightInCtbsY[i] SubpicWidthInMinCbsY[i] = SubpicWidthInLumaSamples[i] / MinCbSizeY SubpicHeightInMinCbsY[i] = SubpicHeightInLumaSamples[i] / MinCbSizeY SubpicSizeInMinCbsY[i] = SubpicWidthInMinCbsY[i] * SubpicHeightInMinCbsY[i] SubpicSizeInSamplesY[i] = SubpicWidthInLumaSamples[i] * SubpicHeightInLumaSamples[i] SubpicWidthInSamplesC[i] = SubpicWidthInLumaSamples[i] / SubWidthC SubpicHeightInSamplesC[i] = SubpicHeightInLumaSamples[i] / SubHeightC

[0132] The syntax and semantics of the sub-picture parameter set RBSP are as follows as follows

Table 4

[0133] spps_subpic_id identifies the sub-picture to which the SPPS belongs. The length of spps_subpic_id is It is subpic_id_len_minus1 + 1 bits. The spps_subpic_parameter_set_id identifies the SPPS for reference by other syntax elements. The value of spps_subpic_parameter_set_id shall be in the range of 0 to 63, inclusive at both ends. The spps_seq_parameter_set_id specifies the value of sps_seq_parameter_set_id for the active SPS. The value of spps_seq_parameter_set_id shall be in the range of 0 to 15, inclusive at both ends. Equal to 1, single_tile_in_subpic_fl ag specifies that there is only one tile in each subpicture that references the SPPS. Equal to 0 single_tile_in_subpic_flag specifies that there is more than one tile in each subpicture that references the SPPS. Adding 1 to num_tile_columns_minus1 specifies the number of tile columns that divide the subpic ture. num_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY[spps_subpic_id] - 1, inclusive at both ends. When it does not exist, the value of num_t ile_columns_minus1 is presumed to be equal to 0. Adding 1 to num_tile_rows_minus1 specifies the number of tile rows that divide the subpicture. num_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY[spps_subpic_id] - 1, inclusive at both ends. When it does not exist , the value of num_tile_rows_minus1 is presumed to be equal to 0. The variable NumTilesInPic is (n ). Adding 1 to num_tile_rows_minus1 specifies the number of tile rows that divide the subpicture. num_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY[spps_subpic_id] - 1, inclusive at both ends. When it does not exist , the value of num_tile_rows_minus1 is presumed to be equal to 0. The variable NumTilesInPic is (n ). (num_tile_columns_minus1 + 1) * (num_tile_rows_minus1 + 1). When single_tile_in_subpic_flag is 1, NumTilesInPic is 1. When single_tile_in_subpic_flag is 0, NumTilesInPic is (num_tile_columns_minus1 + 1) * (num_tile_rows_minus1 + 1). The value of spps_subpic_id shall be in the range of 0 to subpic_id_len_minus1, inclusive at both ends. The syntax of the sub-picture parameter set shall be as follows: (spps_subpic_parameter_set_id It is set to be equal to (um_tile_columns_minus1 + 1) * (num_tile_rows_minus1 + 1).

[0134] When single_tile_in_subpic_flag is equal to 0, NumTilesInPic is greater than 0. A uniform_tile_spacing_flag equal to 1 specifies that the boundaries of tile columns, as well as the boundaries of tile rows, are uniformly distributed across the entire subpicture. A uniform_tile_spacing_flag equal to 0 specifies that the boundaries of tile columns, as well as the boundaries of tile rows, are not uniformly distributed across the entire subpicture and are explicitly signaled using the syntax elements tile_column_width_minus1[i] and tile_row_height_minus1[i]. When it does not exist, the value of uniform_tile_spacing_flag is assumed to be equal to 1. Adding 1 to tile_column_width_minus1[i] specifies the width of the i-th tile column in units of CTB. Adding 1 to tile_row_height_minus1[i] specifies the height of the i-th tile row in units of CTB. A uniform_tile_spacing_flag equal to 0 specifies that the boundaries of tile columns, as well as the boundaries of tile rows, are not uniformly distributed across the entire subpicture and are explicitly signaled using the syntax elements tile_column_width_minus1[i] and tile_row_height_minus1[i]. When it does not exist, the value of uniform_tile_spacing_flag is assumed to be equal to 1. Adding 1 to tile_column_width_minus1[i] specifies the width of the i-th tile column in units of CTB. Adding 1 to tile_row_height_minus1[i] specifies the height of the i-th tile row in units of CTB. not uniformly distributed across the entire subpicture and are explicitly signaled using the syntax elements tile_column_width_minus1[i] and tile_row_height_minus1[i]. When it does not exist, the value of uniform_tile_spacing_flag is assumed to be equal to 1. Adding 1 to tile_column_width_minus1[i] specifies the width of the i-th tile column in units of CTB. Adding 1 to tile_row_height_minus1[i] specifies the height of the i-th tile row in units of CTB. not uniformly distributed across the entire subpicture and are explicitly signaled using the syntax elements tile_column_width_minus1[i] and tile_row_height_minus1[i]. When it does not exist, the value of uniform_tile_spacing_flag is assumed to be equal to 1. Adding 1 to tile_column_width_minus1[i] specifies the width of the i-th tile column in units of CTB. Adding 1 to tile_row_height_minus1[i] specifies the height of the i-th tile row in units of CTB. not uniformly distributed across the entire subpicture and are explicitly signaled using the syntax elements tile_column_width_minus1[i] and tile_row_height_minus1[i]. When it does not exist, the value of uniform_tile_spacing_flag is assumed to be equal to 1. Adding 1 to tile_column_width_minus1[i] specifies the width of the i-th tile column in units of CTB. Adding 1 to tile_row_height_minus1[i] specifies the height of the i-th tile row in units of CTB. Adding 1 to tile_column_width_minus1[i] specifies the width of the i-th tile column in units of CTB. Adding 1 to tile_row_height_minus1[i] specifies the height of the i-th tile row in units of CTB. Adding 1 to tile_row_height_minus1[i] specifies the height of the i-th tile row in units of CTB.

[0135] The following variables, namely, the list ColWidth[i] for i ranging from 0 to num_tile_columns_minus1 inclusive, which specifies the width of the i-th tile column in units of CTB, the list RowHeight[j] for j ranging from 0 to num_tile_rows_minus1 inclusive, which specifies the height of the j-th tile row in units of CTB, and the list that specifies the position of the i-th tile column boundary in units of CTB from 0 to num_tile_columns_minus1 inclusive for i, the list RowHeight[j] for j ranging from 0 to num_tile_rows_minus1 inclusive, which specifies the height of the j-th tile row in units of CTB, and the list that specifies the position of the i-th tile column boundary in units of CTB from 0 to num_tile_rows_minus1 inclusive for j, and the list that specifies the position of the i-th tile column boundary in units of CTB from 0 to num_tile_rows_minus1 inclusive for j, and the list that specifies the position of the i-th tile column boundary in units of CTB , a list ColBd[i] for i ranging from 0 to num_tile_columns_minus1 + 1 including both ends, specifying the position of the j-th tile row boundary in units of CTB, a list RowBd[j] for j ranging from 0 to num_tile_row s_minus1 + 1 including both ends, a list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the tile scan, a list CtbAddrTsToRs[ctbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the tile scan to the CTB address in the CTB raster scan of the picture, a list TileId[c tbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the tile scan to the tile ID, a list NumCtusInTile[tileIdx] for tileIdx ranging from 0 to PicSizeInCtbsY - 1 including both ends, specifying the conversion from the tile index to the number of CTUs in the tile, a list FirstCtbAddrTs[ti leIdx] for tileIdx ranging from 0 to NumTilesInPic - 1 including both ends, specifying the conversion from the tile ID to the CTB address in the tile scan of the first CTB in the tile, a list ColumnWidthInLumaSamples[i] for i ranging from 0 to num_ tile_columns_minus1 including both ends, specifying the width of the i-th tile column in units of luma samples, and a list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the tile scan, a list CtbAddrTsToRs[ctbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the tile scan to the CTB address in the CTB raster scan of the picture, a list TileId[c tbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the tile scan to the tile ID, a list NumCtusInTile[tileIdx] for tileIdx ranging from 0 to PicSizeInCtbsY - 1 including both ends, specifying the conversion from the tile index to the number of CTUs in the tile, a list FirstCtbAddrTs[ti leIdx] for tileIdx ranging from 0 to NumTilesInPic - 1 including both ends, specifying the conversion from the tile ID to the CTB address in the tile scan of the first CTB in the tile, a list ColumnWidthInLumaSamples[i] for i ranging from 0 to num_ tile_columns_minus1 including both ends, specifying the width of the i-th tile column in units of luma samples, and a list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the tile scan, a list CtbAddrTsToRs[ctbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the tile scan to the CTB address in the CTB raster scan of the picture, a list TileId[c tbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the tile scan to the tile ID, a list NumCtusInTile[tileIdx] for tileIdx ranging from 0 to PicSizeInCtbsY - 1 including both ends, specifying the conversion from the tile index to the number of CTUs in the tile, a list FirstCtbAddrTs[ti leIdx] for tileIdx ranging from 0 to NumTilesInPic - 1 including both ends, specifying the conversion from the tile ID to the CTB address in the tile scan of the first CTB in the tile, a list ColumnWidthInLumaSamples[i] for i ranging from 0 to num_ tile_columns_minus1 including both ends, specifying the width of the i-th tile column in units of luma samples, and a list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the tile scan, a list CtbAddrTsToRs[ctbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the tile scan to the CTB address in the CTB raster scan of the picture, a list TileId[c tbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY - 1, specifying the conversion from the CTB address in the tile scan to the tile ID, a list NumCtusInTile[tileIdx] for tileIdx ranging from 0 to PicSizeInCtbsY - 1 including both ends, specifying the conversion from the tile index to the number of CTUs in the tile, a list FirstCtbAddrTs[ti Also, a list RowHeightInLumaSamples[j] for j ranging from 0 to num_ tile_rows_minus1, inclusive, that specifies the height of the j-th tile row in luma samples, is derived by calling the CTB la st and tile scan conversion processes. ColumnWidthInLumaSamples[i] for i ranging from 0 to num_tile_columns_minus1, inclusive, and RowHeightInLumaSamples[j] for j ranging from 0 to num_tile_rows_minus1, inclusive, shall all have values greater than 0. The loop_filter_across_tiles_enabled_flag equal to 1 specifies that the loop filter operation may be performed across tile boundaries in sub-pictures that reference the SPPS. The loop_filter_across_tiles_enabled_flag equal to 0 specifies that the loop filter operation shall not be performed across tile boundaries in sub-pictures that reference the SPPS. The loop filter operation includes deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, the value of the loop_filter_acro

[0136] ss_tiles_enabled_flag is assumed to be equal to 1. The loop_filter_across_subpic_enabled_flag equal to 1 specifies that the loop filter operation may be performed across sub-picture boundaries in sub-pictures that reference the SPPS. The loop_filter_across_subpic_enabled_flag equal to 0 specifies that the loop_filt er operation shall not be performed across sub-picture boundaries in sub-pictures that reference the SPPS. The loop_filter_across_tiles_enabled_flag equal to 0 specifies that the loop filter operation shall not be performed across tile boundaries in sub-pictures that reference the SPPS. The loop filter operation includes deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, the value of the loop_filter_across_tiles_enabled_flag is assumed to be equal to 1. The loop_filter_across_subpic_enabled_flag equal to 1 specifies that the loop filter operation may be performed across sub-picture boundaries in sub-pictures that reference the SPPS. The loop_filter_across_subpic_enabled_flag equal to 0 specifies that the loop_filter operation shall not be performed across sub-picture boundaries in sub-pictures that reference the SPPS. The loop_filter_across_tiles_enabled_flag equal to 0 specifies that the loop filter operation shall not be performed across tile boundaries in sub-pictures that reference the SPPS. The loop filter operation includes deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, the value of the loop_filter_across_tiles_enabled_flag is assumed to be equal to 1. The loop_filter_across_subpic_enabled_flag equal to 1 specifies that the loop filter operation may be performed across sub-picture boundaries in sub-pictures that reference the SPPS. The loop_filter_across_subpic_enabled_flag equal to 0 specifies that the loop_filter operation shall not be performed across sub-picture boundaries in sub-pictures that reference the SPPS. The loop_filter_across_subpic_enabled_flag equal to 1 specifies that the loop filter operation may be performed across sub-picture boundaries in sub-pictures that reference the SPPS. The loop_filter_across_subpic_enabled_flag equal to 0 specifies that the loop_filter operation shall not be performed across sub-picture boundaries in sub-pictures that reference the SPPS. The loop_filter_across_subpic_enabled_flag equal to 1 specifies that the loop filter operation may be performed across sub-picture boundaries in sub-pictures that reference the SPPS. The loop_filter_across_subpic_enabled_flag equal to 0 specifies that the loop_filter operation shall not be performed across sub-picture boundaries in sub-pictures that reference the SPPS. The er_across_subpic_enabled_flag specifies that the in-loop filtering operation is not performed across sub-picture boundaries within the sub-pictures that refer to the SPPS. The in-loop filtering operation includes the deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When it does not exist, the value of the loop_filter_across_subpic_enable d_flag is assumed to be equal to the value of the loop_filter_across_tiles_enabed_flag.

[0137] The syntax and semantics of the general tile group header are as follows .

Table 5

[0138] The values of the tile group header syntax elements tile_group_pic_parameter_set_id and tile_ group_pic_order_cnt_lsb are the same for all tile group headers of the coded picture. The value of the tile group header syntax element tile_g roup_subpic_id is the same for all tile group headers of the coded sub-picture. The tile_group_subpic_id identifies the sub-picture to which the tile group belongs. The length of the tile_group_subpic_id is subpic_id_len_minus1 + 1 bits . The tile_group_subpic_parameter_set_id is the spps_subpic_pa for the active SPPS ​​​​Specify the value of parameter_set_id. The value of tile_group_spps_parameter_set_id shall be in the range from 0 to 63, inclusive at both ends.

[0139] The following variables are derived and overwrite each variable derived from the active SPS. PicWidthInLumaSamples = SubpicWidthInLumaSamples[tile_group_subpic_id] PicHeightInLumaSamples = PicHeightInLumaSamples[tile_group_subpic_id] SubPicRightBorderInPic = SubPicRightBorderInPic[tile_group_subpic_id] SubPicBottomBorderInPic = SubPicBottomBorderInPic[tile_group_subpic_id] PicWidthInCtbsY = SubPicWidthInCtbsY[tile_group_subpic_id] PicHeightInCtbsY = SubPicHeightInCtbsY[tile_group_subpic_id] PicSizeInCtbsY = SubPicSizeInCtbsY[tile_group_subpic_id] PicWidthInMinCbsY = SubPicWidthInMinCbsY[tile_group_subpic_id] PicHeightInMinCbsY = SubPicHeightInMinCbsY[tile_group_subpic_id] PicSizeInMinCbsY = SubPicSizeInMinCbsY[tile_group_subpic_id] PicSizeInSamplesY = SubPicSizeInSamplesY[tile_group_subpic_id]​ PicWidthInSamplesC = SubPicWidthInSamplesC[tile_group_subpic_id] PicHeightInSamplesC = SubPicHeightInSamplesC[tile_group_subpic_id]

[0140] The syntax of the coding tree unit is as follows.

Table 6

Table 7

[0141] The syntax and semantics of the coding quad tree are as follows.

Table 8A

Table 8B

[0142] qt_split_cu_flag[x0][y0] specifies whether the coding unit is split into coding units with half the horizontal size and vertical size. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block considered relative to the top-left luma sample of the picture. When qt_split_cu_flag[x0][y0] does not exist, the following applies. If one or more of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. treeType is equal to DUAL_TREE_CHROMA in the coding unit with half the horizontal size and vertical size. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block considered relative to the top-left luma sample of the picture. When qt_split_cu_flag[x0][y0] does not exist, the following applies. If one or more of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. treeType is equal to DUAL_TREE_CHROMA in the coding unit with half the horizontal size and vertical size. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block considered relative to the top-left luma sample of the picture. When qt_split_cu_flag[x0][y0] does not exist, the following applies. If one or more of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. treeType is equal to DUAL_TREE_CHROMA in the coding unit with half the horizontal size and vertical size. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block considered relative to the top-left luma sample of the picture. When qt_split_cu_flag[x0][y0] does not exist, the following applies. If one or more of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. treeType is equal to DUAL_TREE_CHROMA in the coding unit with half the horizontal size and vertical size. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block considered relative to the top-left luma sample of the picture. When qt_split_cu_flag[x0][y0] does not exist, the following applies. If one or more of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. treeType is equal to DUAL_TREE_CHROMA in the coding unit with half the horizontal size and vertical size. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block considered relative to the top-left luma sample of the picture. When qt_split_cu_flag[x0][y0] does not exist, the following applies. If one or more of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. treeType is equal to DUAL_TREE_CHROMA If it is greater than MaxBtSizeY, or otherwise, x0+(1<<log2CbSize) is greater than SubPicRi ghtBorderInPic, and (1<<log2CbSize) is greater than MaxBtSizeC. If treeType is equal to DUAL_ TREE_CHROMA, or otherwise, if it is greater than MaxBtSizeY, y0+(1<<log2 CbSize) is greater than SubPicBottomBorderInPic, and (1<<log2CbSize) is greater than MaxBtSizeC .

[0143] Otherwise, if all of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. If treeType is equal to DUAL_TREE_CHROMA, or otherwise, if it is greater than MinQtSizeY, x0+(1<<log2CbSize) is greater than SubPicRightBorderInPic , y0+(1<<log2CbSize) is greater than SubPicBottomBorderInPic, and (1<<log2CbSize) is greater than MinQtS izeC. Otherwise, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 0 .

[0144] The syntax and semantics of the multi-type tree are as follows.

Table 9A

Table 9B

Table 9C

[0145] An mtt_split_cu_flag equal to 0 specifies that the coding unit is not split. An mtt_split_cu_flag equal to 1 means that, as indicated by the syntax element mtt_split_cu_binary_flag, the coding unit is split into two coding units using binary splitting or into three coding units using ternary splitting. The binary or ternary splitting can be either vertical or horizontal, as indicated by the syntax element mtt_split_cu_vertical_flag. When the mtt_split_cu_flag is absent, its value is inferred as follows. If one or more of the conditions that x0 + cbWidth is greater than SubPicRightBorderInPic and y0 + cbHeight is greater than SubPicBottomBorderInPic are true, the value of the mtt_split_cu_flag is inferred to be 1. Otherwise, the value of the mtt_split_cu_flag is inferred to be 0. The derivation process for the temporal luma motion vector prediction is as follows. The output of this process is the motion vector prediction mvLXCol with 1 / 16 fractional sample accuracy and the availability flag availableFlagLXCol. The variable currCb specifies the current luma coding block at the luma position (xCb, yCb). The variables mvLXCol and availableFlagLXCol are derived as follows. If the tile_group_temporal_mvp_enabled_flag is equal to 0, or if the reference picture is not present, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the current coding tree unit (CTU) is not the top-left CTU of the tile, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the current CTU is the top-left CTU of the tile, then the process continues as follows. Let prevCb be the luma coding block adjacent to the current luma coding block in the previous picture at the same luma position. If prevCb does not exist, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the reference picture list used for motion compensation for prevCb is different from the reference picture list used for motion compensation for the current picture, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the motion vector of prevCb in the reference picture list used for motion compensation for prevCb is not available, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, mvLXCol is set to the motion vector of prevCb in the reference picture list used for motion compensation for prevCb, and availableFlagLXCol = 1. not present, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the current coding tree unit (CTU) is not the top-left CTU of the tile, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the current CTU is the top-left CTU of the tile, then the process continues as follows. Let prevCb be the luma coding block adjacent to the current luma coding block in the previous picture at the same luma position. If prevCb does not exist, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the reference picture list used for motion compensation for prevCb is different from the reference picture list used for motion compensation for the current picture, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the motion vector of prevCb in the reference picture list used for motion compensation for prevCb is not available, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, mvLXCol is set to the motion vector of prevCb in the reference picture list used for motion compensation for prevCb, and availableFlagLXCol = 1. If one or more of the conditions that x0 + cbWidth is greater than SubPicRightBorderInPic and y0 + cbHeight is greater than SubPicBottomBorderInPic are true, the value of the mtt_split_cu_flag is inferred to be 1. Otherwise, the value of the mtt_split_cu_flag is inferred to be 0. If one or more of the conditions that x0 + cbWidth is greater than SubPicRightBorderInPic and y0 + cbHeight is greater than SubPicBottomBorderInPic are true, the value of the mtt_split_cu_flag is inferred to be 1. Otherwise, the value of the mtt_split_cu_flag is inferred to be 0. If one or more of the conditions that x0 + cbWidth is greater than SubPicRightBorderInPic and y0 + cbHeight is greater than SubPicBottomBorderInPic are true, the value of the mtt_split_cu_flag is inferred to be 1. Otherwise, the value of the mtt_split_cu_flag is inferred to be 0.

[0146] The derivation process for the temporal luma motion vector prediction is as follows. The output of this process is the motion vector prediction mvLXCol with 1 / 16 fractional sample accuracy and the availability flag availableFlagLXCol. The variable currCb specifies the current luma coding block at the luma position (xCb, yCb). The variables mvLXCol and availableFlagLXCol are derived as follows. If the tile_group_temporal_mvp_enabled_flag is equal to 0, or if the reference picture is not present, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the current coding tree unit (CTU) is not the top-left CTU of the tile, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the current CTU is the top-left CTU of the tile, then the process continues as follows. Let prevCb be the luma coding block adjacent to the current luma coding block in the previous picture at the same luma position. If prevCb does not exist, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the reference picture list used for motion compensation for prevCb is different from the reference picture list used for motion compensation for the current picture, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the motion vector of prevCb in the reference picture list used for motion compensation for prevCb is not available, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, mvLXCol is set to the motion vector of prevCb in the reference picture list used for motion compensation for prevCb, and availableFlagLXCol = 1. The variable currCb specifies the current luma coding block at the luma position (xCb, yCb). The variables mvLXCol and availableFlagLXCol are derived as follows. If the tile_group_temporal_mvp_enabled_flag is equal to 0, or if the reference picture is not present, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the current coding tree unit (CTU) is not the top-left CTU of the tile, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the current CTU is the top-left CTU of the tile, then the process continues as follows. Let prevCb be the luma coding block adjacent to the current luma coding block in the previous picture at the same luma position. If prevCb does not exist, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the reference picture list used for motion compensation for prevCb is different from the reference picture list used for motion compensation for the current picture, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the motion vector of prevCb in the reference picture list used for motion compensation for prevCb is not available, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, mvLXCol is set to the motion vector of prevCb in the reference picture list used for motion compensation for prevCb, and availableFlagLXCol = 1. not present, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the current coding tree unit (CTU) is not the top-left CTU of the tile, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the current CTU is the top-left CTU of the tile, then the process continues as follows. Let prevCb be the luma coding block adjacent to the current luma coding block in the previous picture at the same luma position. If prevCb does not exist, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the reference picture list used for motion compensation for prevCb is different from the reference picture list used for motion compensation for the current picture, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, if the motion vector of prevCb in the reference picture list used for motion compensation for prevCb is not available, then mvLXCol = (0, 0) and availableFlagLXCol = 0. Otherwise, mvLXCol is set to the motion vector of prevCb in the reference picture list used for motion compensation for prevCb, and availableFlagLXCol = 1. If it is the current picture, both components of mvLXCol are set equal to 0, and availableFlagL XCol is set equal to 0. Otherwise (tile_group_temporal_mvp_enabled_flag is equal to 1 and the reference picture is not the current picture), the following steps in order are applied . The motion vectors at the same position in the lower right are derived as follows. xColBr = xCb + cbWidth (8 - 355) yColBr = yCb + cbHeight (8 - 356)

[0147] If yCb >> CtbLog2SizeY is equal to yColBr >> CtbLog2SizeY, yColBr is less than SubPicBottomBorderInP ic, and xColBr is less than SubPicRightBorderInPic, the following applies . The variable colCb specifies the luma coding block that encompasses the modified position given by ((xColBr >> 3 ) << 3, (yColBr >> 3) << 3) within the picture at the same position specified by ColPic . The luma position (xColCb, yColCb) is set equal to the top - left sample of the luma coding block at the same position specified by colCb, relative to the top - left luma sample of the picture at the same position specified by ColPic . The derivation process of the motion vector at the same position is called with currCb, colCb, (xColCb, yColCb), r efIdxLX, and sbFlag set equal to 0 as inputs, and the outputs are assigned to mvLXCol and availableF lagLXCol. Otherwise, both components of mvLXCol are set equal to 0 ​​, availableFlagLXCol is set equal to 0.

[0148] The derivation process for the temporal triangular integration candidates is as follows. The variables mvLXColC0, mvLXC olC1, availableFlagLXColC0 and availableFlagLXColC1 are derived as follows. ti If tile_group_temporal_mvp_enabled_flag is equal to 0, both components of mvLXColC0 and mvLXColC1 are set equal to 0, and availableFlagLXColC0 and availableFlagLXColC1 are equal to 0 are set. Otherwise (tile_group_temporal_mvp_enabled_flag is equal to 1), the following steps in order are applied. The motion vector mvLXColC0 at the same position in the lower right is as follows derived. xColBr = xCb + cbWidth (8 - 392) yColBr = yCb + cbHeight (8 - 393)

[0149] If yCb >> CtbLog2SizeY is equal to yColBr >> CtbLog2SizeY, yColBr is less than SubPicBottomBorderInP ic, and xColBr is less than SubPicRightBorderInPic, then the following applies . The variable colCb specifies the luma coding block that includes the modified position given by ((xColBr >> 3 ) << 3, (yColBr >> 3) << 3) within the picture at the same position specified by ColPic. The luma position (xColCb, yColCb) is at the same position as the top - left luma sample of the picture specified by ColPic, at the same position specified by colCb for the luma at the same position specified by colCb. It is set equal to the top-left sample of the macro coding block. The derivation process of the motion vector at the same position is called using currCb, colCb, (xColCb, yColCb), refIdxLXC set equal to 0, and sbFlag as inputs, and the outputs are mvLXColC0 and availableFlagL assigned to XColC0. Otherwise, both components of mvLXColC0 are set equal to 0 , and availableFlagLXColC0 is set equal to 0.

[0150] The derivation process of the constructed affine control point motion vector integration candidates is as follows. With X being 0 and 1, the fourth (bottom-right at the same position) control point motion vector cpMvLXCorner[3] , reference index refIdxLXCorner[3], prediction list usage flag predFlagLXCorner[3], and availability flag availableFlagCorner[3] are derived as follows. With X being 0 or 1, the reference index refIdxLXCorner[3] for the temporal integration candidate is set equal to 0. With X being 0 or 1, the variables mvLXCol and availableFlagLXCol are derived as follows. When tile_group_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are set equal to 0 , and availableFlagLXCol is set equal to 0. Otherwise (when tile_group _temporal_mvp_enabled_flag is equal to 1), the following holds. xColBr = xCb + cbWidth (8 - 566) yColBr = yCb + cbHeight (8 - 567)

[0151] yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, and when yColBr is less than SubPicBottomBorderInP ic and xColBr is less than SubPicRightBorderInPic, the following applies . The variable colCb specifies a luma coding block that includes the modified position given by ((xColBr>>3 )<<3,(yColBr>>3)<<3) within the picture at the same position specified by ColPic . The luma position (xColCb,yColCb) is set equal to the top-left luma sample of the picture at the same position specified by ColPic and the top-left sample of the luma coding block at the same position specified by colCb. The derivation process of the motion vector at the same position is called using currCb, colCb, (xColCb,yColCb), refIdxLX set equal to 0 and sbFlag as inputs, and the output is assigned to mvLXCol and availableFlagLXCo l. Otherwise, both components of mvLXCol are set equal to 0 and avail ableFlagLXCol is set equal to 0. Replace all occurrences of pic_width_in_luma_samples with P icWidthInLumaSamples. Replace all occurrences of pic_height_in_luma_samples with Pi cHeightInLumaSamples.

[0152] In a second exemplary embodiment, the syntax and semantics of the sequence parameter set RBSP are as follows.

Table 10

[0153] Adding 1 to subpic_id_len_minus1 is the number of bits used to represent the syntax element subpic_id[i] in the SPS, the spps_subpic_id in the SPPS that refers to the SPS, and also the tile_group_subpic_id in the tile group header that refers to the SPS. subpic_id_len_minus1's value should be in the range from Ceil(Log2(num_subpic_minus1 + 3)) to 8, inclusive. It is a bitstream conformity requirement that there is no overlap among subpicture[i] for i from 0 to num_subpic_minus1, inclusive. Each subpicture may be a temporal motion constrained subpicture. and the tile_group_subpic_id in the tile group header that refers to the SPS. subpic_id_len_minus1's value should be in the range from Ceil(Log2(num_subpic_minus1 + 3)) to 8, inclusive. It is a bitstream conformity requirement that there is no overlap among subpicture[i] for i from 0 to num_subpic_minus1, inclusive. Each subpicture may be a temporal motion constrained subpicture. and the tile_group_subpic_id in the tile group header that refers to the SPS. subpic_id_len_minus1's value should be in the range from Ceil(Log2(num_subpic_minus1 + 3)) to 8, inclusive. It is a bitstream conformity requirement that there is no overlap among subpicture[i] for i from 0 to num_subpic_minus1, inclusive. Each subpicture may be a temporal motion constrained subpicture. c_id_len_minus1's value should be in the range from Ceil(Log2(num_subpic_minus1 + 3)) to 8, inclusive. It is a bitstream conformity requirement that there is no overlap among subpicture[i] for i from 0 to num_subpic_minus1, inclusive. Each subpicture may be a temporal motion constrained subpicture. c_id_len_minus1's value should be in the range from Ceil(Log2(num_subpic_minus1 + 3)) to 8, inclusive. It is a bitstream conformity requirement that there is no overlap among subpicture[i] for i from ...... to num_subpic_minus1, inclusive. Each subpicture may be a temporal motion constrained subpicture. to num_subpic_minus1, inclusive. Each subpicture may be a temporal motion constrained subpicture.

[0154] The semantics of a general tile group header are as follows. tile_group_subpic_id identifies the subpicture to which the tile group belongs. The length of tile_group_subpic_id is subpic_id_len_minus1 + 1 bits. A tile_group_subpic_id equal to 1 indicates that the tile group does not belong to any subpicture. The semantics of a general tile group header are as follows. tile_group_subpic_id identifies the subpicture to which the tile group belongs. The length of tile_group_subpic_id is subpic_id_len_minus1 + 1 bits. A tile_group_subpic_id equal to 1 indicates that the tile group does not belong to any subpicture. The semantics of a general tile group header are as follows. tile_group_subpic_id identifies the subpicture to which the tile group belongs. The length of tile_group_subpic_id is subpic_id_len_minus1 + 1 bits. A tile_group_subpic_id equal to 1 indicates that the tile group does not belong to any subpicture. The semantics of a general tile group header are as follows. tile_group_subpic_id identifies the subpicture to which the tile group belongs. The length of tile_group_subpic_id is subpic_id_len_minus1 + 1 bits. A tile_group_subpic_id equal to 1 indicates that the tile group does not belong to any subpicture.

[0155] In a third exemplary embodiment, the syntax and semantics of the NAL unit header are as follows. In a third exemplary embodiment, the syntax and semantics of the NAL unit header are as follows. [Table 11]

[0156] nuh_subpicture_id_len represents the syntax element that specifies the subpicture ID. Specify the number of bits used. When the value of nuh_subpicture_id_len is greater than 0, the first nuh_subpicture_id_len bits after nuh_reserved_zero_4bits specify the ID of the subpicture to which the payload of the NAL unit belongs. When nuh_subpicture_id_len is greater than 0, the value of nuh_subpicture_id_len shall be equal to the value of subpic_id_len_minus1 in the active SPS. The value of nuh_subpicture_id_len for non-VCL NAL units is restricted as follows. When nal_unit_type is equal to SPS_NUT or PPS_NUT, nuh_subpicture_id_len shall be equal to 0. nuh_reserved_zero_3bits shall be equal to "000". The decoder shall ignore (e.g., remove and discard from the bitstream) NAL units for which the value of nuh_reserved_zero_3bits is not equal to "000". After nuh_reserved_zero_4bits, the first nuh_subpicture_id_len bits specify the ID of the subpicture to which the payload of the NAL unit belongs. When nuh_subpicture_id_len is greater than 0, the value of nuh_subpicture_id_len shall be equal to the value of subpic_id_len_minus1 in the active SPS. When nuh_subpicture_id_len is greater than 0, the value of nuh_subpicture_id_len shall be equal to the value of subpic_id_len_minus1 in the active SPS. For non-VCL NAL units, the value of nuh_subpicture_id_len is restricted as follows. When nal_unit_type is equal to SPS_NUT or PPS_NUT, nuh_subpicture_id_len shall be equal to 0. nuh_reserved_zero_3bits shall be equal to "000". The decoder shall ignore (e.g., remove and discard from the bitstream) NAL units for which the value of nuh_reserved_zero_3bits is not equal to "000". When nal_unit_type is equal to SPS_NUT or PPS_NUT, nuh_subpicture_id_len shall be equal to 0. nuh_reserved_zero_3bits shall be equal to "000". The decoder shall ignore (e.g., remove and discard from the bitstream) NAL units for which the value of nuh_reserved_zero_3bits is not equal to "000". When nal_unit_type is equal to SPS_NUT or PPS_NUT, nuh_subpicture_id_len shall be equal to 0. nuh_reserved_zero_3bits shall be equal to "000". The decoder shall ignore (e.g., remove and discard from the bitstream) NAL units for which the value of nuh_reserved_zero_3bits is not equal to "000". When nal_unit_type is equal to SPS_NUT or PPS_NUT, nuh_subpicture_id_len shall be equal to 0. nuh_reserved_zero_3bits shall be equal to "000". The decoder shall ignore (e.g., remove and discard from the bitstream) NAL units for which the value of nuh_reserved_zero_3bits is not equal to "000". When nal_unit_type is equal to SPS_NUT or PPS_NUT, nuh_subpicture_id_len shall be equal to 0. nuh_reserved_zero_3bits shall be equal to "000". The decoder shall ignore (e.g., remove and discard from the bitstream) NAL units for which the value of nuh_reserved_zero_3bits is not equal to "000".

[0157] In a fourth exemplary embodiment, the subpicture nesting syntax is as follows. [Table 12]

[0158] An all_sub_pictures_flag equal to 1 indicates that the nested SEI messages apply to all subpictures. An all_sub_pictures_flag equal to 1 indicates that the subpictures to which the nested SEI messages apply are explicitly signaled by subsequent syntax elements. An all_sub_pictures_flag equal to 1 indicates that the nested SEI messages apply to all subpictures. An all_sub_pictures_flag equal to 1 indicates that the subpictures to which the nested SEI messages apply are explicitly signaled by subsequent syntax elements. An all_sub_pictures_flag equal to 1 indicates that the nested SEI messages apply to all subpictures. An all_sub_pictures_flag equal to 1 indicates that the subpictures to which the nested SEI messages apply are explicitly signaled by subsequent syntax elements. nesting_num_sub_pictures_minus1 plus 1 specifies the number of nesting pictures. Specifies the number of subpictures to which the specified SEI message applies. _id[i] is the subpicture of the ith subpicture to which the nested SEI message applies The nesting_sub_picture_id[i] syntax element is Ceil(Log2(nesting_num_sub _pictures_minus1+1)) bits. sub_picture_nesting_zero_bit is equal to 0. It shall be considered as new.

[0159] 9 is a schematic diagram of an example video coding device 900. The device 900 is suitable for implementing the disclosed examples / embodiments as described herein. The video coding device 900 transmits data upstream over a network. and / or a transmitter and / or a receiver for communicating the , downstream port 920, upstream port 950, and / or transceiver The video coding device 900 also includes a transceiver unit (Tx / Rx) 910. a processor 930 including a logic unit and / or central processing unit (CPU) for and memory 932 for storing the data. Video coding device 900 also includes an electrical Components, Optical-Electrical (OE) Components, Electro-Optical (EO) Components, and / or or telecommunications networks, optical communications networks, or wireless communications networks Upstream port 950 and / or downstream port 960 for communication of data over the network It may include a wireless communication component coupled to the Ream port 920. The video - deing device 900 may also include an input and / or output (I / O) device 960 for communicating data with the user. The I / O device 960 may include output devices such as a display for displaying video data and speakers for outputting audio data. The I / O device 960 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0160] The processor 930 is implemented by hardware and software. The processor 930 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 930 communicates with the downstream port 920, Tx / Rx 910, the upstream port 950, and the memory 932. The processor 930 includes a coding module 914. The coding module 914 may utilize the bit stream 500, the picture 600, and / or the picture 800, and implement the disclosed embodiments described above, such as the methods 100, 1 000, 1100, and / or the mechanism 700. The coding module 914 may also implement any other method / mechanism described herein. Furthermore, the coding module 914 may be a codec system 200, an encoder ​Encoder 300 and / or decoder 400 may be implemented. For example, coding module rule 914 may be used to signal and / or obtain sub-picture positions and sizes in the SPS. In another example, coding module 914 may constrain the sub-picture width and height to be multiples of the CTU size, unless such sub-pictures are located at the right or bottom boundaries of the picture, respectively. In another example, coding module 914 may constrain the sub-pictures so as to cover the picture without gaps or overlaps. In another example, coding module 914 may be used to signal and / or obtain data indicating that some sub-pictures are time motion constrained sub-pictures and other sub-pictures are not. In another example, coding module 914 may signal a complete set of sub-picture IDs in the SPS and include sub-picture IDs in each slice header to indicate the sub-pictures including the corresponding slices. In another example, coding module 914 may signal the level of each sub-picture. Thus, coding module 914 enables the video coding device 900 to provide additional functions, reduce processing overhead when segmenting and coding video data, and / or avoid certain processing to improve coding efficiency. Thus, coding module 914 improves the functionality of the video coding device 900 and addresses problems specific to video coding techniques. Further, coding module 914 addresses different situations. In addition, coding module 914 addresses problems specific to video coding techniques. Perform the conversion of the video coding device 900 to the state. Alternatively, the coding module ule 914 may be implemented as instructions stored in the memory 932 and executed by the processor 930 (for example, as a computer program product stored on a non-transitory medium).

[0161] The memory 932 includes one or more memory types such as a disk, a tape drive, a solid state drive, a read-only memory (ROM), a random access memory (RAM), a flash memory, a ternary content addressable memory (TCAM) , a static random access memory (SRAM), etc. The memory 932 stores such a program when a program is selected for execution, and may be used as an overflow data storage device to store the instructions and data read during program execution.

[0162] Figure 10 is a flowchart of an exemplary method 1000 for encoding a bitstream 500 and / or a sub-bitstream 501, etc., of sub-pictures such as sub-pictures 522, 523, 622, 722, and / or 822, with adaptive size constraints. <000208=28]]The method 1000 may be utilized by an encoder such as the codec system 200, the encoder 300, and / or the video coding device 900 when performing the method 100.

[0163] The method 1000 may begin when an encoder receives a video sequence including a plurality of pictures and determines to encode the video sequence into a bitstream, for example, based on user input. The video sequence is partitioned further for pictures before encoding.​ / Image / Frame. In step 1001, the picture is divided into a plurality of sub-pictures . When dividing the sub-pictures, adaptive size constraints are applied. Each sub-picture includes a sub-picture width and a sub-picture height. When the current sub-picture includes a right boundary that does not coincide with the right boundary of the picture, the sub-picture width of the current sub-picture is constrained to be an integer multiple of the CTU (for example, sub-picture 822 other than 822b) . Therefore, at least one of the sub-pictures (for example, sub-picture 822b) has a sub-picture width that is not an integer multiple of the CTU size when each sub-picture includes a right boundary that coincides with the right boundary of the picture. When the current sub-picture includes a bottom boundary that does not coincide with the bottom boundary of the picture, the sub-picture height of the current sub-picture is constrained to be an integer multiple of the CTU (for example, sub-picture 822 other than 822c). Therefore , at least one of the sub-pictures (for example, sub-picture 822c) may include a sub-picture height that is not an integer multiple of the CTU size when each sub-picture includes a bottom boundary that coincides with the bottom boundary of the picture. This adaptive size constraint allows sub-pictures to be divided from a picture that includes a picture width that is not an integer multiple of the CTU size and / or a picture height that is not an integer multiple of the CTU size. The CTU size may be measured in units of luma samples . In step 1003, one or more of the sub-pictures are encoded into a bitstream . In step 1005, the bitstream is stored for communication to the decoder .

[0164] ​​​​​​The bitstream may then be transmitted to a decoder as desired. In this example, the sub-bitstreams may be extracted from the encoded bitstream. In such cases, the transmitted bitstream is a sub-bitstream. In the example, the encoded bitstream is used to extract the sub-bitstreams at the decoder. In yet another example, the encoded bitstream may be transmitted for sub-sampling. In any of these examples, the bitstream may be decoded and displayed without extraction. However, adaptive size constraints are not supported for pictures with heights or widths that are not multiples of the CTU size. This enhances the performance of the encoder by allowing sub-pictures to be partitioned from one another.

[0165] FIG. 11 shows the sub-pictures 522, 523, 622, 722, and / or 822, the bitstream 500 and / or the sub-bitstreams 11 is a flowchart of an example method 1100 for decoding a bitstream such as stream 501. The method 1100, when performing the method 100, includes the codec system 200, the decoder 400, and / or may be utilized by a decoder such as video coding device 900. For example, , method 1100 is applied to decode the bitstream produced as a result of method 1000. This may be done.

[0166] The method 1100 begins when the decoder begins receiving a bitstream containing subpictures. The bitstream may contain a complete video sequence, or it may contain only a few bits. The sub-bitstream contains a reduced set of sub-pictures for separate extraction. In step 1101, a bitstream is received. A picture stream consists of one or more subpictures partitioned from a picture according to adaptive size constraints. Each subpicture comprises a subpicture width and a subpicture height. When the current subpicture contains a right border that does not coincide with the right border of the picture, The subpicture width of a picture is constrained to be an integer multiple of the CTU (e.g., for formats other than 822b). Subpicture 822). Therefore, at least one of the subpictures (e.g., Picture 822b) is C when each subpicture contains a right border that coincides with the right border of the picture. It may contain subpicture widths that are not integer multiples of the TU size. contains a bottom border that does not coincide with the bottom border of the picture, the current subpicture's subpicture The height of a subpicture is constrained to be an integer multiple of the CTU (e.g., subpictures 82c other than 822c). 2). Therefore, at least one of the subpictures (e.g., subpicture 822c) , when each subpicture contains a bottom border that matches the bottom border of the picture, the CTU size is This adaptive size constraint may include subpicture heights that are not multiples of the CTU size. picture width that is not an integer multiple of the CTU size and / or picture height that is not an integer multiple of the CTU size. The CTU size is the number of luma samples. It may be measured in units of

[0167] In step 1103, the bitstream is processed to obtain one or more sub-pictures. In step 1105, one or more The sub-picture is decoded. The video sequence can then be transferred for display. Thus, the adaptive size constraint enables sub-pictures to be segmented from pictures with a height or width that is not a multiple of the CTU size. Therefore, the decoder can utilize sub-picture-based functions such as separate sub-picture extraction and / or display on pictures with a height or width that is not a multiple of the CTU size. Thus, the application of the adaptive size constraint enhances the functionality of the decoder.

[0168] FIG. 12 is a schematic diagram of an exemplary system 1200 for signaling a bitstream 500 and / or a sub-bitstream 501, such as sub-pictures 522, 523, 622, 722, and / or 822, with an adaptive size constraint. System 1200 may be implemented by an encoder and a decoder, such as codec system 200, encoder 300, decoder 400, and / or video coding device 900. Further, system 1200 may be utilized when implementing methods 100, 1000, and / or 1100.

[0169] System 1200 includes video encoder 1202. Video encoder 1202 includes a segmentation module 1201 for segmenting a picture into a plurality of sub-pictures such that each sub-picture includes a sub-picture width that is an integer multiple of the CTU size when each sub-picture includes a right boundary that does not coincide with the right boundary of the picture. Video encoder 1202 further includes an encoding module 1203 for encoding one or more of the sub-pictures into a bitstream. 。The video encoder 1202 further comprises a storage module 1205 for storing the bitstream for communication to the decoder. The video encoder 1202 further comprises a transmission module 1207 for transmitting the bitstream including sub-pictures to the decoder. The video encoder 1202 may further be configured to execute any of the steps of method 1000. The system 1200 also includes a video decoder 1210. When each sub-picture includes a right boundary that does not coincide with the right boundary of the picture, the video decoder 1210 includes a reception module 1211 for receiving a bitstream including one or more sub-pictures divided from the picture, such that each sub-picture includes a sub-picture width that is an integer multiple of the coding tree unit (CTU) size. The video decoder 1210 further comprises an analysis module 1213 for analyzing the bitstream to obtain one or more sub-pictures. The video decoder 1210 further comprises a decoding module 1215 for decoding one or more sub-pictures to create a video sequence. The video decoder 1210 further comprises a transfer module 1217 for transferring the video sequence for display. The video decoder 1210 may further be configured to execute any of the steps of method 1100. When there is no intervening component, except for a line, wiring, or another medium between the first component and the second component, the first component is directly coupled to the second component. When there is a line, wiring, or another medium between the first component and the second component,

[0170]

[0171] ​​​​​​​​​​​​​​​When there are intervening components other than the specified medium, the first component is indirectly coupled to the second component. The term "coupled" and its variations include both direct coupling and indirect

[0172] coupling. The use of the term "about" means a range that includes ±10% of the number that follows, unless otherwise stated. The steps of the exemplary methods described herein are not necessarily required to be executed in the

[0173] order described, and it should be understood that the order of the steps of such methods is merely exemplary. Similarly, additional steps may be included in such methods, and some steps may be omitted or combined in a manner consistent with various embodiments of the present disclosure.

[0174] Although several embodiments have been provided in the present disclosure, it should be understood that the disclosed systems and methods may be embodied in many Other examples of variations, substitutions, and alterations are ascertainable by one of ordinary skill in the art and are disclosed herein. may be made without departing from the spirit and scope thereof. [Explanation of symbols]

[0175] 200 Codec System 201 segmented video signal 211 General-Purpose Coder Control Component 213 Transform Scaling and Quantization Components 215 Intra-picture Estimation Component 217 Intra-picture Prediction Component 219 Motion Compensation Component 221 Motion Estimation Component 223 Decoded Picture Buffer Component 225 In-Loop Filter Components 227 Filter Control Analysis Component 229 Scaling and Inverse Transformation Components 231 Header Formatting and CABAC Components 301 Segmented Video Signal 313 Transform and Quantize Components 317 Intra-picture Prediction Component 321 Motion Compensation Component 323 Decoded Picture Buffer Component 325 In-Loop Filter Components 329 Inverse Transform and Quantization Components 331 Entropy Coding Component 417 Intra-picture Prediction Component 421 Motion Compensation Component 423 Decoded Picture Buffer Component 425 In-Loop Filter Components 429 Inverse Transform and Quantization Components 433 Entropy Decoding Component 500 Bitstream 501 Sub-bitstream 510 SPS 512 PPS 514 Slice Header 515 SEI Message 520 Image Data 521 Picture 522 Sub-picture 523 Sub-picture 524 Slice 525 Slice 531 Sub-picture Size 532 Sub-picture Position 533 Sub-picture ID 534 Motion Constraint Sub-picture Flag 535 Sub-picture Level 600 Picture 622 Sub-picture 631 Sub-picture Size 631a Sub-picture Width 631b Sub-picture Height 632 Position 633 Sub-picture ID 634 Temporal Motion Constraint Sub-picture 642 Top-Left Sample 700 Mechanism 722 Sub-picture 724 Slice 733 Sub-picture ID 742 Top-Left Corner 801 Right Boundary of Picture 802 Bottom Boundary of Picture 822 Sub-picture 822b, 822c Sub-picture 825 CTU 900 Video Coding Device 910 Transmitter / Receiver 914 Coding Module 920 Downstream Port 930 Processor 932 Memory 950 Upstream Port 960 I / O Device 1200 System 1201 Partition Module 1202 Video Encoder 1203 Encoding Module 1205 Storage Module 1207 Transmission Module 1210 Video Decoder 1211 Reception Module 1213 Analysis Module 1215 Decryption Module 1217 Transfer Module

Claims

1. A method implemented in a decoder, comprising: receiving, by a receiver of the decoder, a bitstream comprising one or more sub-pictures segmented from a picture, wherein when a first sub-picture includes a right boundary that coincides with the right boundary of the picture, the first sub-picture has a sub-picture width comprising an incomplete coding tree unit (CTU); obtaining, by a processor of the decoder, the one or more sub-pictures by analyzing the bitstream; decoding, by the processor, the one or more sub-pictures to create a video sequence; and transferring, by the processor, the video sequence for display.

2. The method according to claim 1, wherein when a second sub-picture includes a bottom boundary that does not coincide with the bottom boundary of the picture, the second sub-picture has a sub-picture height comprising an integer number of complete CTUs.

3. The method according to claim 1 or 2, wherein when a third sub-picture includes a right boundary that does not coincide with the right boundary of the picture, the third sub-picture has a sub-picture width comprising an integer number of complete CTUs.

4. The method according to any one of claims 1 to 3, wherein when a fourth sub-picture includes a bottom boundary that does not coincide with the bottom boundary of the picture, the fourth sub-picture has a sub-picture height comprising an incomplete CTU.

5. A method implemented in a decoder, comprising: receiving, by a receiver of the decoder, a bitstream comprising one or more sub-pictures segmented from a picture, wherein when each sub-picture includes a right boundary that does not coincide with the right boundary of the picture, each sub-picture has a sub-picture width that is an integer multiple of a coding tree unit (CTU) size; obtaining, by a processor of the decoder, the one or more sub-pictures by analyzing the bitstream; decoding, by the processor, the one or more sub-pictures to create a video sequence; and transferring, by the processor, the video sequence for display.

6. The method according to claim 5, wherein when each sub-picture includes a bottom boundary that does not coincide with the bottom boundary of the picture, each sub-picture... ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ The method according to claim 5, comprising a sub-picture height in which the picture is an integer multiple of the CTU size. Method. **Claim 7** When each of the sub-pictures includes a right boundary that coincides with the right boundary of the picture, at least one of the sub-pictures has a sub-picture width that is not an integer multiple of the CTU size. The method according to claim 5 or 6. **Claim 8** When each of the sub-pictures includes a lower boundary that coincides with the lower boundary of the picture, at least one of the sub-pictures has a sub-picture height that is not an integer multiple of the CTU size. The method according to any one of claims 5 to 7. **Claim 9** The method according to any one of claims 5 to 8, wherein the picture includes a picture width that is not an integer multiple of the CTU size. **Claim 10** The method according to any one of claims 5 to 9, wherein the picture includes a picture height that is not an integer multiple of the CTU size. **Claim 11** The method according to any one of claims 5 to 10, wherein the CTU size is measured in units of luma samples. **Claim 12** A method implemented in an encoder, comprising: Dividing the picture into a plurality of sub-pictures by a processor of the encoder such that when each of the sub-pictures includes a right boundary that does not coincide with the right boundary of the picture, each of the sub-pictures includes a sub-picture width that is an integer multiple of a coding tree unit (CTU) size; Encoding one or more of the sub-pictures into a bitstream by the processor; Storing the bitstream in a memory of the encoder for communication to a decoder. **Claim 13** The method according to claim 12, wherein when each of the sub-pictures includes a lower boundary that does not coincide with the lower boundary of the picture, each of the sub-pictures includes a sub-picture height that is an integer multiple of the CTU size. **Claim 14** When each of the sub-pictures includes a right boundary that coincides with the right boundary of the picture, at least one of the sub-pictures includes a sub-picture width that is not an integer multiple of the CTU size. The method according to claim 12 or 13. **Claim 15** When each of the sub-pictures includes a lower boundary that coincides with the lower boundary of the picture, at least one of the sub-pictures includes a sub-picture height that is not an integer multiple of the CTU size. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ The method according to any one of claims 12 to 14, comprising

16. The method according to any one of claims 12 to 15, wherein the picture comprises a picture width that is not an integer multiple of the CTU size.

17. The method according to any one of claims 12 to 16, wherein the picture comprises a picture height that is not an integer multiple of the CTU size.

18. The method according to any one of claims 12 to 17, wherein the CTU size is measured in units of luma samples.

19. A video coding device comprising a processor, a memory, a receiver coupled to the processor, and a transmitter coupled to the processor, wherein the processor, memory, receiver, and transmitter are configured to execute the method according to any one of claims 1 to 18.

20. A non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to execute the method according to any one of claims 1 to 18.

21. Receiving means for receiving a bitstream comprising one or more sub-pictures partitioned from the picture such that each sub-picture has a sub-picture width that is an integer multiple of a coding tree unit (CTU) size when each sub-picture includes a right boundary that does not coincide with the right boundary of the picture; Analyzing means for analyzing the bitstream to obtain the one or more 、 sub-pictures; Decoding means for decoding the one or more sub-pictures to create a video sequence; and Transferring means for transferring the video sequence for display.

22. The decoder according to claim 17, wherein the decoder is further configured to execute the method according to any one of claims 1 to 11.

23. Partitioning means for partitioning the picture into a plurality of sub-pictures such that each sub-picture has a sub-picture width that is an integer multiple of a coding tree unit (CTU) size when each sub-picture includes a right An encoding step for encoding one or more of the sub-pictures into a bit stream and storage means for storing the bit stream for communication to a decoder , an encoder.

24. The encoder according to claim 23, further configured to execute the method according to any one of claims 12 to 18 .

Citation Information

Patent Citations

  • Concept for picture / video data streams allowing efficient reducibility or efficient random access

    WO2017137444A1

  • Advanced video data stream extraction and multi-resolution video transmission

    WO2018172234A2