VIDEO ENCODER, VIDEO DECODER, AND CORRESPONDING METHODS - Patent application

By enforcing subpicture size constraints and signaling layout information in the SPS, the inefficiencies in handling non-CTU-sized picture layouts are addressed, leading to improved encoding and decoding efficiency and reduced resource usage in video coding systems.

JP7813400B2Active Publication Date: 2026-02-12HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025071560
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-01-09
Filing Date
2025-04-23
Publication Date
2026-02-12
Estimated Expiration
2040-01-09

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently handling picture layouts that are not multiples of the Coding Tree Unit (CTU) size, leading to improper functioning of CTUs, increased resource usage, and reduced coding efficiency.

Method used

Implementing subpicture size constraints that ensure widths and heights are multiples of the CTU size, allowing for efficient encoding and decoding by restricting subpictures to these dimensions, especially at picture borders, and signaling layout information in the Sequence Parameter Set (SPS) to reduce redundancy and resource usage.

Benefits of technology

This approach enhances encoder and decoder capabilities, reduces resource consumption, and improves coding efficiency by allowing subpictures to be used with any picture layout, minimizing network, memory, and processing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007813400000016
    Figure 0007813400000016
  • Figure 0007813400000017
    Figure 0007813400000017
  • Figure 0007813400000018
    Figure 0007813400000018
Patent Text Reader

Abstract

To provide a decoder, an encoder, a method, and a non-temporary computer readable medium which improve a compression ratio with little or no sacrifice in image quality.SOLUTION: A method of decoding a bit stream includes: receiving 1101 a bitstream comprising one or more sub-pictures partitioned from a picture, each sub-picture including a sub-picture that is an integer multiple of a coding tree unit (CTU) size when each sub-picture includes a right border that does not coincide with picture's right border; analyzing the bitstream to obtain 1103 the one or more sub-pictures; decoding the one or more sub-pictures to create a video sequence; and forwarding 1105 the video sequence for display.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0002] TECHNICAL FIELD This disclosure relates generally to video coding, and more particularly to subroutines in video coding. Regarding picture management. [Background technology]

[0003] The amount of video data required to render even a relatively short video is quite large. This is because data is streamed over a communications network with limited bandwidth capacity. This can cause difficulties when the data is to be sent or otherwise communicated. Video data is typically compressed before being transmitted over modern telecommunications networks. Memory resources may be limited, so if the video is stored on a storage device, The size of the video may also be an issue when it is compressed. Video compression devices often Use the software and / or hardware at the source to transmit or store to code video data beforehand to represent a digital video image This reduces the amount of data required to decode the video data. Received at the destination by a video decompression device. Limited resources and ever-increasing demand for higher video quality This improves the compression ratio with little or no sacrifice in image quality. A well-defined compression and decompression technique is desirable. Summary of the Invention [Means for solving the problem]

[0004] In an embodiment, the present disclosure includes a method implemented in a decoder, the method comprising: The decoder receiver determines whether the first subpicture is positioned at the right border of the picture. When the first subpicture contains an incomplete coding tree unit (CTU), One or more sub-pictures separated from the picture, each having a width of one picture. receiving a bitstream; and decoding the bitstream by a processor of the decoder. analyzing the frame to obtain one or more sub-pictures; decoding the sub-pictures to create a video sequence; and transferring the video sequence for display. The system may restrict subpictures to contain heights and widths that are multiples of the CTU size. However, a picture may contain heights and widths that are not multiples of the CTU size. Therefore, the subpicture size constraints are applicable to many picture layouts. This prevents the CTU from working properly. In the example shown, , the subpicture width and subpicture height are constrained. These constraints are If the bottom and right subpictures have heights and widths that are not multiples of the CTU size, they will be removed. By allowing the inclusion of any picture, Subpictures may also be used for both the video and audio streams. This allows for improved encoder and decoder capabilities. Additionally, the improvements allow the encoder to code pictures more efficiently. This allows for network resources, memory resources, and / or Reduces the use of processing resources in the encoder and decoder.

[0005] Optionally, in any of the preceding aspects, another implementation of the aspect is When a picture contains a bottom border that does not coincide with the bottom border of the picture, the second subpicture is aligned. It specifies having a sub-picture height that contains several complete CTUs.

[0006] Optionally, in any of the preceding aspects, another implementation of the aspect further comprises: When a picture contains a right border that does not coincide with the right border of the picture, the third subpicture is aligned. It specifies having a sub-picture width that contains several complete CTUs.

[0007] Optionally, in any of the preceding aspects, another implementation of the aspect further comprises: The fourth subpicture is incomplete when the picture contains a bottom border that coincides with the bottom border of the picture. It specifies that the subpicture height includes the entire CTU.

[0008] In an embodiment, the present disclosure includes a method implemented in a decoder, the method comprising: The decoder receiver may detect that each subpicture has a right boundary that does not coincide with the right boundary of the picture. When the subpicture size is an integer multiple of the coding tree unit (CTU) size, With one or more sub-pictures segmented from the picture, such that the sub-picture width is included receiving a bitstream; and decoding the bitstream by a processor of the decoder. analyzing the frame to obtain one or more sub-pictures; decoding the sub-pictures to create a video sequence; and transferring the video sequence for display. The system may restrict subpictures to contain heights and widths that are multiples of the CTU size. However, a picture may contain heights and widths that are not multiples of the CTU size. Therefore, the subpicture size constraints are applicable to many picture layouts. This prevents the CTU from working properly. In the example shown, , the subpicture width and subpicture height are constrained. These constraints are If the bottom and right subpictures have heights and widths that are not multiples of the CTU size, they will be removed. By allowing the inclusion of any picture, Subpictures may also be used for both the video and audio streams. This allows for improved encoder and decoder capabilities. Additionally, the improvements allow the encoder to code pictures more efficiently. This allows for network resources, memory resources, and / or Reduces the use of processing resources in the encoder and decoder.

[0009] Optionally, in any of the preceding aspects, another implementation of the aspect is When a subpicture contains a bottom border that does not coincide with the bottom border of the picture, each subpicture is It specifies that the subpicture height must be an integer multiple of

[0010] Optionally, in any of the preceding aspects, another implementation of the aspect is When a subpicture contains a right border that coincides with the right border of the picture, at least one of the subpictures It also specifies that one of the subpicture widths is not an integer multiple of the CTU size.

[0011] Optionally, in any of the preceding aspects, another implementation of the aspect is When a subpicture contains a bottom border that coincides with the bottom border of the picture, at least one of the subpictures Also, it specifies that one of the subpictures may contain a subpicture height that is not an integer multiple of the CTU size.

[0012] Optionally, in any of the preceding aspects, another implementation of the aspect is further characterized in that the picture is a CT It specifies that picture widths that are not integer multiples of the U size are included.

[0013] Optionally, in any of the preceding aspects, another implementation of the aspect is further characterized in that the picture is a CT It specifies that picture heights that are not integer multiples of the U size are included.

[0014] Optionally, in any of the preceding aspects, another implementation of the aspect is It specifies that it is measured in units of luma samples.

[0015] In an embodiment, the present disclosure includes a method implemented in an encoder, the method comprising: ,The encoder's processor detects that each subpicture does not match the right boundary of the picture. When including the right border, each subpicture contains a subpicture width that is an integer multiple of the CTU size. partitioning the picture into a plurality of sub-pictures, such as encoding one or more of the sub-pictures into a bitstream; and storing the bitstream in a memory of the decoder for communication to the decoder. Some video systems are sub-pistoned to include heights and widths that are multiples of the CTU size. However, pictures may be limited in size and height that are not multiples of the CTU size. Therefore, the subpicture size constraints may include the number of picture layers. This prevents the subpicture from working properly with the CTU size. The subpicture width and height are constrained to be multiples of the size. while the subpicture is located at the right border of the picture or the bottom border of the picture, respectively. These constraints are removed when the bottom and right subpictures are not multiples of the CTU size. By allowing the height and width to be included separately, it is possible to Subpictures may be used with any picture. Furthermore, the improved functionality allows the encoder to process the picture more efficiently. This allows for efficient coding, saving network and memory resources. Reduce the use of bandwidth and / or processing resources in the encoder and decoder.

[0016] Optionally, in any of the preceding aspects, another implementation of the aspect is When a subpicture contains a bottom border that does not coincide with the bottom border of the picture, each subpicture is It specifies that the subpicture height must be an integer multiple of

[0017] Optionally, in any of the preceding aspects, another implementation of the aspect is When a subpicture contains a right border that coincides with the right border of the picture, at least one of the subpictures It also specifies that one of the subpicture widths is not an integer multiple of the CTU size.

[0018] Optionally, in any of the preceding aspects, another implementation of the aspect is When a subpicture contains a bottom border that coincides with the bottom border of the picture, at least one of the subpictures Also, it specifies that one of the subpictures may contain a subpicture height that is not an integer multiple of the CTU size.

[0019] Optionally, in any of the preceding aspects, another implementation of the aspect is further characterized in that the picture is a CT It specifies that picture widths that are not integer multiples of the U size are included.

[0020] Optionally, in any of the preceding aspects, another implementation of the aspect is further characterized in that the picture is a CT It specifies that picture heights that are not integer multiples of the U size are included.

[0021] Optionally, in any of the preceding aspects, another implementation of the aspect is It specifies that it is measured in units of luma samples.

[0022] In one embodiment, the present disclosure provides a method for implementing a multi-threaded ... signal processing system, comprising: a processor; a memory; and a receiver coupled to the processor. a video coding device comprising a processor; and a transmitter coupled to the processor; The processor, memory, receiver, and transmitter are adapted to perform the method of any of the preceding aspects. It is configured to:

[0023] In some embodiments, the present disclosure provides a computer program for use by a video coding device. a non-transitory computer-readable medium having a computer program product thereon, The program product, when executed by a processor, precedes a video coding device. A computer program stored on a non-transitory computer-readable medium that performs any of the methods of the present invention. It comprises computer-executable instructions.

[0024] In some embodiments, the present disclosure provides a method for creating a right sub-picture that does not coincide with the right border of the picture. When including borders, each subpicture must contain a subpicture width that is an integer multiple of the CTU size. receiving a bitstream comprising one or more sub-pictures partitioned from a picture, receiving means for parsing the bitstream to obtain one or more sub-pictures; and a decoding means for decoding one or more sub-pictures to create a video sequence. a decoding means for decoding a video sequence and a transmitting means for transmitting the video sequence for display; Includes the driver.

[0025] Optionally, in any of the preceding aspects, another implementation of the aspect is It is provided that the method is further configured to perform any of the methods of the aspects described above.

[0026] In some embodiments, the present disclosure provides a method for creating a right sub-picture that does not coincide with the right border of the picture. When including borders, each subpicture must contain a subpicture width that is an integer multiple of the CTU size. a dividing means for dividing the picture into a plurality of sub-pictures; encoding means for encoding one or more of the above into a bitstream; and and storage means for storing the bitstream for encoding the encoded data.

[0027] Optionally, in any of the preceding aspects, another implementation of the aspect is further characterized in that the encoder It is provided that the device is further configured to perform the method of any of the preceding aspects.

[0028] For clarity, any one of the above embodiments may be considered a new implementation within the scope of this disclosure. The present invention may be combined with any one or more of the other aforementioned embodiments to create an embodiment.

[0029] These and other features will become more apparent from the following detailed description taken in conjunction with the accompanying drawings and claims. This will be more clearly understood.

[0030] For a more complete understanding of the present disclosure, reference is now made to the accompanying drawings and detailed description, in which: Reference is made to the following brief description, in which like reference numerals represent like parts. [Brief explanation of the drawings]

[0031] [Figure 1] 1 is a flowchart of an exemplary method for coding a video signal. [Figure 2] 1 is a schematic diagram of an example encoding and decoding (codec) system for video coding. [Figure 3] FIG. 1 is a schematic diagram illustrating an exemplary video encoder. [Figure 4] FIG. 1 is a schematic diagram illustrating an exemplary video decoder. [Figure 5] FIG. 2 is a schematic diagram illustrating an exemplary bitstream and sub-bitstreams extracted from the bitstream. [Figure 6] FIG. 2 is a schematic diagram illustrating an exemplary picture partitioned into sub-pictures. [Figure 7] FIG. 1 is a schematic diagram illustrating an exemplary mechanism for associating slices with subpicture layouts. [Figure 8] FIG. 10 is a schematic diagram illustrating another exemplary picture partitioned into sub-pictures. [Figure 9] 1 is a schematic diagram of an exemplary video coding device. [Figure 10] 1 is a flowchart of an exemplary method for encoding a sub-picture bitstream with adaptive size constraints. [Figure 11] 1 is a flowchart of an exemplary method for decoding a bitstream of sub-pictures with adaptive size constraints. [Figure 12]FIG. 1 is a schematic diagram of an example system for signaling a sub-picture bitstream with adaptive size constraints. DETAILED DESCRIPTION OF THE INVENTION

[0032] Illustrative implementations of one or more embodiments are provided below, but the disclosed system Any system and / or method, whether currently known or existing, It should be understood at the outset that this disclosure may be implemented using any number of techniques. including the example designs and implementations illustrated and described herein. The illustrative implementations, diagrams, and techniques illustrated below should not be considered limiting. Modifications may be made within the scope of the appended claims, along with their full range of equivalents.

[0033] Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coding Video Sequence (CVS), Joint Video Expert Team JVET, Motion Constrained Tile Set (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Level Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sequence Sequence Parameter Set (SPS), Versatile Video Coding (VVC), and Working Various acronyms are utilized herein, such as Draft (WD).

[0034] Many video formats are used to reduce the size of video files while minimizing data loss. For example, video compression techniques can be used to compress the spatial (e.g., image) performing intrapicture prediction and / or temporal (e.g., interpicture) prediction, This can include reducing or removing data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video picture) A video image (or a portion of a video picture) may be partitioned into video blocks, which Tree Block, Coding Tree Block (CTB), Coding Tree Unit (CTU), They may also be called coding units (CUs) and / or coding nodes. Video blocks within an intra-coded (I) slice of a picture are The image is coded using spatial prediction with respect to reference samples in neighboring blocks in the image. Inter-coded unidirectionally predicted (P) or bidirectionally predicted (B) slices of a picture The video blocks in are spatially related to reference samples in neighboring blocks of the same picture. prediction, or temporal prediction with respect to reference samples in other reference pictures. A picture may be called a frame and / or an image. Instead, a reference picture may also be called a reference frame and / or a reference image. The residual data is the sum of the original image data and the residual block. It represents the pixel difference between the block and the predicted block. The predicted block is a motion vector pointing to a block of reference samples that form the predicted block. and residual data indicating the difference between the coded block and the predicted block. The intra-coded blocks are coded according to the intra-coding model. The residual data is coded according to the pixel code and the residual data. For further compression, the residual data is These may be transferred from the cell domain to the transform domain, where they are the residual transform coefficients, which may be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The resulting transform coefficients may be scanned to produce a one-dimensional vector of transform coefficients. Tropic coding may be applied to achieve further compression. Video compression techniques are discussed in more detail below.

[0035] To ensure that the encoded video can be correctly decoded, the corresponding video codec Video is encoded and decoded according to a video coding standard. International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC Advanced Video Coding (AVC), also known as MPEG-4 Part 10, and IT High Efficiency Video Coding (HEVC), also known as UT H.265 or MPEG-H Part 2 AVC includes Scalable Video Coding (SVC), Multiview Video Coding (MVC), VC) and Multiview Video Coding plus Depth (MVC+D), as well as three-dimensional (3D) AV C (3D-AVC). HEVC is a standard for scalable HEVC (SHVC), multi-view HEVC (MV-HEVC), and other extensions. C), and extensions such as 3D HEVC (3D-HEVC). Joint Video The JVET Expert Team is developing a new video coding technology called Versatile Video Coding (VVC). The development of a coding standard has begun. VVC is included in the Working Draft (WD), and This includes JVET-L1001-v9.

[0036] To code a video image, the image is first partitioned and the partitions are converted into a bitstream Various picture partitioning schemes are available. For example, an image is coded as , into ordinary slices, dependent slices, tiles, and / or using wavefront parallel processing (WPP). For simplicity, we will divide the CTB into groups for video coding. When dividing slices, there are normal slices, subordinate slices, tiles, WPPs, and HEVC constrains the encoder so that a combination of these may be used. Supports maximum transmission unit (MTU) size adaptation, parallel processing, and reduced end-to-end latency The MTU is the maximum amount of data that can be sent in a single packet. If a packet payload exceeds the MTU, the payload is called fragmented. The data is split into two packets through a process called

[0037] Ordinary slicing, also known simply as slicing, is a loop filtering operation. Reconstructed independently of other ordinary slices in the same picture, despite their degree of interdependence Each regular slice is assigned a unique network address for transmission. It is encapsulated in Network Abstraction Layer (NAL) units. (intra-sample prediction, motion information prediction, coding mode prediction) and slice boundaries The entropy coding dependency is disabled to support independent reconstruction. Such independent reconstructions support parallelization. For example, Rice-based parallelism utilizes minimal inter-processor or inter-core communication. However, each normal slice is independent, so each slice has a separate slice header. The normal use of slices is to associate the slice header bits with each slice. The cost of coding and the lack of prediction across slice boundaries results in significant coding overhead. Furthermore, regular slices may require more memory to meet MTU size requirements. Specifically, a normal slice is a separate NAL unit. Since they are encapsulated in a single slice and can be coded independently, each ordinary slice is To avoid breaking slices into multiple packets, the MTU must be smaller than the MTU in the MTU scheme. Therefore, the goal of parallelization and the goal of adapting the MTU size are These may impose conflicting requirements on the slice layout in the

[0038] A dependent slice is similar to a normal slice, but has an abbreviated slice header and This allows for the division of image tree block boundaries without breaking intra-channel prediction. Generic slices allow a regular slice to be fragmented into multiple NAL units, This is because part of a normal slice is sent before the entire normal slice has finished encoding. This results in a reduction in end-to-end delay by enabling

[0039] A tile is an image created by horizontal and vertical boundaries that create columns and rows of tiles. The tiles are coded in raster scan order (right to left and top to bottom). The scan order of the CTB is local within a tile. The CTB of a tile is coded in raster scan order before proceeding to the CTB of the next tile. Similar to slices, tiles have dependencies for intra-picture prediction as well as for entropy decoding. However, tiles may not be contained in individual NAL units. In some cases, tiles may not be used due to MTU size compatibility. Each tile is assigned to one process. The pixels can be processed by the processor / core, and the pixels between the processing units that decode neighboring tiles are The inter-processor / inter-core communication used for intra-chamber prediction is a shared slice header. (when adjacent tiles are in the same slice), and the reconstructed slice. It may be limited to performing loop filtering-related sharing of samples and metadata. When more than one tile is contained in a slice, the first entry point in the slice is The entry point byte offset for each tile other than the entry point offset is For each slice and tile, 1) the slice 1) All coded treeblocks in a slice belong to the same slice, and 2) The condition that all coded blocks in a tile belong to the same slice At least one of the following should be satisfied.

[0040] In WPP, the image is partitioned into a single row of CTB. The entropy decoding and prediction mechanism Data from CTBs in other rows may be used. Parallel processing is achieved through parallel decoding of CTB rows. For example, the current row may be decoded in parallel with the previous row. However, the decoding of the current row lags behind the decoding process of the preceding row by the two CTBs. The delay causes the data for the CTB above and to the right of the current CTB on the current line to be This ensures that data will be available before the current CTB is coded. The law appears as a wavefront when graphically represented. This offset onset can be up to It allows parallelization using as many processors / cores as the number of CTB rows the image contains. Since intra-picture prediction between neighboring treeblock rows in a pixel is allowed, we can use intra-picture prediction. There can be a lot of inter-processor / inter-core communication to make this possible. Therefore, WPP does not support MTU size adaptation. However, some coding overhead is required to accommodate MTU size on demand. With the -head, regular slices can be used with WPP.

[0041] The tiles may also include motion constrained tile sets. A motion constrained tile set (MCTS) is a set of associated The motion vectors are integer sample positions inside the MCTS and integer positions inside the MCTS for interpolation. The sample position is constrained to point to a fractional sample position, which requires only the sample position. Furthermore, the time motion vector derived from the external block of MCTS is The use of motion vector candidates for vector prediction is not allowed. In this way, each MCTS , tiles not included in the MCTS may be decoded independently. The Enhancement Information (SEI) message indicates the presence of an MCTS in the bitstream and signals the MCTS. The MCTS SEI message may be used to nullify the conformance to the MCTS set. To generate the bitstream, MCTS sub-bitstream extraction (SEI message) provides additional information that may be used in the context of The information includes several extracted information sets, each of which defines a number of MCTS sets and A replacement video parameter set (VPS) to be used during the bitstream extraction process. ), Sequence Parameter Set (SPS), and Picture Parameter Set (PPS) MCTS Sub-Bitstream Extraction Process When extracting the sub-bitstream according to the parameter sets (VPS, SPS, and PPS ) may be rewritten or replaced, slice headers may be updated, It is the slice address related syntax elements (first_slice_segment_in_pic_flag and One or all of the following (including slice_segment_address and slice_segment_address) are extracted sub-bits: This is because different values ​​may be used in the stream.

[0042] A picture may also be divided into one or more sub-pictures. A sub-picture is equal to 0. The length of the tile group / slice, starting with the tile group with tile_group_address. Each subpicture may refer to a different PPS, and therefore a separate tile. Sub-pictures are treated like pictures in the decoding process. The reference sub-picture for decoding the current sub-picture may be stored in the decoded picture buffer. Extracting an area in the image that is at the same position as the current sub-picture from the reference picture The extracted area is treated as a decoded sub-picture. Inter prediction is performed between a sub-picture of the same size and the same location within a picture. A tile group, also known as a slice, is a portion of a picture or subpicture. A sequence of related tiles in a picture. Determines the position of a subpicture within a picture. For example, each current subpicture is a picture CTU raster run within a picture large enough to fit the current subpicture within its boundaries. may be placed in the next unoccupied position in the scanning order.

[0043] Furthermore, picture partitioning is based on picture-level tiles and sequence-level tiles. Sequence level tiles may contain MCTS functionality and may be used as subpictures. For example, picture-level tiling may be implemented as a specific tile within a picture. It is defined as the rectangular region of the coding tree block within a column and within a particular tile row. Sequence-level tiles may be used to represent coding tree fragments contained in different frames. A set of rectangular regions of a block may be defined, each of which may further define one or more With picture-level tiles, the set of rectangular regions of the coding tree block is similar Sequence-level tiles are decodable independently from any other set of rectangular regions. A group set (STGPS) is a group of such sequence-level tiles. The NAL unit header contains the non-video coding layer along with its associated identifier (ID). It may be signaled in the VCL NAL unit.

[0044] Previous sub-picture based partitioning schemes may be associated with certain problems. For example, when subpictures are enabled, tiling within the subpicture (subpicture to tile) A partition of the architecture can be used to support parallel processing. The tile partitioning of a picture can vary from picture to picture (e.g., depending on the balance of parallel processing loads). for the purposes of interpretation), and therefore even if it is managed at the picture level (e.g. in PPS) However, subpicture partitioning (partitioning a picture into subpictures) is not possible in areas of interest. It is also used to support ROI and subpicture-based picture access. In such cases, signaling sub-pictures or MCTS in the PPS is not efficient. There is no.

[0045] In another example, any sub-picture in a picture can be coded as a temporal motion constrained sub-picture. When a picture is loaded, all subpictures in the picture are temporal motion constrained subpictures. Such picture partitioning may be restrictive. For example, coding a subpicture as a temporal motion constrained subpicture: Additional functionality may be traded for reduced coding efficiency. In the base application, typically only one or a few of the sub-pictures are subject to the temporal motion constraints. Uses subpicture-based functionality, so the remaining subpictures are It incurs a loss in coding efficiency without providing any benefit.

[0046] In another example, the syntax element for specifying the size of a subpicture is the luma CTU subpicture. Therefore, both the width and height of the subpicture may be specified in units of CtbSiz This mechanism for specifying the subpicture width and height solves a variety of problems. For example, subpicture partitioning may cause problems if the subpicture partitioning is an integer multiple of CtbSizeY. This is only applicable to pictures with subpicture width and / or picture height. Disables picture partitioning for pictures with dimensions that are not an integral multiple of CTbSizeY. When the picture dimensions are not an integer multiple of CtbSizeY, the subpicture partitions are and / or height, the rightmost and bottommost subpictures Derivation of subpicture width and / or subpicture height in luma samples for In some coding tools, such an inaccurate derivation may lead to incorrect results. Brings fruit.

[0047] In another example, the position of the sub-picture in the picture may not be signaled. Instead, the position is derived using the following rules: The current subpicture is the picture Next in CTU raster scan order within a picture large enough to fit the subpicture within its boundaries In some cases, the sub Deriving the sub-picture position may lead to errors. If a sub-picture is lost in transmission, then the positions of other sub-pictures will be derived incorrectly, Decoded samples are misplaced when subpictures arrive out of order. , the same issues apply.

[0048] In another example, decoding a sub-picture may involve decoding a sub-picture at the same location in a reference picture. This may require the extraction of a sub-picture, which can be expensive in terms of processor and memory resource usage. This can impose additional complexity and resulting burden on the

[0049] In another example, when a subpicture is designated as a temporal motion constrained subpicture, The loop filter that scans tile boundaries is disabled. This occurs whether or not group filters are enabled. Such restrictions are This can be a visual artifact for video pictures that utilize multiple subpictures. This can result in artifacts.

[0050] In another example, the relationship between SPS, STGPS, PPS, and Tile Group Header is as follows: STGPS refers to SPS, PPS refers to STGPS, and tile group header / slice header However, the STGPS and the PPS are not related to the PPS, but rather to the STGPS. , must be orthogonal. The above configuration ensures that all tile groups of the same picture It may also not be acceptable for multiple PPSs to refer to the same PPS.

[0051] In another example, each STGPS may contain IDs for the four sides of the subpicture. Uses unique IDs to identify subpictures that share the same boundary and determines their relative spatial relationship. However, in some cases, such information may be defined as a sequence-level tag. It may not be sufficient to derive position and size information for a set of file groups. In other cases, signaling location and size information may be redundant. do.

[0052] In another example, the STGPS ID is stored in the NAL unit header of a VCL NAL unit using 8 bits. This may aid in sub-picture extraction. Signaling may unnecessarily increase the length of the NAL unit header. Unless sequence level tile group sets are constrained to prevent overlaps, A tile group may be associated with multiple sequence-level tile group sets. That is what it means.

[0053] To address one or more of the above-mentioned problems, various mechanisms are disclosed herein. In the first example, the layout information for the subpictures is included in the SPS instead of the PPS. The picture layout information includes the subpicture position and the subpicture size. The pixel position is the offset between the top left sample of the subpicture and the top left sample of the picture. The subpicture size is the height and width of the subpicture as measured in luma samples. As mentioned above, tiles may change from picture to picture, so some systems The system includes tiling information in the PPS. However, ROI application and subpicture-based Subpictures may be used to support access to: It does not change from picture to picture. Furthermore, a video sequence is a single SPS (or video segment). It may contain a PPS (one per frame) and may contain as many as one PPS per picture. Placing the layout information of the sub-pictures means that the layout is redundantly signed for each PPS. Signaled only once per sequence / segment instead of being nulled Therefore, signaling the subpicture layout in the SPS This increases coding efficiency, saving network resources, memory resources, and and / or reduce the use of processing resources in the encoder and decoder. The system has sub-picture information derived by the decoder. Signaling reduces the probability of error when a packet is lost and Therefore, SPS supports additional functionality for extracting subpictures. Signaling the channel layout is a way to control the functionality of the encoder and / or decoder. Increase.

[0054] In the second example, the subpicture width and subpicture height are constrained to be multiples of the CTU size. However, these constraints are not enforced if the subpicture is positioned on the right border of the picture or on the left border of the picture. As mentioned above, some The video system resizes the subpicture to contain heights and widths that are multiples of the CTU size. This allows subpictures to work correctly with many picture layouts. The bottom and right subpictures have heights and widths that are not multiples of the CTU size. By allowing the inclusion of sub-pictures, the sub-pictures can be decoded without causing decoding errors. It may be used with any picture. This is a function of the encoder and decoder functions. Additionally, the improved functionality allows the encoder to code pictures more efficiently. This allows for network linking at the encoder and decoder. Reduce the use of sources, memory resources, and / or processing resources.

[0055] In the third example, the subpicture is constrained to encompass the picture without gaps or overlaps. As mentioned above, some video coding systems use the Allows for gaps and overlaps, which means that a tile group / slice can have multiple sub-tiles. This creates the possibility of associating a picture with the image, if this is allowed by the encoder. , the decoder will be able to use such a codec even when the decoding scheme is rarely used. Subpicture gaps and overlaps must be supported. By not allowing multiple subpictures, the decoder avoids latency when determining the size and position of the subpicture. Reduces decoder complexity since potential gaps and overlaps do not need to be considered. Furthermore, by not allowing gaps and overlaps between subpictures, the video sequence can be This allows the encoder to avoid considering gap and overlap cases when selecting the encoding for a sequence. This reduces the complexity of the rate-distortion optimization (RDO) process in the encoder. Therefore, avoiding gaps and overlaps reduces memory usage in the encoder and decoder. May reduce resource and / or processing resource usage.

[0056] In the fourth example, a frame is added to indicate when a subpicture is a temporal motion constrained subpicture. The lag may be signaled in the SPS. As mentioned above, some systems , by setting all subpictures together as temporal motion constrained subpictures, or However, it may be necessary to completely disallow the use of temporal motion constrained sub-pictures. Subpictures provide independent extraction capabilities at the expense of reduced coding efficiency. However, in region-of-interest based applications, the regions of interest are coded for independent extraction. The area outside the region of interest does not require such functionality. The remaining subpictures are then coded in a manner that reduces coding efficiency without providing any practical benefit. Therefore, this flag is used to improve coding efficiency when independent extraction is not desired. To enhance the image quality, temporal motion constrained and non-motion constrained subpictures are provided with independent extraction functions. Therefore, this flag is used to improve functionality and / or code This allows for improved packet efficiency, saving network resources, memory resources, and and / or reduce the use of processing resources in the encoder and decoder.

[0057] In the fifth example, the complete set of subpicture IDs is signaled in the SPS and The slice header contains a subpicture ID that indicates the subpicture that contains the corresponding slice. As mentioned, some systems may use a picture position relative to other sub-pictures. This is the case when sub-pictures are lost or extracted separately. This can cause problems if you specify each subpicture by its ID. It can be positioned and sized without reference to other subpictures. This allows for error correction and extracting only a portion of a subpicture to avoid transmitting other subpictures. A complete list of all subpicture IDs is provided for each associated subpicture. Each slice header can be transmitted in the SPS along with the slice size information. A subpicture ID indicating the subpicture it contains may also be included. and the corresponding slice is extracted and positioned without reference to other subpictures. Therefore, subpicture IDs can provide improved functionality and / or coding efficiency. This may be due to network resources, memory resources, and / or encoder and decoder Reduces the use of processing resources in the coder.

[0058] In the sixth example, a level is signaled for each subpicture. In a coding system, levels are signaled for pictures. This indicates the hardware resources required to decode the image. Additionally, in some cases, different sub-pictures may have different functions, so coding may be treated differently during the rendering process. Thus, the picture-based levels Therefore, this disclosure provides a method for decoding each sub-frame. In this way, each sub-picture contains a level for the setting decoding requirements too high for subpictures coded according to a different scheme This allows the subpictures to be displayed independently of other subpictures without unnecessarily burdening the decoder. The signaled sub-picture level information allows for enhanced functionality and and / or improve coding efficiency, which reduces network resources, memory resources, Reduce the use of processing resources at the source and / or encoder and decoder.

[0059] FIG. 1 is a flow chart of an exemplary operational method 100 of coding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process includes the following steps: Compress the video signal by utilizing various mechanisms to reduce the video file size. The smaller the file size, the more compressed the video file will be sent to the user. The decoder then decodes the compressed data while reducing the associated bandwidth overhead. Decodes the compressed video file and restores the original video signal for display to the end user. The decoding process generally allows the decoder to reconstruct the video signal reliably. The encoding process is mirrored to allow for

[0060] In step 101, a video signal is input to an encoder. For example, may be an uncompressed video file stored in memory. ,Video files are captured by video capture devices such as video cameras, It may be encoded to support live streaming of video. A file may contain both audio and video components. The video components are viewed sequentially. It contains a sequence of image frames that, when viewed together, give the visual effect of motion. The frames are separated by a luma component (or A pixel is represented in terms of light, referred to herein as a luma sample (or luma sample), and a chrominance component. Contains pixels that are described in terms of colors called fractions (or color samples). In an example, the frames may also include depth values ​​to support three-dimensional viewing.

[0061] In step 103, the video is partitioned into blocks. The partitions are made up of the pictures of each frame. including subdividing cells into square and / or rectangular blocks for compression For example, High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2) In a coding tree, a frame can first be divided into coding tree units (CTUs). The CTU is divided into blocks of a predetermined size (e.g., 64 pixels by 64 pixels). A CTU contains both luma and chroma samples. The coding tree is into blocks and then divide the blocks until a configuration is achieved that supports further encoding. For example, the luma component of a frame may be , may be subdivided until each block contains a relatively uniform illumination value. The chroma components of the image may be subdivided until each block contains relatively uniform color values. Therefore, the segmentation scheme varies depending on the content of the video frame.

[0062] In step 105, the image block partitioned in step 103 is compressed. Various compression mechanisms are used, for example, inter-prediction and / or intra-prediction. Inter prediction is used to predict when objects in a common scene appear in consecutive frames. It is designed to take advantage of the fact that people tend to The blocks depicting the object do not need to be described repeatedly in adjacent frames. In practice, an object such as a table may remain in a constant position across multiple frames. Thus, the table is written once and adjacent frames refer back to the reference frame. To match objects across multiple frames, a pattern matching mechanism can be used. Furthermore, a moving object may be captured by, for example, the object's motion or the camera's motion. A specific example is a video may show a car moving around the screen over multiple frames. may be used to describe such motion. A motion vector is a vector that represents the time A two-dimensional vector that gives the offset from the coordinates of the object in the reference frame to the coordinates of the object in the reference frame. Therefore, inter prediction is based on the corresponding block in the reference frame. We represent the image blocks in the current frame as a set of motion vectors that indicate the offsets of It can be encoded.

[0063] Intra prediction encodes blocks within a common frame. It takes advantage of the fact that the chromatic and chromatic components tend to be densely packed in a frame. For example, green spots in a piece of wood tend to be positioned next to similar green spots. Intra prediction supports multiple directional prediction modes (e.g., 33 in HEVC), planar modes, and Direct Current (DC) mode. Directional mode uses the current block in the corresponding direction. The planar mode indicates that the sample in the neighboring block is similar / same as the sample in the row / column ( For example, a series of blocks along a plane are interpolated based on neighboring blocks at the ends of the rows. The planar mode is essentially a relatively constant gradient by varying the value. By utilizing the DC mode, it shows smooth transitions of light / color across rows / columns. All neighbors associated with the angular direction of the directional prediction mode are used for field smoothing. Indicates that the block is similar / same as the mean value associated with the block's samples. Therefore, intra-predicted blocks are coded as various related prediction mode values ​​instead of actual values. Furthermore, inter-predicted blocks can be represented by a In either case, the predicted block can be represented as a motion vector value. The blocks may not represent image blocks exactly in some cases. The residual blocks are then transformed to compress the file further. may be used.

[0064] Various filtering techniques may be applied in step 107. In HEVC, The filters are applied according to the in-loop filtering method discussed above. Prediction of the block may result in the creation of blocky images in the decoder. Block-based prediction schemes encode a block and then use it later as a reference block. The in-loop filtering scheme may reconstruct the coded blocks to Noise suppression filter, deblocking filter, adaptive loop filter, and sample adaptation Applying a set of offset (SAO) filters to a block / frame repeatedly. These filters are Such blocking artefacts are used to ensure that the encoded file can be accurately reconstructed. Furthermore, these filters reduce the artifacts in the reconstructed reference blocks. Reduce artifacts so that the artifacts are based on the reconstructed reference block. This can lead to additional artifacts in subsequent blocks that are coded based on It will be lower.

[0065] Once the video signal has been segmented, compressed and filtered, in step 109: The resulting data is encoded in a bitstream, which is the same as the bitstream discussed above. to support the encoded data and reconstruction of the appropriate video signal at the decoder. For example, such data may be data, prediction data, residual blocks, and coding instructions to the decoder. The bitstream may contain flags for transmission to a decoder on demand. The bitstream may also be broadcast to multiple decoders. The creation of the bitstream may be iterative and / or multicast. Therefore, steps 101, 103, 105, 107, and 109 are performed to The time series may occur sequentially and / or simultaneously across multiple time series and blocks. The order in which the video is recorded is presented for clarity and ease of discussion. It is not intended to limit the coding process to any particular order.

[0066] In step 111, the decoder receives the bitstream and begins the decoding process. Specifically, the decoder uses an entropy decoding method to decode the bitstream. In step 111, the decoding is performed by converting the received data into the corresponding syntax and video data. The decoder uses syntax data from the bitstream to determine the division into frames. This division must match the block division result in step 103. The entropy encoding / decoding as utilized in step 111 is not described herein. The encoder makes several possible choices based on the spatial arrangement of values ​​in the input image. Many choices are made during the compression process, such as choosing from a range of block partitioning schemes The signaling of strict selection may utilize multiple bins. Herein, a bin is defined as: It is a binary value (e.g., a bit value that may change depending on the situation) that is treated as a variable. Entropy coding is a method for avoiding any choice that is not clearly feasible for a particular case. Allows the encoder to discard choices, leaving a set of acceptable options. Each allowable choice is assigned a codeword. The length of the codeword is determined by the length of the allowable choice. based on the number of choices (e.g., one bin for two choices, one for three to four choices, (For example, two bins). The encoder then encodes the codeword for the selected choice. This scheme reduces the size of the codeword, which allows for a large probability of all possible choices. Rather than uniquely representing a choice from a set of The decoder is able to decode the codewords in a way that makes them as large as desired to uniquely represent a selection from is then calculated by determining the set of allowable choices in a manner similar to the encoder. , and decodes the selection. By determining the set of allowable choices, the decoder The code word can be read to determine the selection made by the encoder.

[0067] In step 113, the decoder performs block decoding. Specifically, the decoder , and uses the inverse transform to generate a residual block. The decoder then The corresponding predicted block is used to reconstruct the image block according to the partition. The block is an intra-predicted block as generated by the encoder in step 105 and an intra-predicted block. The reconstructed image block may then include both a super-prediction block and a super-prediction block. The segment data determined in 111 is positioned into the frames of the reconstructed video signal. The syntax for step 113 is also as discussed above. It may be signaled in the bitstream via bitrate coding.

[0068] In step 115, the encoder reconstructs the image in a similar manner to step 107. Filtering is performed on frames of the video signal. For example, noise suppression filters The blocking filter, deblocking filter, adaptive loop filter, and SAO filter may be applied to the frame to remove filtering artifacts. Once filtered, the video signal is sent to step 117 for viewing by the end user. The image data can be output to a display.

[0069] FIG. 2 illustrates an exemplary encoding and decoding (codec) system 2 for video coding. 1 is a schematic diagram of a codec system 200. Specifically, the codec system 200 supports the implementation of the operating method 100. The codec system 200 provides functionality for both the encoder and decoder. It is generalized to describe the components used in a codec system. The system 200 may be configured to operate in a manner similar to that discussed with respect to steps 101 and 103 in the method of operation 100. The signal is received and segmented, which results in a segmented video signal 201. System 200 then performs the engine operations as discussed with respect to steps 105, 107, and 109 of method 100. When operating as a coder, the segmented video signal 201 is converted into a coded bitstream. When operating as a decoder, the codec system 200 operates in the following manner: the bitstream as discussed with respect to steps 111, 113, 115, and 117 of modulus 100. The codec system 200 generates an output video signal from a general-purpose coder control component. 211, transform scaling and quantization component 213, intra-picture estimation component component 215, intra-picture prediction component 217, motion compensation component 219 , the motion estimation component 221, the scaling and inverse transform component 229, the filter Control analysis component 227, in-loop filter component 225, decoded picture buffer Header formatting and context adaptive bidirectional filtering The CABAC component 231 is shown. In Figure 2, the black lines indicate the movement of data to be coded / decoded. The dashed lines indicate the movement of control data that controls the operation of other components. The components of system 200 may all reside within the encoder. The decoder may include a subset of the components of the decoder system 200. For example, The data includes an intra-picture prediction component 217, a motion compensation component 219, a scaling component 220, and a the decoding and inverse transform component 229, the in-loop filter component 225, and the These components may include a picture buffer component 223. is explained in.

[0070] The segmented video signal 201 is segmented into blocks of pixels by a coding tree. The coding tree is a representation of the various segments of a captured video sequence. Use the divide mode to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. A link may be referred to as a node on the coding tree. A larger parent node may have a smaller parent node. The number of times a node is subdivided depends on the number of nodes in the node / coding tree. In some cases, the divided blocks are called coding units (CUs). For example, a CU may include a luma block, a red differential chroma (Cr) block, and a blue differential Below the CTU, which contains the minute chroma (Cb) blocks along with the corresponding syntax instructions for the CU. The division modes can be two, three, or four, with the shape changing depending on the division mode used. Binary Tree (BT), which is used to partition a node into one or four child nodes, respectively. The segmented video signal 201 may include a ternary tree (TT) and a quad tree (QT). To achieve this, a general purpose coder control component 211, a transform scaling and quantization component component 213, intra-picture estimation component 215, filter control analysis component 22 7, as well as to the motion estimation component 221.

[0071] The generic coder control component 211 processes the images of a video sequence according to the constraints of the application. It is configured to make decisions regarding the coding of the image into a bitstream, e.g. , the generic coder control component 211 controls the bit rate / bit stream size vs. reconstruction Manage quality optimization. Such decisions depend on storage space / bandwidth availability and image quality. This may be done based on resolution requirements. The general coder control component 211 also To mitigate buffer underrun and overrun issues, consider the transmission speed. Manage buffer utilization. To manage these issues, a general purpose coder control component is used. The component 211 manages the segmentation, prediction, and filtering by the other components. For example, the general purpose coder control component 211 dynamically increases the compression complexity and reduces the resolution. Increase the compression complexity to improve bandwidth utilization, or decrease the resolution and bandwidth Therefore, the general coder control component 211 may other components of the video system 200 to control the quality and bitrate of the video signal reconstruction. The general purpose coder control component 211 creates control data This controls the behavior of other components. The control data is stored in the header format. and CABAC component 231 to generate the parameters for decoding at the decoder. It is encoded in the bitstream for signaling the meter.

[0072] The segmented video signal 201 also includes a motion estimation component 221 and a and motion compensation component 219. Each slice may be divided into multiple video blocks. 1 and the motion compensation component 219 calculates one or more blocks in one or more reference frames. performing inter-predictive coding of the received video block relative to The codec system 200 performs multiple coding passes to, for example, For example, an appropriate coding mode may be selected for each block of video data.

[0073] The motion estimation component 221 and the motion compensation component 219 are highly integrated. Although it is possible to perform the motion estimation by the motion estimation component 221, it is shown separately for conceptual purposes. Motion estimation is the process of generating motion vectors that estimate the motion of video blocks. A motion vector is, for example, a vector of a coded object relative to a prediction block. The predicted block may indicate the difference in the pixel difference between the coded block and the image. The predicted block is the block that is found to be a good match for the block that should be referenced. Such pixel differences may be calculated using the sum of absolute differences (SAD), squared It may be determined by sum of difference (SSD), or other difference measure. Several coded objects, including the Coding Tree Block (CTB), and the CU For example, a CTU can be split into CTBs, which are then split into CTs for inclusion in a CU. A CU can be divided into a prediction unit (PU) containing prediction data and / or a CB. can be coded as a transform unit (TU) containing the transformed residual data for the CU. The motion estimation component 221 performs rate-distortion optimization as part of the rate-distortion optimization process. By using motion analysis, motion vectors, PUs, and TUs are generated. The estimation component 221 estimates multiple reference blocks for the current block / frame, The reference block with the best rate-distortion performance may be determined, for example, by the number of motion vectors, The best rate-distortion performance is achieved by selecting the quality of the video reconstruction (e.g. the amount of data lost due to compression) and coding efficiency (e.g., the size of the final encoded Balance the size.

[0074] In some examples, the codec system 200 includes a decoded picture buffer component 2 23, values ​​for sub-integer pixel positions of the reference picture stored in . For example, the video codec system 200 may include a quarter pixel position, an eighth pixel position, and a , or other fractional pixel positions in the reference picture. The estimation component 221 performs motion search for integer and fractional pixel positions. The motion estimation component 22 may perform the motion estimation to output fractional pixel precision motion vectors. 1. By comparing the position of the PU with the position of the predicted block in the reference picture, Calculate motion vectors for PUs of video blocks in a coded slice. The motion estimation component 221 converts the calculated motion vectors into motion data for encoding. and output to the CABAC component 231 as a header formatting and motion compensation. The compensation component 219 outputs the compensation data.

[0075] The motion compensation performed by the motion compensation component 219 is 21 to fetch or generate a prediction block based on the motion vector determined by Again, the motion estimation component 221 and the motion compensation component 219 In some examples, the motion vectors for the PU of the current video block may be functionally integrated. Upon receiving the motion vector, the motion compensation component 219 calculates the predicted motion vector to which the motion vector points. The residual video block may then be located. Subtract the pixel values ​​of the predicted block from the pixel values ​​of the current video block, and Generally, the motion estimation component 221 is The motion compensation component 219 performs motion estimation on the chroma and luma components. For both the predicted block and the non-predicted block, a motion vector calculated based on the luma component is used. The residual block is forwarded to the transform scaling and quantization component 213. do.

[0076] The segmented video signal 201 includes an intra-picture estimation component 215 and an intra-picture estimation component 216. It is also sent to the picture prediction component 217. The intra-picture estimation component 215 and the intra-picture compensation component 219 The picture prediction component 217 may be highly integrated, but for conceptual purposes they are kept separate. The intra-picture estimation component 215 and the intra-picture prediction component The motion estimation component 217 performs the motion estimation between frames as described above. As an alternative to inter prediction performed by the motion compensation component 219, The current block is intra-predicted with respect to blocks in the previous frame. The trap picture estimation component 215 determines the picture to be used to encode the current block. Determines the intra prediction mode. In some examples, the intra picture estimation component The control 215 selects one of the multiple tested intra-prediction modes for encoding the current block. The selected intra prediction mode is then coded as The resulting data is forwarded to the header formatting and CABAC component 231 for encoding.

[0077] For example, the intra-picture estimation component 215 may use various tested intra-picture predictions. Rate-distortion analysis for the measured modes is used to calculate the rate-distortion values ​​for the modes tested. The intra prediction mode with the best rate-distortion performance is selected from the set of intra prediction modes. In general, the encoded block and the original data that was encoded to produce the encoded block are the amount of distortion (or error) between the coded block and the uncoded block, Determine the bit rate (e.g., number of bits) used to produce the resulting block. The intra picture estimation component 215 determines which intra prediction mode is used for the block. The quantization of the various coded blocks is performed to determine which one gives the best rate-distortion value. In addition, the intra-picture estimation component 21 calculates the ratio from the distortion and rate. 5 uses depth modeling mode (DMM) based on rate-distortion optimization (RDO). The method may be configured to code the depth block of the frame.

[0078] The intra-picture prediction component 217, when implemented on an encoder, The selected intra-prediction mode is determined by the picture estimation component 215. Generate a residual block from the predicted block based on , the residual block may be read from the bitstream. The residual block is represented as a matrix The residual block then contains the difference in values ​​between the predicted block and the original block, represented as , and forwarded to the transform scaling and quantization component 213. The constant component 215 and the intra-picture prediction component 217 are used to predict the luma and chrominance components. It may operate on both the .lambda. and .lambda. components.

[0079] The transform scaling and quantization component 213 further compresses the residual block. The transform scaling and quantization component 213 is configured to A transform such as the discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform, is added to the residual block. to produce a video block with residual transform coefficient values. A digital transform, a subband transform, or other types of transforms may also be used. It may be transformed from the pixel value domain to a transform domain such as the frequency domain. The quantization component 213 also quantizes the transformed residual information, e.g., based on frequency. Such scaling is configured to scale different frequency information differently. This involves applying a scale factor to the residual information so that it is quantized to a granularity of Transform scaling and may affect the final visual quality of the reconstructed video. The quantization component 213 also quantizes the transform coefficients to further reduce the bit rate. The quantization process is configured to quantize some or all of the coefficients. The bit depth may be reduced. The degree of quantization can be adjusted by adjusting the quantization parameter. In some examples, the transform scaling and quantization components 213 may then perform a scan of the matrix containing the quantized transform coefficients. The transform coefficients are forwarded to the header formatting and CABAC component 231 for encoding. It is encoded in a bit stream.

[0080] The scaling and inverse transform component 229 performs the transformation to support motion estimation. Apply the inverse operation of the transform scaling and quantization component 213. and inverse transform component 229 applies inverse scaling, transformation, and / or quantization. For example, a reference block may be a prediction block for another current block. The residual block is then reconstructed in the pixel domain for later use. The motion compensation component 221 and / or the motion compensation component 219 may Adding the residual block back to the corresponding predicted block for use in motion estimation The reference block may be calculated by A filter is applied to the reconstructed reference block to reduce artifacts. Such artifacts are caused by the fact that subsequent blocks would otherwise be predicted. can cause inaccurate predictions (and create further artifacts) when .

[0081] The filter control analysis component 227 and the in-loop filter component 225 Apply filters to the residual blocks and / or reconstructed image blocks. For example, The transformed residual block from the scaling and inverse transform component 229 is In order to reconstruct the image blocks, an intra-picture prediction component 217 and / or It may be combined with the corresponding prediction block from the motion compensation component 219. The filter may then be applied to the reconstructed image block. The filter may alternatively be applied to the residual block. Additionally, the filter control analysis component 227 and the in-loop filter component 225 are highly Although they may be integrated and implemented together, they are shown separately for conceptual purposes. The filter applied to the constructed reference block is applied to a specific spatial region, such that Contains several parameters for adjusting how the filter is applied. The control analysis component 227 determines where such filters should be applied. The reconstructed reference block is analyzed to set the corresponding parameters. The data is sent as filter control data for encoding using header formatting and The in-loop filter component 225 forwards the filter to the CABAC component 231. Such a filter is applied based on the data control data. may include a filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such filters can be used in the spatial / pixel domain (e.g., the reconstructed pixel) depending on the example. It may be applied in the frequency domain (on a block of cells) or in the frequency domain.

[0082] When acting as an encoder, the filtered reconstructed image blocks, residual The difference block and / or the predicted block are used later in the motion estimation as discussed above. The decoded picture is stored in the decoded picture buffer component 223 for use by the decoder and When operating as a decoded picture buffer component 223, a portion of the output video signal is The reconstructed and filtered blocks are stored and transferred to the display as The decoded picture buffer component 223 stores the predicted blocks, residual blocks, and / or can be any memory device capable of storing reconstructed image blocks. stomach.

[0083] The header formatting and CABAC component 231 is part of the codec system 200. receives data from various components of the The data is encoded into a coded bitstream. The 231 general control data and filter control data It generates various headers to encode control data such as intra-frame data. It includes prediction and motion data, as well as residual data in the form of quantized transform coefficient data. All prediction data is coded in the bitstream. The system performs all the steps desired by the decoder to reconstruct the original segmented video signal 201. Such information includes the intra-prediction mode index table (codeword (also called mapping table), definition of coding context for various blocks , an indication of the most probable intra prediction mode, an indication of segmentation information, etc. Such data may be encoded using entropy coding. For example, information can be provided by Context-Adaptive Variable Length Coding (CAVLC), CABAC, syntax Space-Context Adaptive Binary Arithmetic Coding (SBAC), Probability Interval Segment Entropy (PIPE) coding, or by using another entropy coding technique. Following entropy coding, the coded bits may be The stream may be sent to another device (e.g., a video decoder) or may be further It may be archived for later transmission or retrieval.

[0084] 3 is a block diagram illustrating an example video encoder 300. to implement the encoding functionality of the codec system 200 and / or the operating method 10. 0 may be used to implement steps 101, 103, 105, 107, and / or 109. The encoder 300 segments the input video signal and generates a video signal substantially identical to the segmented video signal 201. The segmented video signal 301 is then compressed. and encoded into a bitstream by components of the encoder 300.

[0085] Specifically, the partitioned video signal 301 is subjected to intra-picture prediction for intra prediction. The intra-picture prediction component 317 then forwards the intra-picture prediction data to the intra-picture prediction component 317. The sub-picture estimation component 215 and the intra-picture prediction component 217 The segmented video signal 301 may also be stored in a decoded picture buffer. A motion compensation component is used for inter prediction based on a reference block in component 323. The motion compensation component 321 transfers the motion estimation component 221 and It may be substantially similar to the motion compensation component 219. The prediction block and the residual block from the motion compensation component 321 are is transferred to the transform and quantization component 313 for transforming and quantizing the residual block. The transform and quantization component 313 performs transform scaling and quantization. The transformed and quantized residual block and The corresponding predicted blocks (along with associated control data) are coded into the bitstream. The image is forwarded to the entropy coding component 331 for encoding. The coding component 331 is responsible for the header formatting and CABAC components. It may be substantially similar to port 231.

[0086] The transformed and quantized residual block and / or the corresponding prediction block are subjected to motion compensation. Transformation and quantization for reconstruction into a reference block used by component 321 The transform component 313 also forwards the image to the inverse transform and quantization component 329. and the quantization component 329 is substantially the same as the scaling and inverse transform component 229. The in-loop filter in the in-loop filter component 325 may be similar. also applies to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component 325 is connected to the filter control analysis component 227 and the loop The in-loop filter component may be substantially similar to the in-loop filter component 225. The in-loop filter component 325 may be any of a number of filters such as those discussed with respect to the in-loop filter component 225. The filtered block may then be filtered by a motion compensation component. Component 321 of the decoded picture buffer for use as a reference block. The decoded picture buffer component 323 stores the decoded picture It may be substantially similar to the component 223.

[0087] 4 is a block diagram illustrating an example video decoder 400. To implement the decoding functionality of the codec system 200 and / or the steps of the operating method 100 It may be utilized to perform steps 111, 113, 115, and / or 117. The encoder 400 receives the bitstream, for example from the encoder 300, and displays it to the end user. To do this, a reconstructed output video signal is generated based on the bitstream.

[0088] The bitstream is received by the entropy decoding component 433. The Tropey Decoding Component 433 supports CAVLC, CABAC, SBAC, PIPE coding, or other configured to implement an entropy decoding scheme, such as the entropy coding technique of For example, the entropy decoding component 433 may decode the encoded data in the bitstream. To provide a context for interpreting additional data that is coded as words, The decoded information may include general control data, filter control data, and data, segmentation information, motion information, prediction data, and quantized transform coefficients from the residual block Quantized transform coefficients is forwarded to the inverse transform and quantization component 429 for reconstruction into a residual block. The inverse transform and quantization component 429 is the same as the inverse transform and quantization component 329. It may be similar.

[0089] The reconstructed residual block and / or predicted block are based on an intra prediction operation. and forwarded to the intra-picture prediction component 417 for reconstruction into image blocks. The intra-picture prediction component 417 is an intra-picture estimation component. The components may be similar to the components 215 and 217 for intra-picture prediction. The intra-picture prediction component 417 locates reference blocks within a frame. The prediction mode is used to calculate the residual block and the result is the intra predicted image. Reconstruct the block. Reconstructed intra-predicted image block and / or residual The difference block and the corresponding inter-prediction data are fed to an in-loop filter component 42. 5 to the decoded picture buffer component 423, which The buffer component 223 and the in-loop filter component 225 are substantially The in-loop filter component 425 filters the reconstructed image blocks. blocks, residual blocks, and / or predicted blocks, and such information The decoded picture buffer component 423 stores the picture. The reconstructed image block from component 423 is used as a motion compensation component for inter prediction. The motion compensation component 421 forwards the motion estimation component 221 to the motion compensation component 421. and / or may be substantially similar to the motion compensation component 219. Specifically, The motion compensation component 421 uses the motion vector from the reference block to calculate the predicted block. The resulting reconstruction is then applied to the residual block to reconstruct the image block. The decoded blocks are also filtered through the in-loop filter component 425. The decoded picture buffer component 423 may then forward the , and continues to store additional reconstructed image blocks, which are then added to the frame via the partition information. Such frames may also be arranged in sequences. The resulting signal is output to the display as a reconstructed output video signal.

[0090] FIG. 5 illustrates an exemplary bitstream 500 and substreams extracted from the bitstream 500. 5 is a schematic diagram illustrating a bitstream 501. For example, the bitstream 500 may be a codec for decoding by the codec system 200 and / or the decoder 400. and / or may be generated by the encoder 300. As another example, the bitstream 500 is generated in step 109 of method 100 for use by the decoder in step 111. It may be generated by an encoder.

[0091] The bitstream 500 includes a sequence parameter set (SPS) 510, a number of picture parameters, and A meter set (PPS) 512, a plurality of slice headers 514, image data 520, and one or more SEIs The SPS 510 includes a message 515 in the video sequence included in the bitstream 500. Contains sequence data common to all pictures in a picture sequence. This may include size, bit depth, coding tool parameters, bit rate limitations, etc. The PPS 512 contains parameters specific to one or more corresponding pictures. Each picture in the sequence may point to one PPS 512. The PPS 512 points to the corresponding picture. Coding tools, quantization parameters, and offsets available for tiles in a picture-specific coding tool parameters (e.g., filter controls) The slice header 514 may indicate the location of one or more corresponding slices 524 in the picture. Therefore, each slice 524 in the video sequence contains unique parameters. The slice header 514 may refer to the slice header 514. The slice header 514 contains slice type information, picture Order Count (POC), Reference Picture List, Prediction Weight, Tile Entry Point, Devlo In some examples, slices 524 may include tile grouping parameters, etc. In such a case, the slice header 514 may be referred to as a tile group header. The SEI message 515 may also be referred to as a block decryption message. This is an optional message containing the picture output timing, display settings, loss detection, etc. It can be used for related purposes, such as to show fraud, concealment of losses, etc.

[0092] The image data 520 is a video signal that is coded according to inter-prediction and / or intra-prediction. Such image data includes the corresponding transformed and quantized residual data. The data 520 is classified according to the partition used to partition the image before encoding. For example, a video sequence is divided into pictures 521, which are further divided into sub-pictures. A sub-picture 522 may be divided into sub-pictures 522, which are divided into slices 524. The packet 524 may be further divided into tiles and / or CTUs. The coding blocks are divided into coding blocks based on the coding tree. For example, picture 521 may be divided into one or more subpictures. A sub-picture 522 may contain one or more slices 524. Subpicture 521 references PPS 512, and slice 524 references slice header 514. 22 is a stable division across the entire video sequence (also known as a segment) Each slice 524 may contain one or more tiles, so may refer to SPS 510. Each slice 524, and therefore each picture 521 and subpicture 522, may also contain multiple CTUs. It may include.

[0093] Each picture 521 contains visual data associated with the video sequence for the corresponding moment in time. However, in some applications, the image may contain a complete set of pixels. It may be desirable to display only a portion of the image 521. For example, in a virtual reality (VR) system may display a user-selected area of ​​picture 521, which may It creates the feeling of being in the scene depicted in the The region is unknown at the time the bitstream 500 is encoded. Picture 521 is a sub-picture 522, each of which may be viewed by the user. These regions may be separately decoded and displayed based on user input. In other applications, the region of interest may be displayed separately. A television with a picture function can capture a specific area from a video sequence, and therefore A user may wish to display a picture 522 over a picture 521 from an unrelated video sequence. In yet another example, the teleconferencing system may display a general picture of the user currently speaking. A sub-picture 521 of the user not currently speaking and a sub-picture 522 of the user not currently speaking may be displayed. A sub-picture 522 may contain a defined region of picture 521. The sub-picture 522 may be separately decodable from the rest of the picture 521. The temporal motion constrained subpicture does not refer to samples outside the temporal motion constrained subpicture. Since it is coded without reference to the rest of picture 521, there is enough information to perform a complete decoding. Includes information.

[0094] Each slice 524 has a length defined by a CTU in the upper left corner and a CTU in the lower right corner. In some examples, the slices 524 are arranged in a left-to-right and top-to-bottom order. In another example, slice 52 includes a series of tiles and / or CTUs in a raster scan order including: 4 is a rectangular slice. A rectangular slice is a rectangle that is the width of the picture in raster scan order. Instead, rectangular slices are scanned over CTUs and / or tile rows. , and picture 521 and / or The sub-picture 522 may include rectangular and / or square regions. is the smallest unit that can be individually displayed by a decoder. These slices 524 are divided into different sub-pictures to separately depict desired regions of the picture 521. may be assigned to channel 522.

[0095] The decoder may display one or more sub-pictures 523 of a picture 521. The sub-pictures 523 are user-selected subgroups of the sub-pictures 522 or predefined sub-groups. For example, picture 521 is divided into nine sub-pictures 522. However, the decoder may select only a single sub-picture 523 from a group of sub-pictures 522. Subpicture 523 includes slice 525, which is a selection of slice 524. Subpicture 52 is a selected or predefined subgroup. Sub-bitstream 501 is derived from bitstream 500 to allow for three separate displays. The extraction 529 may be performed if the decoder receives only the sub-bitstream 501. In other cases, the entire bitstream 500 may be is sent to the decoder, which extracts the sub-bitstreams 501 for separate decoding. The sub-bitstream 501 is sometimes also referred to generically as the bitstream (529). Note that the sub-bitstream 501 may include SPS 510, PPS 512, and The sub-picture 523 and the slice header 514 and the sub-picture 523 and / or Or includes an SEI message 515 related to the slice 525.

[0096] This disclosure provides a method for selecting and displaying subpicture 523 in a decoder. Signaling various data to support efficient coding of 22. SP S510 determines the subpicture size 531, the subpicture position 532, and the completeness of the subpicture 522. The subpicture size 531 includes the subpicture ID 533 for the corresponding subpicture. Subpicture height in luma samples and subpicture size in luma samples for image 522 The subpicture position 532 is the top left sample of the corresponding subpicture 522. and the top left sample of picture 521. The subpicture size 531 defines the layout of the corresponding subpicture 522. The picture ID 533 contains data that uniquely identifies the corresponding subpicture 522. ID 533 is the raster scan index or other defined value of subpicture 522. Therefore, the decoder reads the SPS 510 and determines the size, position, and Some video coding systems allow the user to determine the location and ID of the subpictures. Data relating to the subpicture 522 may be included in the PPS 512 because the subpicture 522 However, to create subpicture 522, The division used for this is based on the video sequence / segmentation, such as ROI-based applications, VR applications, etc. This may be used by applications that rely on consistent partitioning of sub-pictures 522 across the entire image. Therefore, the partitioning of the sub-pictures 522 generally does not change from picture to picture. Putting the layout information of the subpicture 522 in the PPS 512 (which is the case in signaled redundantly for each picture 521) signaled only once for a sequence / segment, rather than Also, rather than relying on the decoder to derive such information, Signaling picture 522 information reduces the probability of error in the event of packet loss. , supports additional functionality with respect to extracting sub-pictures 523. Thus, SPS Signaling the layout of the sub-pictures 522 at 510 includes the encoder and and / or improve the functionality of the decoder.

[0097] The SPS 510 also generates a motion-constrained subpicture frame associated with the complete set of subpictures 522. The motion constraint subpicture flag 534 indicates whether each subpicture 522 is a temporal motion constraint subpicture. Therefore, the decoder must use the motion constraint subpicture flag to indicate whether the picture is a subpicture. 534 and any of the subpictures 522 can be decoded without decoding the other subpictures 522. This allows you to determine which sub-items can be extracted and displayed separately. While allowing picture 522 to be coded as a temporal motion constrained subpicture, However, other sub-pictures 522 are coded without such constraints for improved coding efficiency. This allows the device to be pinged.

[0098] The subpicture ID 533 is also included in the slice header 514. The slice header 514 contains data associated with the corresponding set of slices 524. , only the subpicture ID 533 corresponding to the slice 524 associated with the slice header 514 is retrieved. Thus, the decoder receives the slice 524 and extracts the subpictures from the slice header 514. 524. The slice ID 533 can be obtained to determine which subpicture 522 contains the slice 524. The decoder also reads the slice headers to correlate them with the associated data in the SPS 510. Therefore, the decoder can use the sub-picture ID 533 from SPS 510. and the subpictures 522 / 523 and 524 by reading the associated slice headers 514. You can decide how to position the slices 524 / 525. Even if the sub-picture 522 is lost in transmission or is intentionally omitted for coding efficiency, Although graphically omitted, it allows sub-picture 523 and slice 525 to be decoded. .

[0099] The SEI message 515 may also include a sub-picture level 535. The rule 535 indicates the hardware resources required to decode the corresponding sub-picture 522. In this way, each sub-picture 522 is coded independently of the other sub-pictures 522. This ensures that each sub-picture 522 uses the correct amount of hardware resources in the decoder. If there is no such subpicture level 535, Each sub-picture 522 is allocated enough resources to decode the most complex sub-picture 522. Therefore, the subpicture level 535 is the hard level at which the subpicture 522 changes. Decoders may over-utilize hardware resources when combined with hardware resource requirements. Prevent allocation.

[0100] FIG. 6 is a schematic diagram illustrating an example picture 600 partitioned into sub-pictures 622. For example, picture 600 may be encoded by, for example, codec system 200, encoder 300, and / or is encoded in a bitstream 500 by the decoder 400, 0. Furthermore, picture 600 supports encoding and decoding according to method 100. The sub-bitstreams 501 may be partitioned and / or included to accommodate the above-mentioned needs.

[0101] Picture 600 may be substantially similar to picture 521. Furthermore, picture 600 may be The image may be partitioned into sub-pictures 622, which are substantially the same as sub-picture 522. Similarly, each subpicture 622 includes a subpicture size 631. 631 may be included in the bitstream 500 as the subpicture size 531. The subpicture size 631 includes a subpicture width 631a and a subpicture height 631b. Width 631a is the width of the corresponding subpicture 622 in units of luma samples. The height 631b is the height of the corresponding subpicture 622 in units of luma samples. 22 each includes a subpicture ID 633, and the subpicture ID 633 is a bit The sub-picture ID 633 may be included in the sub-picture stream 500. In the example shown, subpicture ID 633 is subpicture 6 The subpictures 622 each contain a position 632, which is the index of the subpicture. The position 632 may be included in the bitstream 500 as a subpicture position 532. 622 and the top left sample 642 of picture 600. can be.

[0102] Also shown, some subpictures 622 are temporal motion constrained subpictures 634. In the example shown, sub-picture 5 is The subpicture 622 with picture ID 633 is a temporal motion constrained subpicture 634. , 5, the sub-picture 622 identified as sub-picture 622 is compiled without reference to any other sub-picture 622. The subpictures are loaded and therefore extracted without considering data from other subpictures 622. It shows which sub-pictures 622 are temporal motion constrained sub-pictures. The indication of whether the subpicture is a motion constraint subpicture flag 534 is provided in the bitstream 500. can be signaled in

[0103] As shown, subpicture 622 is arranged to encompass picture 600 without gaps or overlaps. A gap is an area of ​​picture 600 that is not included in any of the sub-pictures 622. An overlap is an area of ​​picture 600 that is included in more than one sub-picture 622. In the example shown in FIG. 6, subpicture 622 is positioned within picture 600 to prevent both gaps and overlaps. The gap leaves samples of picture 600 outside of sub-picture 622. Due to overlap, related slices are included in multiple sub-pictures 622. and overlapping, the samples are coded differently for subpictures 622. This may result in different treatments being applied to the encoder. If permitted, the decoder should be able to decode such codes even when the decoding method is rarely used. Supports the following coding scheme: Allows gaps and overlaps of subpictures 622 By not doing so, the decoder potentially Reduced decoder complexity since it is not necessary to consider significant gaps and overlaps. Furthermore, not allowing gaps and overlaps in the sub-pictures 622 allows the encoder This reduces the complexity of the RDO process in the video sequence. This is because the encoder can omit considering gap and overlap cases when selecting Therefore, avoiding gaps and overlaps is a costly approach to memory resources in the encoder and decoder. This may reduce the use of memory and / or processing resources.

[0104] FIG. 7 illustrates an exemplary mechanism 724 for associating slices 724 with the layout of subpictures 722. 7 is a schematic diagram showing the image 600. For example, the mechanism 700 may be applied to the image 600. Furthermore, The mechanism 700 may be implemented, for example, in the codec system 200, the encoder 300, and / or the decoder 40. 0 may be applied based on the data in the bitstream 500. The 0's may be utilized to support encoding and decoding according to method 100.

[0105] The mechanism 700 divides sub-pictures, such as slices 524 / 525 and sub-pictures 522 / 523, respectively. In the example shown, the subpicture 722 may be applied to a slice 724 in the first subpicture 722. Each of the slices 724 includes a first slice 724a, a second slice 724b, and a third slice 724c. Each slice header contains the subpicture ID 733 of the subpicture 722. Subpicture ID 733 from the device header can be matched with subpicture ID 733 in the SPS. The decoder then finds position 7 of subpicture 722 from the SPS based on the subpicture ID 733. Using the position 732, the subpicture 722 can be positioned relative to the picture. The sample may be positioned relative to the top left corner 742 of the chart. The size may be used to set the height and width of the relative subpicture 722. Therefore, slice 724 may be included in subpicture 722 without any other In the correct subpicture 722 based on subpicture ID 733 without referencing the subpicture This means that other lost sub-pictures can be located in the sub-picture 722. This also makes it possible to extract only the subpicture 722. This helps the example and avoids sending other subpictures. Therefore, subpicture ID 733 is improves performance and / or coding efficiency, which reduces network resources , memory resources, and / or processing resource usage in the encoder and decoder. reduce.

[0106] FIG. 8 is a schematic diagram illustrating another exemplary picture 800 partitioned into sub-pictures 822. Picture 800 may be substantially similar to picture 600. In addition, picture 800 may be , for example, by the codec system 200, the encoder 300, and / or the decoder 400. , can be encoded in and decoded from the bitstream 500. Furthermore, the picture 800 supports encoding and decoding according to the method 100 and / or the mechanism 700. In order to do.

[0107] Picture 800 includes subpicture 822, which is a subset of subpictures 522, 523, and 622. , and / or may be substantially similar to 722. The sub-picture 822 is divided into multiple CTUs 825. CTU825 is a basic codec in a standardized video coding system. The CTU825 is subdivided into coding blocks by the coding tree. The coding blocks are coded according to inter prediction or intra prediction. As shown, some sub-pictures 822a are sub-pictures that are multiples of the size of the CTU 825. It is constrained to include the picture width and the subpicture height. Char 822a has a height of 6 CTUs 825 and a width of 5 CTUs 825. This constraint is Sub-picture 822b located on the right border 801 of the picture and sub-picture 822c located on the bottom border 802 of the picture. In the example shown, subpicture 822b is Subpicture 822b has a width of between 5 and 6 CTUs 825. Therefore, subpicture 822b has 5 complete CTUs. There are 1 complete CTU 825 and 1 incomplete CTU 825. However, as shown in the picture below, Subpictures 822b that are not located on boundary 802 still have subpictures that are multiples of the size of CTU 825. In the example shown, subpicture 822c is constrained to maintain the image height. Subpicture 822c therefore has a height of six complete CTUs. 825 and one incomplete CTU 825. However, the right border of the picture Subpicture 822c that is not located in 801 is still a subpicture that is a multiple of the size of CTU 825. The right border 801 of the picture and the bottom border 802 of the picture are constrained to maintain the pixel width. Note that these may also be referred to as the picture right and picture bottom boundaries, respectively. Also note that the size of the CTU 825 is a user-defined value. The size may be any value between the minimum CTU 825 size and the maximum CTU 825 size. For example, the smallest CTU825 size is 16 luma samples high and 16 luma samples wide. Additionally, the maximum CTU825 size is 128 luma samples high and 128 luma samples wide. It may also be a luma sample.

[0108] As mentioned above, some video systems have heights and dimensions that are multiples of the size of the CTU825. The subpicture 822 may be limited to include the size and width of the subpicture. Multiple picture layouts, e.g. CTU825, with an overall width or height that is not a multiple of the size This may prevent the picture 800 from working properly. and right subpicture 822b each have a height and width that are not a multiple of the size of CTU 825. By allowing this, the subpicture 822 can be decoded arbitrarily without causing a decoding error. This may be used with any picture 800. This improves the functionality of the encoder and decoder. Furthermore, the improved functionality allows the encoder to code pictures more efficiently. This allows for network resources, memory resources, and and / or reduce the use of processing resources in the encoder and decoder.

[0109] As described herein, this disclosure provides a method for implementing subpicture coding in video coding. This paper describes the design of a picture partitioning scheme based on sub-pictures. A rectangular area within a picture that can be independently decoded using a decoding process similar to that used for This disclosure relates to coded video sequences and / or bitstreams. SUBPICTURE SIGNALING IN A SYSTEM AND A PROCESS FOR SUBPICTURE EXTRACTION - Patent application The description of the technique is based on VVC by JVET of ITU-T and ISO / IEC. The techniques also apply to other video codec specifications. Such embodiments may be applied individually or in combination. It is possible.

[0110] Information about sub-pictures that may be present in a coded video sequence (CVS) is stored in S It may be signaled in a sequence level parameter set such as PS. Such signaling may include the following information: The number of sub-pictures present in each picture of the CVS The number of chats may be signaled in the SPS. In the context of an SPS or CVS, all Sub-pictures at the same position relative to an access unit (AU) are collectively called sub-pictures. It may also be called a sequence. A loop for this purpose may also be included in the SPS. This information includes subpicture identification information, subpicture the position of the subpicture luma sample (for example, the top left corner luma sample of the subpicture and the top left corner luma sample of the picture) The SPS may also include the offset distance between the image and the subpicture, and the size of the subpicture. signals whether each subpicture is a motion constrained subpicture (including the MCTS feature). The profile, layer, and level information for each subpicture may also be stored in Such information may be signaled or derivable at the decoder. , a bitstream created by extracting subpictures from the original bitstream. It may be used to determine profile, tier, and level information for the The profile and layer of each subpicture is the same as the profile of the entire bitstream. and hierarchy. Each subpicture level is explicitly Such signaling may be present in a loop contained in the SPS. Sequence-level hypothetical reference decoder (HRD) parameters are or equivalently, each sub-picture sequence) It may be signaled in the section.

[0111] When a picture is not partitioned into two or more sub-pictures, the characteristics of the sub-pictures (for example, position, size, etc.) are not present / in the bitstream, except for the subpicture ID. The subpictures of a picture in a CVS may not be signaled in the stream. When the images are extracted, each access unit in the new bitstream is a subpicture. In this case, the pictures in each AU in the new bitstream Therefore, the position and size of the sub-pictures in the SPS are not specified. There is no need to signal any subpicture characteristics, because such information is stored in the picture. However, the subpicture identification information is still not available in the system. The ID may be specified in the VCL NAL unit data included in the extracted subpicture. This allows the subpicture I to be referenced by the unit / tile group. D may be allowed to remain the same when extracting sub-pictures.

[0112] The position (x and y offset) of the subpicture within the picture is expressed in luma samples. The position can be signaled in units of the top left corner luma sample of the subpicture and the picture. Alternatively, the position of the subpicture within the picture The position may be signaled in units of the minimum coding luma block size (MinCbSizeY). Alternatively, the units of the subpicture position offset may be syntaxed in the parameter set. The units may be explicitly indicated by the CtbSizeY, MinCbSizeY, and luma sample It may be a sine wave, or some other value.

[0113] The subpicture size (subpicture width and subpicture height) is the unit of luma samples. Alternatively, the size of the subpicture can be signaled in units of the minimum coding luma. It can be signaled in units of block size (MinCbSizeY). Alternatively, it can be signaled in units of subpicture subpicture size (MinCbSizeY). The units of the size value may be explicitly indicated by syntax elements in the parameter set. The units may be CtbSizeY, MinCbSizeY, luma samples, or other values. When the right border of the picture does not coincide with the right border of the picture, the width of the subpicture is It may be required to be an integer multiple of the CTU size (CtbSizeY). When the bottom border of the subpicture does not coincide with the bottom border of the picture, the subpicture height is The width of the subpicture may be required to be an integer multiple of the luma CTU If the size is not an integral multiple, the subpicture is placed at the rightmost position within the picture. Similarly, if the height of the subpicture is not an integer multiple of the luma CTU size, If not, the subpicture is required to be located at the bottommost position in the picture. In some cases, the subpicture width is signaled in units of luma CTU size. However, the width of the subpicture is not an integer multiple of the luma CTU size. The actual width in pixels can be derived based on the offset position of the subpicture. The width of the picture can be derived based on the luma CTU size, and the height of the picture can be derived based on the luma sample size. Similarly, the subpicture height can be derived based on the luma CTU size. The subpicture height may be signaled in increments of 1, but the subpicture height may not be an integer multiple of the luma CTU size. In such cases, the actual height in luma samples is determined by the offset position of the subpicture. Deriving the sub-picture height based on the luma CTU size and the picture height can be derived based on the luma samples.

[0114] For any given subpicture, the subpicture ID is different from the subpicture index. The subpicture index may be used as a signature in a subpicture loop in an SPS. The subpicture ID may be the index of the subpicture to be nulled. The index of the subpicture in the picture's subpicture raster scan order. The value of the subpicture ID of each subpicture is the same as the subpicture index. When a subpicture is generated, the subpicture ID may be signaled or derived. When the picture ID differs from the subpicture index, the subpicture ID must be explicitly signaled. The number of bits for signaling the subpicture ID depends on the subpicture feature may be signaled in the same parameter set (e.g., in the SPS) that contains Some values ​​for the subpicture ID may be reserved for certain purposes. For example, a tile group header specifies which subpicture contains the tile group. Prevents the accidental inclusion of anti-emulation code when including subpicture IDs for To ensure that the first few bits of the tile group header are not all zeros, Additionally, the value 0 is reserved for subpictures and may not be used. In the optional case where the image does not encompass the entire area of ​​the picture without gaps and overlaps, A value (e.g., value 1) is reserved for tile groups that are not part of any subpicture. Alternatively, the sub-picture IDs of the remaining areas may be explicitly signaled. The number of bits for signaling the subpicture ID may be constrained as follows: The range of values ​​is for all subpictures in the picture, including the reserved value for the subpicture ID. It must be sufficient to uniquely identify the picture. For example, a picture ID for a subpicture ID. The minimum number of bits is Ceil(Log2(number of subpictures in a picture + spare subpicture IDs) The value can be any number of

[0115] The union of subpictures must encompass the entire picture without gaps or overlaps. When this constraint is applied, for each subpicture, There may also be a flag to specify whether the picture is a motion constrained subpicture. , which indicates that sub-pictures can be extracted. Alternatively, the sub-picture union can be It may not encompass the entire picture, but overlaps may not be allowed.

[0116] Subpicture extraction without requiring the extractor to parse the rest of the NAL unit bits To aid the extraction process, a subpicture ID may be present immediately after the NAL unit header. For VCL NAL units, the subpicture ID is the first bit of the tile group header. For non-VCL NAL units, the following may apply: SPS In PPS, the subpicture ID does not need to be located immediately after the NAL unit header. If all tile groups of the same picture are constrained to refer to the same PPS, The picture ID does not need to be located immediately after the NAL unit header. If a picture group is allowed to refer to different PPSs, the subpicture ID must be the first It may be present in the first bit (e.g., immediately after the NAL unit header). Any group of tiles in a block may be allowed to share the same PPS. Additionally, tile groups of the same picture are allowed to refer to different PPSs, and the same picture Subpicture IDs are also allowed when different tile groups in a pixel share the same PPS. may not be present in the PPS syntax. Alternatively, tiling of the same picture It is allowed for loops to reference different PPSs, and even different tile groups of the same picture. When sharing the same PPS is allowed, the list of subpicture IDs is included in the PPS syntax. This list indicates the subpictures to which the PPS applies. For NAL units, non-VCL units (e.g., access unit delimiter, end of sequ If a criterion (such as ence, end of bitstream) applies at picture level or above, then In this case, the subpicture ID does not have to be located immediately after the NAL unit header. In this case, the subpicture ID may be located immediately after the NAL unit header.

[0117] Using the SPS signaling described above, tile partitioning within each sub-picture is performed in PPS. Tile groups within the same picture may refer to different PPSs. In this case, tile grouping is allowed only within each subpicture. The concept of tile grouping is the division of a subpicture into tiles.

[0118] Alternatively, a set of parameters for describing tile partitioning within each sub-picture may be Such a parameter set is called a Sub-Picture Parameter Set (SPPS). The SPPS references the SPS. The syntax element that references the SPS ID must exist in the SPPS. The SPPS may contain a subpicture ID. For the purpose of subpicture extraction, the subpicture ID is used. The syntax element that references the ID is the first syntax element in the SPPS. The tiling structure (e.g., number of columns, number of rows, uniform tile spacing, etc.) A flag to indicate whether the subpicture filter is valid across the associated subpicture boundary. Alternatively, the subpicture properties for each subpicture may be The tile partitioning within each sub-picture may be signaled in the SPPS instead of the sub-picture itself. Tile groups within the same picture may also be signaled in the PPS. It is allowed to refer to different PPSs. When SPPS is enabled, the SPPSs are consecutively arranged in decoding order. However, SPPS may not be used in AUs that are not the first in the CVS. Activated / may be activated. With multiple sub-pictures in any AU At any moment during the decoding process of a single-layer bitstream, multiple SPPSs may be active. SPPS may be shared by different sub-pictures of an AU. Alternatively, the SPPS and PPS may be combined into one parameter set. It is not required that all tile groups in the same picture refer to the same PPS. All tile groups in the same subpicture are merged between SPPS and PPS. Constraints may be applied such that all the parameters may refer to the same set of parameters.

[0119] The number of bits used to signal the subpicture ID is determined by the NAL unit header. When present in the NAL unit header, Such information is stored at the beginning of the NAL unit payload (e.g., immediately after the NAL unit header). The subpicture extraction process is performed by parsing the subpicture ID value (the first few bits). For such signaling, a spare bit in the NAL unit header is used. Some of the bits (e.g., 7 spare bits) are used to avoid increasing the length of the NAL unit header. The number of bits of such signaling is determined by the sub-picture-ID-b It may contain a value for it-len, e.g., the 7 spare bits of the VVC NAL unit header. Four of the bits may be used for this purpose.

[0120] When decoding a sub-picture, the position of each coding tree block (e.g., xCtb and yCtb) are the actual luma sample locations in the picture, not the luma sample locations in the subpicture. The coding tree blocks may be adjusted to the sub-sample position. Since the image is decoded by referring to the picture rather than the original, the image is decoded at the same position from each reference picture. This avoids the extraction of subpictures that are not needed. To do this, the variables SubpictureXOffset and SubpictureYOffset are used to determine the position of the subpicture (s The value of the variable can be derived based on the subpic_x_offset and subpic_y_offset. The values ​​of the x and y coordinates of the luma sample position of each coding tree block in the These may be added together.

[0121] The subpicture extraction process can be defined as follows: The input to the process is the extracted This can be in the form of a subpicture ID or a subpicture position. When the input is a sub-picture position, the associated sub-picture ID can be This can be resolved by analyzing the subpicture information in the non-VCL NAL units. The following applies: The syntax in the SPS for picture size and level The box element may be updated with the size and level information of the subpicture. Non-VCL NAL units, i.e., PPS, Access Unit Delimiter (AUD), End of Sequence ( End of Bitstream (EOB), and any applicable at picture level or above Other non-VCL NAL units in the same subpicture remain unchanged. Any remaining non-VCL NAL units with new subpicture IDs may be removed. VCL NAL units with subpicture IDs that are not equal to the picture ID may also be removed.

[0122] Sequence level subpicture nesting SEI messages are used to organize a set of subpictures. For nesting of AU level SEI messages or subpicture level SEI messages This may be used for buffering periods, picture timing, and non-HRD The syntax of this subpicture nesting the SEI message may contain The formats and semantics can be as follows: Omni-Directional Media Format (OMA) F) In system operation in an environment, etc., a subpicture sequence that includes a viewport may be requested and decoded by an OMAF player. The level SEI message contains sub-picture sequences that collectively encompass a rectangular picture area. This information is used by the system to carry information about a set of This information can be used to determine the required decoding capabilities as well as the set of sub-picture sequences. This information is only available for the set of subpicture sequences. This information also indicates the level of the bitstream containing the subpicture sequence. Indicates the bit rate of the bitstream containing only the sub-bitstream. The frame extraction process may be specified for a set of sub-picture sequences. The advantage of doing this is that bitstreams containing only a set of subpicture sequences can also be adapted. The drawback is that it can be difficult to consider different viewport sizes. When considering the individual sub-picture sequences, which may already be numerous, there are many The implication is that many such sets may exist.

[0123] In an exemplary embodiment, one or more of the disclosed examples may be implemented as follows. A subpicture is defined as a rectangular region of one or more tile groups within a picture. An admissible bisection process may be defined as follows: The inputs are the bisection mode btSplit, the coding block width cbWidth, and the coding block height cbHeight, the coding block being considered relative to the top-left luma sample of the picture. The position (x0,y0) of the top left luma sample of the block, the multi-type tree depth mttDepth, and the offset. The maximum multitype tree depth maxMttDepth, the maximum binary tree size maxBtSize, and the partition index The output of this process is the variable allowBtSplit. [Table 1]

[0124] The variables parallelTtSplit and cbSize are derived as specified above. Spit is derived as follows: If the following conditions are met, namely, cbSize is less than or equal to MinBtSizeY, Width is greater than maxBtSize, cbHeight is greater than maxBtSize, and mttDepth is greater than maxMt If one or more of the conditions of being greater than or equal to tDepth is true, allowBtSplit is equal to FALSE. Instead, the following condition is met: btSplit equals SPLIT_BT_VER , and y0+cbHeight is greater than SupPicBottomBorderInPic allowBtSplit is set equal to FALSE. Otherwise, the following condition is met: , btSplit is equal to SPLIT_BT_HOR, and x0+cbWidth is greater than SupPicRightBorderInPic and y0+cbHeight is less than or equal to SubPicBottomBorderInPic are all true If so, then allowBtSplit is set equal to FALSE. Otherwise, the following condition is met: That is, mttDepth is greater than 0, partIdx is equal to 1, and MttSplitMode[x0][y0][mttD If all of the conditions are true, then allowBtSplit is Otherwise, the following condition is met: btSplit is set equal to SPLIT_BT_VER , cbWidth is less than or equal to MaxTbSizeY, and cbHeight is greater than MaxTbSizeY. If all of the conditions above are true, then allowBtSplit is set equal to FALSE. and the following conditions are met: btSplit equals SPLIT_BT_HOR, cbWidth is greater than MaxTbSizeY If all of the conditions are true: greater than, and cbHeight is less than or equal to MaxTbSizeY, then lowBtSplit is set equal to FALSE, otherwise allowBtSplit is set equal to TRUE It is set.

[0125] An acceptable trichotomy process may be defined as follows: The inputs to this process are , 3-way split mode ttSplit, coding block width cbWidth, coding block height cb Height, the considered coding block relative to the top-left luma sample of the picture The position (x0,y0) of the top-left luma sample of the multi-type tree, mttDepth, and the maximum value with the offset. The maximum multitype tree depth is maxMttDepth, and the maximum binary tree size is maxTtSize. The output of the process is the variable allowTtSplit. [Table 2]

[0126] The variable cbSize is derived as specified above. The variable allowTtSplit is derived as follows: The following conditions are met: cbSize is less than or equal to 2*MinTtSizeY, cbWidth is less than or equal to Min(MaxTbSize Y,maxTtSize), cbHeight is greater than Min(MaxTbSizeY,maxTtSize), mttDepth is greater than or equal to maxMttDepth, x0+cbWidth is greater than SupPicRightBoderInPic, and y0+c If one or more of the following conditions is true: bHeight is greater than SubPicBottomBorderInPic If so, allowTtSplit is set equal to FALSE; otherwise, allowTtSplit is set equal to TRUE. are set equal.

[0127] The syntax and semantics of the sequence parameter set RBSP are as follows: do. [Table 3]

[0128] pic_width_in_luma_samples is the width of each decoded picture in units of luma samples pic_width_in_luma_samples must not be equal to 0 and must be an integer value greater than or equal to MinCbSizeY. pic_height_in_luma_samples is the height of each decoded pixel in units of luma samples. Specifies the height of the displayed picture. pic_height_in_luma_samples must not be equal to 0. and is an integer multiple of MinCbSizeY. num_subpicture_minus1 plus 1 is The number of sub-pictures to be partitioned in the coded picture is specified by the subpic_id_len_minus1 plus 1 belongs to the video sequence The syntax element subpic_id[i] in the PS, spps_subpic_id in the SPPS that references the SPS, and used to represent the tile_group_subpic_id in the tile group header that references the SPS. Specifies the number of bits used. The value of subpic_id_len_minus1 must be Ceil(Log2( num_subpic_minus1+2) to 8. subpic_id[i] is the pic that refers to the SPS. Specifies the subpicture ID of the ith subpicture in the picture. The length of subpic_id[i] is pic_id_len_minus1+1 bits. The value of subpic_id[i] must be greater than 0. _level_idc[i] is the CVS-specified resource requirement resulting from the extraction of the ith subpicture. The bitstream may not use subpic_level_idc[i] other than the one specified. Other values ​​of subpic_level_idc[i] are reserved. In this case, the value of subpic_level_idc[i] is inferred to be equal to the value of general_level_idc.

[0129] subpic_x_offset[i] is the offset of the ith subpicture relative to the top-left corner of the picture. Specifies the horizontal offset of the top left corner. If not present, the value of subpic_x_offset[i] is set to 0. The offset value of subpicture x is assumed to be equal to SubpictureXOffset[i]=subpic _x_offset[i] is derived as subpic_y_offset[i] relative to the top left corner of the picture. Specifies the relative vertical offset of the top-left corner of the ith subpicture. If not present , the value of subpic_y_offset[i] is inferred to be equal to 0. The value of the subpicture y offset is derived as follows: SubpictureYOffset[i] = subpic_y_offset[i]. uma_samples[i] is the width of the ith decoded subpicture for which this SPS is the active SPS The sum of SubpictureXOffset[i] and subpic_width_in_luma_samples[i] is pic_wid When it is less than th_in_luma_samples, the value of subpic_width_in_luma_samples[i] is CtbSizeY If not present, the value of subpic_width_in_luma_samples[i] shall be an integer multiple of , which is assumed to be equal to the value of pic_width_in_luma_samples. es[i] specifies the height of the ith decoded subpicture for which this SPS is the active SPS. The sum of SubpictureYOffset[i] and subpic_height_in_luma_samples[i] is When it is less than luma_samples, the value of subpic_height_in_luma_samples[i] is an integer in CtbSizeY. When not present, the value of subpic_height_in_luma_samples[i] shall be the same as that of pic_ It is assumed to be equal to the value of height_in_luma_samples.

[0130] The union of sub-pictures should encompass the entire area of ​​the picture without overlaps or gaps This is a bitstream conformance requirement. subpic_motion_constraint equal to 1 ed_flag[i] specifies that the ith subpicture is a temporal motion constrained subpicture subpic_motion_constrained_flag[i] equal to 0 means that the ith subpicture is temporal motion constrained. Specifies whether the image is a subpicture or not. If not present, subpic_motion The value of n_constrained_flag is inferred to be equal to 0.

[0131] Variables SubpicWidthInCtbsY, SubpicHeightInCtbsY, SubpicSizeInCtbsY, SubpicWidthInM inCbsY, SubpicHeightInMinCbsY, SubpicSizeInMinCbsY, SubpicSizeInSamplesY, Subpic WidthInSamplesC and SubpicHeightInSamplesC are derived as follows: SubpicWidthInLumaSamples[i]=subpic_width_in_luma_samples[i] SubpicHeightInLumaSamples[i]=subpic_height_in_luma_samples[i] SubPicRightBorderInPic[i]=SubpictureXOffset[i]+PicWidthInLumaSamples[i] SubPicBottomBorderInPic[i]=SubpictureYOffset[i]+PicHeightInLumaSamples[i] SubpicWidthInCtbsY[i]=Ceil(SubpicWidthInLumaSamples[i]÷CtbSizeY) SubpicHeightInCtbsY[i]=Ceil(SubpicHeightInLumaSamples[i]÷CtbSizeY) SubpicSizeInCtbsY[i]=SubpicWidthInCtbsY[i]*SubpicHeightInCtbsY[i] SubpicWidthInMinCbsY[i]=SubpicWidthInLumaSamples[i] / MinCbSizeY SubpicHeightInMinCbsY[i]=SubpicHeightInLumaSamples[i] / MinCbSizeY SubpicSizeInMinCbsY[i]=SubpicWidthInMinCbsY[i]*SubpicHeightInMinCbsY[i] SubpicSizeInSamplesY[i]=SubpicWidthInLumaSamples[i]*SubpicHeightInLumaSamples[i] SubpicWidthInSamplesC[i]=SubpicWidthInLumaSamples[i] / SubWidthC SubpicHeightInSamplesC[i]=SubpicHeightInLumaSamples[i] / SubHeightC

[0132] The syntax and semantics of the subpicture parameter set RBSP are as follows: be. [Table 4]

[0133] spps_subpic_id identifies the subpicture to which the SPPS belongs. The length of spps_subpic_id is subpic_id_len_minus1+1 bits. spps_subpic_parameter_set_id is the same as other syntax Identifies the SPPS for reference by subpic elements. The value of spps_subpic_parameter_set_id is The range of the spps_seq_parameter_set_id is 0 to 63, inclusive. Specifies the value of sps_seq_parameter_set_id for the SPS. spps_seq_parameter_set_id value shall be in the range 0 to 15 inclusive. single_tile_in_subpic_fl equal to 1 ag specifies that there is only one tile in each subpicture that references the SPPS. Equal to 0 The new single_tile_in_subpic_flag specifies that there is more than one tile in each subpicture that references the SPPS. num_tile_columns_minus1 plus 1 is the number of sub-pictures Specifies the number of tile columns that separate the tiles. num_tile_columns_minus1 can range from 0 to 1, including both ends. If it does not exist, num_t The value of tile_columns_minus1 is inferred to be equal to 0. The value of num_tile_rows_minus1 plus 1 is specifies the number of tile rows that divide the subpicture. The range shall be from 0 to PicHeightInCtbsY[spps_subpic_id]-1, inclusive. When the value of num_tile_rows_minus1 is not specified, the value of num_tile_rows_minus1 is inferred to be equal to 0. The variable NumTilesInPic is num_tile_columns_minus1+1)*(num_tile_rows_minus1+1).

[0134] When single_tile_in_subpic_flag is equal to 0, NumTilesInPic shall be greater than 0. uniform_tile_spacing_flag equal to 1 specifies that tile column boundaries are equal to tile row boundaries as well. uniform_tile_spac equal to 0 specifies that the tile spacing is uniformly distributed across the subpicture. ing_flag specifies whether tile column boundaries, and similarly tile row boundaries, are uniform across the entire subpicture. No distribution, syntax elements tile_column_width_minus1[i] and tile_row_height_minus 1[i] to specify that it is explicitly signaled. When not present, uniform_ The value of tile_spacing_flag is inferred to be equal to 1. Add 1 to tile_column_width_minus1[i]. tile_row_height_minus1[i] specifies the width of the i-th tile row in CTB units. plus 1 specifies the height of the ith tile row in CTB units.

[0135] The following variable specifies the width of the ith tile column in CTB units, inclusive: list ColWidth[i] for i ranging from to num_tile_columns_minus1, in units of CTB Specifies the height of the jth tile row, in the range 0 to num_tile_rows_minus1, inclusive. A list RowHeight[j] for barrel j, specifying the position of the ith tile column boundary in CTB units , the list ColBd[i ], specifies the position of the jth tile row boundary in CTB units, from 0 to num_tile_row inclusive List RowBd[j] for j ranging from s_minus1+1, in the CTB raster scan of the picture Specifies the conversion from the CTB address in the tile scan to the CTB address in the tile scan. List CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs ranging from PicSizeInCtbsY to PicSizeInCtbsY-1, From the CTB address in the file scan to the CTB address in the CTB raster scan of the picture A list for ctbAddrTs ranging from 0 to PicSizeInCtbsY-1, inclusive, specifying the conversion CtbAddrTsToRs[ctbAddrTs], specifies the conversion from CTB address to tile ID in tile traversal. A list of TileId[c tbAddrTs], which specifies the conversion from tile index to the number of CTUs in the tile, inclusive First, a list NumCtusInTile[tileIdx] for tileIdx ranging from 0 to PicSizeInCtbsY-1, Specifies the conversion from a tile ID to a CTB address in tile scanning of the first CTB in a tile. , a list FirstCtbAddrTs[ti leIdx], specifies the width of the ith tile column in units of luma samples, from 0 to num_ The list ColumnWidthInLumaSamples[i] for i ranging from tile_columns_minus1, and and a range from 0 to num_j, inclusive, that specifies the height of the jth tile row in units of luma samples. The list RowHeightInLumaSamples[j] for j over the range tile_rows_minus1 is It is derived by invoking the star and tile scan conversion process. ColumnWidthInLumaSamples[i], for i ranging from 0 to num_tile_columns_minus1, and RowHeightInLumaSamples[j] for j ranging from 0 to num_tile_rows_minus1, inclusive All values ​​of must be greater than 0.

[0136] loop_filter_across_tiles_enabled_flag equal to 1 indicates that in-loop filtering behavior is enabled. Specifies that the SPPS reference may span tile boundaries within the subpicture loop_filter_across_tiles_enabled_flag equal to 0 disables in-loop filtering operation. Specifies that the operation will not span tile boundaries within subpictures that reference SPPS. The in-loop filtering operations include a deblocking filter, sample adaptive offset loop_filter_acro The value of ss_tiles_enabled_flag is inferred to be equal to 1. bpic_enabled_flag specifies whether in-loop filtering operations are performed in subpictures that refer to SPPS. loop_filt equal to 0 specifies that the loop may run across subpicture boundaries. er_across_subpic_enabled_flag specifies whether in-loop filtering operations refer to the subpic Specifies that the loop will not run across subpicture boundaries within a picture. The filtering operations include a deblocking filter, a sample adaptive offset filter, and and adaptive loop filter operation. When not present, loop_filter_across_subpic_enable The value of d_flag is inferred to be equal to the value of loop_filter_across_tiles_enabed_flag.

[0137] The syntax and semantics of a general tile group header are as follows: . [Table 5]

[0138] Tile group header syntax elements tile_group_pic_parameter_set_id and tile_ The value of group_pic_order_cnt_lsb is the sum of all tile groups in a coded picture. The tile group header syntax element tile_g The group_subpic_id value is the number of tile group headers in all coded subpictures. The tile_group_subpic_id is the subpic to which the tile group belongs. The length of tile_group_subpic_id is subpic_id_len_minus1+1 bits. The tile_group_subpic_parameter_set_id is the spps_subpic_parameter for the SPPS in use. The value of tile_group_spps_parameter_set_id must be inclusive. The range is from 0 to 63.

[0139] The following variables are derived and override the respective variables derived from the active SPS: PicWidthInLumaSamples=SubpicWidthInLumaSamples[tile_group_subpic_id] PicHeightInLumaSamples=PicHeightInLumaSamples[tile_group_subpic_id] SubPicRightBorderInPic=SubPicRightBorderInPic[tile_group_subpic_id] SubPicBottomBorderInPic=SubPicBottomBorderInPic[tile_group_subpic_id] PicWidthInCtbsY=SubPicWidthInCtbsY[tile_group_subpic_id] PicHeightInCtbsY=SubPicHeightInCtbsY[tile_group_subpic_id] PicSizeInCtbsY=SubPicSizeInCtbsY[tile_group_subpic_id] PicWidthInMinCbsY=SubPicWidthInMinCbsY[tile_group_subpic_id] PicHeightInMinCbsY=SubPicHeightInMinCbsY[tile_group_subpic_id] PicSizeInMinCbsY=SubPicSizeInMinCbsY[tile_group_subpic_id] PicSizeInSamplesY=SubPicSizeInSamplesY[tile_group_subpic_id] PicWidthInSamplesC=SubPicWidthInSamplesC[tile_group_subpic_id] PicHeightInSamplesC=SubPicHeightInSamplesC[tile_group_subpic_id]

[0140] The coding tree unit syntax is as follows.

Table 6

Table 7

[0141] The syntax and semantics of the coding quad-tree are as follows.

Table 8A

Table 8B

[0142] qt_split_cu_flag[x0][y0] specifies whether the coding unit is split into coding units with half the horizontal size and vertical size. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block considered relative to the top-left luma sample of the picture. When qt_split_cu_flag[x0][y0] does not exist, the following applies. If one or more of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. If treeType is equal to DUAL_TREE_CHROMA or, otherwise, greater than MaxBtSizeY, x0+(1<<log2CbSize) is greater than SubPicRightBorderInPic and (1<<log2CbSize) is greater than MaxBtSizeC. If treeType is equal to DUAL_TREE_CHROMA or, otherwise, greater than MaxBtSizeY, y0+(1<<log2 ​​​​​​​​CbSize) is larger than SubPicBottomBorderInPic, and (1 << log2CbSize) is larger than MaxBtSizeC Yes.

[0143] In other cases, if all of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. If treeType is equal to DUAL_TREE_CHROMA, or otherwise if it is larger than MinQtSizeY, x0 + (1 << log2CbSize) is larger than SubPicRightBorderInPic , y0 + (1 << log2CbSize) is larger than SubPicBottomBorderInPic, and (1 << log2CbSize) is larger than MinQtS izeC. Otherwise, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 0 .

[0144] The syntax and semantics of the multi-type tree are as follows.

Table 9A

Table 9B

Table 9C

[0145] An mtt_split_cu_flag equal to 0 specifies that the coding unit is not split. An mtt_split_cu_flag equal to 1 means that, as indicated by the syntax element mtt_split_cu_binary_flag, the coding unit is split into two coding units using a binary split and two coding units to specify that the video is split into three coding units, or into three coding units using trisection. The division into two or three parts is indicated by the syntax element mtt_split_cu_vertical_flag. It can be either vertical or horizontal, as desired. When not specified, the value of mtt_split_cu_flag is inferred as follows: x0+cbWidth is SubPicRight BorderInPic and y0+cbHeight is greater than SubPicBottomBorderInPic If one or more of the following conditions are true, the value of mtt_split_cu_flag is inferred to be equal to 1: Otherwise, the value of mtt_split_cu_flag is inferred to be equal to 0.

[0146] The derivation process for temporal luma motion vector prediction is as follows: The output of the 1 / 16 fractional sample precision motion vector prediction mvLXCol and availability flag a The variable currCb is the current luma code at luma position (xCb, yCb). The variables mvLXCol and availableFlagLXCol are derived as follows: If tile_group_temporal_mvp_enabled_flag is equal to 0 or the reference picture is If it is the current picture, both components of mvLXCol are set equal to 0 and availableFlagL XCol is set equal to 0. Otherwise (tile_group_temporal_mvp_enabled_flag is 1 and the reference picture is not the current picture), the following steps are applied: The motion vector at the same position in the bottom right is derived as follows: xColBr=xCb+cbWidth (8-355) yColBr=yCb+cbHeight (8-356)

[0147] yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY and yColBr is equal to SubPicBottomBorderInP If xColBr is less than SubPicRightBorderInPic, and xColBr is less than SubPicRightBorderInPic, then the following is true: The variable colCb is the internal (xColBr>>3) of the picture at the same position specified by ColPic. )<<3,(yColBr>>3)<<3) Specifies that the luma position (xColCb, yColCb) is locked to the same position specified by ColPic. The same position specified by colCb relative to the top-left luma sample of a picture The upper left sample of the luma coding block at the same position is set equal to the upper left sample of the luma coding block at the same position. The derivation process of the vector begins with currCb, colCb, (xColCb, yColCb), r It is called with efIdxLX and sbFlag as inputs and outputs mvLXCol and availableF Otherwise, both components of mvLXCol are set equal to 0. , availableFlagLXCol is set equal to 0.

[0148] The derivation process for the temporal triangle integration candidates is as follows: olC1, availableFlagLXColC0 and availableFlagLXColC1 are derived as follows: If le_group_temporal_mvp_enabled_flag is equal to 0, both mvLXColC0 and mvLXColC1 The component is set equal to 0, and availableFlagLXColC0 and availableFlagLXColC1 are equal to 0. Otherwise (tile_group_temporal_mvp_enabled_flag is equal to 1), The following steps are applied: The motion vector mvLXColC0 at the same position in the bottom right is It is derived as follows. xColBr=xCb+cbWidth (8-392) yColBr=yCb+cbHeight (8-393)

[0149] yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY and yColBr is equal to SubPicBottomBorderInP If xColBr is less than SubPicRightBorderInPic, and xColBr is less than SubPicRightBorderInPic, then the following is true: The variable colCb is the internal (xColBr>>3) of the picture at the same position specified by ColPic. )<<3,(yColBr>>3)<<3) Specifies that the luma position (xColCb, yColCb) is locked to the same position specified by ColPic. The top-left luma sample of a picture is compared to the co-located luma sample specified by colCb. The upper left sample of the coding block is set equal to the The derivation process of the rule begins with currCb, colCb, (xColCb, yColCb), refIdxLXC set equal to 0. 0, and sbFlag as inputs, and the output is mvLXColC0 and availableFlagL XColC0. Otherwise, both components of mvLXColC0 are set equal to 0. , availableFlagLXColC0 is set equal to 0.

[0150] The process of deriving the constructed affine control point motion vector integration candidates is as follows. The fourth (co-located bottom-right) control point motion vector cpMvLXCorner[3] with X at 0 and 1 , reference index refIdxLXCorner[3], prediction list usage flag predFlagLXCorner[3], and and the availability flag availableFlagCorner[3] are derived as follows: X is 0 or 1. Then, the reference index for the temporal integration candidate, refIdxLXCorner[3], is set equal to 0. . With X being either 0 or 1, the variables mvLXCol and availableFlagLXCol are derived as follows: If tile_group_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are equal to 0. otherwise (tile_group _temporal_mvp_enabled_flag equals 1), the following is true: xColBr=xCb+cbWidth (8-566) yColBr=yCb+cbHeight (8-567)

[0151] yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY and yColBr is equal to SubPicBottomBorderInP If xColBr is less than SubPicRightBorderInPic, and xColBr is less than SubPicRightBorderInPic, then the following is true: The variable colCb is the internal (xColBr>>3) of the picture at the same position specified by ColPic. )<<3,(yColBr>>3)<<3) Specifies that the luma position (xColCb, yColCb) is locked to the same position specified by ColPic. The top-left luma sample of a picture is compared to the co-located luma sample specified by colCb. The upper left sample of the coding block is set equal to the The derivation process of the rule is performed with currCb, colCb, (xColCb, yColCb), refIdxLX set equal to 0. , and sbFlag as input, and the output is mvLXCol and availableFlagLXCol l. Otherwise, both components of mvLXCol are set equal to 0 and avail pic_width_in_luma_samples is set equal to 0. Replace all occurrences of pic_height_in_luma_samples with Pi icWidthInLumaSamples. Replace with cHeightInLumaSamples.

[0152] In a second exemplary embodiment, the syntax of the sequence parameter set RBSP and The semantics are as follows: [Table 10]

[0153] subpic_id_len_minus1 plus 1 is the syntax element subpic_id[i] in SPS. the number of bits used to represent the SPS, the spps_subpic_id in the SPPS that references the SPS, and Specifies the tile_group_subpic_id in the tile group header that references the subpic and SPS. The value of c_id_len_minus1 must be in the range Ceil(Log2(num_subpic_minus1+3)) to 8, inclusive. The overlapping between subpicture[i] for i between 0 and num_subpic_minus1 inclusive The absence of overlapping is a requirement for bitstream conformance. Each subpicture is It may be a sub-picture.

[0154] The general semantics of a tile group header are: tile_group_subpi c_id identifies the subpicture to which the tile group belongs. Length of tile_group_subpic_id is subpic_id_len_minus1+1 bits. A tile_group_subpic_id equal to 1 means that the tile Indicates that the group does not belong to any subpicture.

[0155] In the third exemplary embodiment, the syntax and semantics of the NAL unit header are as follows: As stated above. [Table 11]

[0156] nuh_subpicture_id_len is used to represent the syntax element that specifies the subpicture ID. Specifies the number of bits used. When the value of nuh_subpicture_id_len is greater than 0, The first nuh_subpicture_id_len bits after reserved_zero_4bits are used to denote the NAL unit Specifies the ID of the subpicture to which the payload belongs. If nuh_subpicture_id_len is greater than 0, When the value of nuh_subpicture_id_len is less than the value of subpic_id_len_minus1 in the active SPS, The value of nuh_subpicture_id_len for non-VCL NAL units shall be equal to the following: If nal_unit_type is equal to SPS_NUT or PPS_NUT, nuh_subpicture e_id_len shall be equal to 0. nuh_reserved_zero_3bits shall be equal to "000". The decoder shall ignore NAL units whose nuh_reserved_zero_3bits value is not equal to "000". shall be ignored (e.g., removed from the bitstream and discarded).

[0157] In a fourth exemplary embodiment, the subpicture nesting syntax is as follows: do. [Table 12]

[0158] all_sub_pictures_flag equal to 1 indicates that the nested SEI message includes all subpictures. all_sub_pictures_flag equal to 1 indicates that the subpictures apply to all nested SEI The subpicture to which the message applies is explicitly signaled by a subsequent syntax element. nesting_num_sub_pictures_minus1 plus 1 specifies the number of nesting pictures. Specifies the number of subpictures to which the specified SEI message applies. _id[i] is the subpicture of the ith subpicture to which the nested SEI message applies The nesting_sub_picture_id[i] syntax element is Ceil(Log2(nesting_num_sub _pictures_minus1+1)) bits. sub_picture_nesting_zero_bit is equal to 0. It shall be considered as new.

[0159] 9 is a schematic diagram of an example video coding device 900. The device 900 is suitable for implementing the disclosed examples / embodiments as described herein. The video coding device 900 transmits data upstream over a network. and / or a transmitter and / or a receiver for communicating the , downstream port 920, upstream port 950, and / or transceiver The video coding device 900 also includes a transceiver unit (Tx / Rx) 910. a processor 930 including a logic unit and / or central processing unit (CPU) for and memory 932 for storing the data. Video coding device 900 also includes an electrical Components, Optical-Electrical (OE) Components, Electro-Optical (EO) Components, and / or or telecommunications networks, optical communications networks, or wireless communications networks Upstream port 950 and / or downstream port 960 for communication of data over the network The video camera may also include a wireless communication component coupled to the wireless port 920. Reading device 900 also includes input and / or output ports for communicating data to and from a user. The system may also include an output (I / O) device 960. The I / O device 960 is used to display video data. output devices such as a display for outputting the audio data and a speaker for outputting the audio data. The I / O devices 960 may also include input devices such as a keyboard, mouse, trackball, etc. device and / or a corresponding interface for interacting with such output device. It may include a face.

[0160] The processor 930 is implemented in hardware and software. A processor consists of one or more CPU chips, cores (for example, as a multi-core processor), and field processors. Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), and Digital Signal Processors The processor 930 may be implemented as a digital signal processor (DSP). 920, Tx / Rx 910, upstream port 950, and memory 932. The decoder 930 includes a coding module 914. The coding module 914 The method 100 may utilize the stream 500, the picture 600, and / or the picture 800. Implementing the disclosed embodiments described above, such as 000, 1100, and / or mechanism 700 The coding module 914 may also implement any other method / mechanism described herein. Furthermore, the coding module 914 may be implemented as a codec system 200, an encoder, For example, the coding module may implement the decoder 300 and / or the decoder 400. The controller 914 signals and / or obtains the location and size of the subpictures in the SPS. In another example, the coding module 914 may be utilized to Unless the subpicture is located at the right border of the picture or the bottom border of the picture, respectively You may constrain the subpicture width and height to be multiples of the CTU size. In another example, the coding module 914 may include pictures without gaps or overlaps. In another example, the coding module 914 may constrain the subpicture to , some subpictures are temporal motion constrained subpictures and others are not. is not utilized to signal and / or obtain data indicating In another example, the coding module 914 may use the complete subpicture ID in the SPS. a subpicture for signaling a subpicture set containing the corresponding slice; In another example, the coding module 914 may include a slice ID in each slice header. , the level of each sub-picture may be signaled. The controller 914 provides additional functionality to the video coding device 900 and To reduce processing overhead when coding separately and / or To improve the coding efficiency, some processing is avoided. Module 914 improves the functionality of video coding device 900 and Furthermore, the coding module 914 can be used in different situations. Alternatively, the coding module may be implemented as a conversion module. Module 914 may be written as instructions stored in memory 932 and executed by processor 930 (e.g., For example, as a computer program product stored on a non-transitory medium.

[0161] The memory 932 can be a disk, tape drive, solid state drive, read-only Memory (ROM), Random Access Memory (RAM), Flash Memory, Ternary Content Addressable Memory (TCAM) , static random access memory (SRAM), and more. The memory 932 stores programs when they are selected for execution. for storing instructions and data that are read during program execution , may be used as an overflow data storage device.

[0162] FIG. 10 shows the sub-pictures 522, 523, 622, 722, and / or 822, the bitstream 500 and / or the sub-bitstreams 10 is a flowchart of an example method 1000 for encoding a bitstream, such as stream 501. The method 1000 includes a codec system 200, an encoder 300, and / or a Or it may be utilized by an encoder such as video coding device 900.

[0163] The method 1000 includes an encoder receiving a video sequence including a plurality of pictures, for example, For example, determining to encode a video sequence into a bitstream based on user input. The video sequence may be divided into pictures for further segmentation before encoding. In step 1001, a picture is divided into a number of sub-pictures / images / frames. When partitioning the subpictures, adaptive size constraints are applied. Each subpicture contains a subpicture width and a subpicture height. contains a right border that does not coincide with the right border of the picture, the current subpicture's subpicture The width of each subpicture is constrained to be an integer multiple of the CTU (e.g., subpicture 822 other than 822b). Therefore, at least one of the subpictures (e.g., subpicture 822b) An integer number of CTUs in size, where each subpicture contains a right border that coincides with the right border of the picture. In addition, the current subpicture may contain subpictures with widths other than double the picture width. The subpicture height of the current subpicture is CT is constrained to be an integer multiple of U (e.g., subpicture 822 other than 822c). , at least one of the subpictures (e.g., subpicture 822c) When a pixel contains a bottom border that coincides with the bottom border of the picture, the pixel size is not an integer multiple of the CTU size. This adaptive size constraint may be a multiple of the CTU size, but not an integer multiple of the CTU size. from pictures containing a picture width and / or a picture height that is not an integer multiple of the CTU size Allows subpictures to be partitioned. CTU size is measured in units of luma samples. This may also be done.

[0164] In step 1003, one or more of the sub-pictures are encoded into a bitstream. In step 1005, the bitstream is stored for communication to a decoder. The bitstream may then be transmitted to a decoder as desired. In this example, the sub-bitstreams may be extracted from the encoded bitstream. In such cases, the transmitted bitstream is a sub-bitstream. In the example, the encoded bitstream is used to extract the sub-bitstreams at the decoder. In yet another example, the encoded bitstream may be transmitted for sub-sampling. In any of these examples, the bitstream may be decoded and displayed without extraction. However, adaptive size constraints are not supported for pictures with heights or widths that are not multiples of the CTU size. This enhances the performance of the encoder by allowing sub-pictures to be partitioned from one another.

[0165] FIG. 11 shows the sub-pictures 522, 523, 622, 722, and / or 822, the bitstream 500 and / or the sub-bitstreams 11 is a flowchart of an example method 1100 for decoding a bitstream such as stream 501. The method 1100, when performing the method 100, includes the codec system 200, the decoder 400, and / or may be utilized by a decoder such as video coding device 900. For example, , method 1100 is applied to decode the bitstream produced as a result of method 1000. This may be done.

[0166] The method 1100 begins when the decoder begins receiving a bitstream containing subpictures. The bitstream may contain a complete video sequence, or it may contain only a few bits. The sub-bitstream contains a reduced set of sub-pictures for separate extraction. In step 1101, a bitstream is received. A picture stream consists of one or more subpictures partitioned from a picture according to adaptive size constraints. Each subpicture comprises a subpicture width and a subpicture height. When the current subpicture contains a right border that does not coincide with the right border of the picture, The subpicture width of a picture is constrained to be an integer multiple of the CTU (e.g., for formats other than 822b). Subpicture 822). Therefore, at least one of the subpictures (e.g., Picture 822b) is C when each subpicture contains a right border that coincides with the right border of the picture. It may contain subpicture widths that are not integer multiples of the TU size. contains a bottom border that does not coincide with the bottom border of the picture, the current subpicture's subpicture The height of a subpicture is constrained to be an integer multiple of the CTU (e.g., subpictures 82c other than 822c). 2). Therefore, at least one of the subpictures (e.g., subpicture 822c) , when each subpicture contains a bottom border that matches the bottom border of the picture, the CTU size is This adaptive size constraint may include subpicture heights that are not multiples of the CTU size. picture width that is not an integer multiple of the CTU size and / or picture height that is not an integer multiple of the CTU size. The CTU size is the number of luma samples. It may be measured in units of

[0167] In step 1103, the bitstream is processed to obtain one or more sub-pictures. In step 1105, one or more The sub-pictures are decoded and the video sequence can then be transmitted for display. Therefore, adaptive size constraints are used to restrict the size of pictures with heights or widths that are not multiples of the CTU size. This allows the decoder to separate the subpictures from the CTU size. Separate subpicture extraction and / or This allows for the use of subpicture-based features such as adaptive filtering or display. Applying a general size constraint improves the decoder's capabilities.

[0168] FIG. 12 shows the sub-pictures 522, 523, 622, 722, and / or 822, the bitstream 500 and / or the sub-bitstreams 1 is a schematic diagram of an example system 1200 for signaling a bitstream, such as system 501. The system 1200 includes a codec system 200, an encoder 300, a decoder 400, and a and / or implemented by an encoder and a decoder, such as video coding device 900. Additionally, system 1200 may, when performing methods 100, 1000, and / or 1100, It may be used for.

[0169] The system 1200 includes a video encoder 1202. The video encoder 1202 encodes each sub-picture When a subpicture contains a right border that does not coincide with the right border of the picture, each subpicture is Partitioning a picture into multiple sub-pictures so that each sub-picture width is an integer multiple of The video encoder 1202 further comprises a partitioning module 1201 for partitioning sub-pictures. a coding module 1203 for coding one or more of them into a bitstream. The video encoder 1202 also stores the bitstream for communication to the decoder. The video encoder 1202 further comprises a storage module 1205 for storing the sub-pictures. a transmission module 1207 for transmitting a bitstream including the video to a decoder. The encoder 1202 may be further configured to perform any of the steps of the method 1000. good.

[0170] The system 1200 also includes a video decoder 1210. The video decoder 1210 is configured to generate a video stream in which each sub-picture is Each subpicture has a coding tree when it contains a right border that does not coincide with the right border of the picture. A subpicture is a separate image from a picture, containing a subpicture width that is an integer multiple of the CTU size. a receiving module for receiving a bitstream comprising one or more divided sub-pictures; The video decoder 1210 further analyzes the bitstream to generate one or more The video decoder 1210 further comprises a parsing module 1213 for obtaining the sub-pictures. and a decoding module for decoding one or more sub-pictures to create a video sequence. The video decoder 1210 also includes a video decoder 1215. The video decoder 1210 further comprises a transfer module 1217 for transferring the data from the video decoder 1210 to the video decoder 1210. The system may be configured to execute any of the following:

[0171] Except for the wire, wiring, or other medium between the first and second components When there are no intervening components, the first component directly connects to the second component. A wire, wiring, or other connection is made between the first and second components. When there is an intervening component other than the medium, the first component The term "coupled" and its variations refer to a directly coupled component. The use of the term "about" includes both indirect and indirect binding. Unless otherwise specified, the range means a range including ±10% of the number that follows.

[0172] The steps of the exemplary methods described herein may not necessarily be performed in the order described. It is understood that no particular order is required and that the order of steps in such methods is merely exemplary. It should also be understood that additional steps may be included in such a method. Optionally, some steps may be omitted or completed in a manner consistent with various embodiments of the present disclosure. They may be combined.

[0173] Although several embodiments have been provided in this disclosure, the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. It may be understood that the present example is for illustrative purposes and not for limitation. and the intention is not to be limited to the details given herein. For example, in another system, various elements or components may be combined or They may be combined or some features may be omitted or not implemented.

[0174] Additionally, various embodiments may be described or illustrated as separate or distinct. The techniques, systems, subsystems, and methods described herein may be modified without departing from the scope of this disclosure. , may be combined or integrated with other systems, components, techniques, or methods. Other examples of variations, substitutions, and alterations are ascertainable by one of ordinary skill in the art and are disclosed herein. may be made without departing from the spirit and scope thereof. [Explanation of symbols]

[0175] 200 Codec System 201 segmented video signal 211 General-Purpose Coder Control Component 213 Transform Scaling and Quantization Components 215 Intra-picture Estimation Component 217 Intra-picture Prediction Component 219 Motion Compensation Component 221 Motion Estimation Component 223 Decoded Picture Buffer Component 225 In-Loop Filter Components 227 Filter Control Analysis Component 229 Scaling and Inverse Transformation Components 231 Header Formatting and CABAC Components 301 Segmented Video Signal 313 Transform and Quantize Components 317 Intra-picture Prediction Component 321 Motion Compensation Component 323 Decoded Picture Buffer Component 325 In-Loop Filter Components 329 Inverse Transform and Quantization Components 331 Entropy Coding Component 417 Intra-picture Prediction Component 421 Motion Compensation Component 423 Decoded Picture Buffer Component 425 In-Loop Filter Components 429 Inverse Transform and Quantization Components 433 Entropy Decoding Component 500 bitstream 501 Sub-Bitstream 510 SPS 512 PPS 514 slice header 515 SEI Message 520 Image Data 521 Pictures 522 Subpictures 523 Subpictures 524 slices 525 slices 531 Subpicture Size 532 Subpicture Position 533 Subpicture ID 534 Motion Constraint Subpicture Flag 535 Subpicture Level 600 pictures 622 Subpicture 631 Subpicture Size 631a Subpicture Width 631b Subpicture Height 632 position 633 Subpicture ID 634 Time-Motion Constrained Subpictures 642 Upper left sample 700 mechanism 722 Subpicture 724 slices 733 Subpicture ID 742 top left corner 801 Picture right border 802 Picture bottom border 822 Subpicture 822b, 822c subpictures 825 CTU 900 Video Coding Device 910 Transmitter / Receiver 914 Coding Module 920 downstream ports 930 processor 932 memory 950 upstream ports 960 I / O devices 1200 System 1201 Division Module 1202 Video Encoder 1203 Encoding Module 1205 Memory Module 1207 Transmitting Module 1210 Video Decoder 1211 Receiver Module 1213 Analysis Module 1215 Decryption Module 1217 Transfer Module

Claims

1. 1. A method implemented in a decoder, comprising: receiving a bitstream comprising coded data of one or more sub-pictures partitioned from a picture and a sequence parameter set (SPS), wherein a coding tree unit (CTU) at a right boundary of the picture is incomplete when the picture has a picture width that is not an integer multiple of a CTU size, or a CTU at a bottom boundary of the picture is incomplete when the picture has a picture height that is not an integer multiple of the CTU size; the SPS comprises a subpicture identifier (ID), a subpicture size, and a subpicture position for a current subpicture of the one or more subpictures; the bitstream further comprising a flag, wherein the flag equal to 1 specifies that in-loop filtering operations across sub-picture boundaries are enabled, and wherein the flag equal to 0 specifies that in-loop filtering operations across sub-picture boundaries are disabled; parsing the subpicture ID, the subpicture size, the subpicture position, and the flag from the bitstream for the current subpicture; decoding the current subpicture based on the subpicture ID, the subpicture size, the subpicture position, and the flag.

2. The method described in claim 1, wherein when a subpicture of the one or more subpictures includes a bottom border that does not coincide with the bottom border of the picture, the subpicture has a subpicture height that is an integer multiple of the CTU size.

3. A method as described in claim 1 or 2, wherein when a subpicture of the one or more subpictures includes a right border that does not coincide with the right border of the picture, the subpicture has a subpicture width that is an integer multiple of the CTU size.

4. A method described in any one of claims 1 to 3, wherein when a subpicture of the one or more subpictures includes a bottom border that coincides with the bottom border of the picture, the subpicture has a subpicture height that is not an integer multiple of the CTU size.

5. 1. A method implemented in an encoder, comprising: partitioning a picture into one or more sub-pictures, wherein a coding tree unit (CTU) at a right boundary of the picture is incomplete when the picture has a picture width that is not an integer multiple of a CTU size, or a CTU at a bottom boundary of the picture is incomplete when the picture has a picture height that is not an integer multiple of the CTU size; encoding the one or more sub-pictures into a bitstream; encoding a sub-picture identifier (ID), a sub-picture size, and a sub-picture position for a current sub-picture of the one or more sub-pictures into a sequence parameter set (SPS) of the bitstream; encoding a flag into the bitstream, wherein the flag equal to 1 specifies that in-loop filtering operations across sub-picture boundaries are enabled, and wherein the flag equal to 0 specifies that in-loop filtering operations across sub-picture boundaries are disabled.

6. The method of claim 5 , further comprising storing the bitstream in a memory of the encoder for communication to a decoder.

7. A method as described in claim 5 or 6, wherein when at least one of the one or more subpictures includes a bottom border that does not coincide with the bottom border of the picture, the at least one subpicture includes a subpicture height that is an integer multiple of the CTU size.

8. A method described in any one of claims 5 to 7, wherein when at least one of the one or more subpictures includes a right border that coincides with the right border of the picture, the at least one subpicture includes a subpicture width that is not an integer multiple of the CTU size.

9. 9. The method of claim 5, wherein when at least one of the one or more subpictures includes a bottom border that coincides with the bottom border of the picture, the at least one subpicture includes a subpicture height that is not an integer multiple of the CTU size.

10. 10. The method of claim 5, wherein the CTU size is measured in units of luma samples.

11. 10. A non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, causes the video coding device to perform the method of any one of claims 1 to 4.

12. A non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to perform a method according to any one of claims 5 to 10.

13. A decoder comprising:

1. A receiving unit configured to receive a bitstream comprising coded data of one or more sub-pictures partitioned from a picture and a sequence parameter set (SPS), wherein a coding tree unit (CTU) at a right boundary of the picture is incomplete when the picture has a picture width that is not an integer multiple of a CTU size, or a CTU at a bottom boundary of the picture is incomplete when the picture has a picture height that is not an integer multiple of the CTU size, the SPS comprises a subpicture identifier (ID), a subpicture size, and a subpicture position for a current subpicture of the one or more subpictures; a receiving unit, the bitstream further comprising a flag, wherein the flag equal to 1 specifies that in-loop filtering operations across sub-picture boundaries are enabled and the flag equal to 0 specifies that in-loop filtering operations across sub-picture boundaries are disabled; a decoding unit, Parsing the subpicture ID, the subpicture size, the subpicture position, and the flag from the bitstream for the current subpicture; a decoding unit configured to decode the current subpicture based on the subpicture ID, the subpicture size, the subpicture position, and the flag.

14. 14. The decoder of claim 13, wherein the decoder is further configured to perform the method of any one of claims 2 to 4.

15. An encoder comprising: a partition unit configured to partition a picture into one or more sub-pictures, wherein a coding tree unit (CTU) at a right boundary of the picture is incomplete when the picture has a picture width that is not an integer multiple of a CTU size, or a CTU at a bottom boundary of the picture is incomplete when the picture has a picture height that is not an integer multiple of the CTU size; A coding unit, comprising: encoding the one or more sub-pictures into a bitstream; encoding a sub-picture identifier (ID), a sub-picture size, and a sub-picture position for a current sub-picture of the one or more sub-pictures into a sequence parameter set (SPS) of the bitstream; configured to encode the flag into the bitstream; a coding unit, wherein the flag equal to 1 specifies that in-loop filtering operations across sub-picture boundaries are enabled, and wherein the flag equal to 0 specifies that in-loop filtering operations across sub-picture boundaries are disabled.

16. The encoder of claim 15, further comprising a storage unit configured to store the bitstream for communication to a decoder.

17. 17. The encoder of claim 15 or 16, wherein the encoder is further configured to perform the method of any one of claims 7 to 10.

18. A method for storing a bitstream, the bitstream comprising coded data of one or more subpictures partitioned from a picture and a sequence parameter set (SPS), wherein when the picture has a picture width that is not an integer multiple of a coding tree unit (CTU) size, a CTU at the right boundary of the picture is incomplete, or when the picture has a picture height that is not an integer multiple of the CTU size, a CTU at the bottom boundary of the picture is incomplete; the SPS comprises a subpicture identifier (ID), a subpicture size, and a subpicture position for a current subpicture of the one or more subpictures; 10. The method of claim 9, wherein the bitstream further comprises a flag, wherein the flag equal to 1 specifies that in-loop filtering operations across sub-picture boundaries are enabled, and the flag equal to 0 specifies that in-loop filtering operations across sub-picture boundaries are disabled.

19. a receiving unit configured to receive a bitstream to be decoded; a transmitting unit coupled to the receiving unit, the transmitting unit configured to transmit the decoded image to a display unit; a storage unit coupled to at least one of the receiving unit or the transmitting unit, the storage unit configured to store instructions; and a processing unit coupled to the storage unit, the processing unit being configured to execute the instructions stored in the storage unit to perform the method of any one of claims 1 to 4.

20. A decoder comprising processing circuitry for implementing the method of any one of claims 1 to 4.

21. An encoder comprising processing circuitry for implementing the method of any one of claims 5 to 10.

22. An apparatus for storing and transmitting a bitstream, comprising: a receiver; a processor; a transmitter; and a storage medium, wherein the receiver is configured to receive the bitstream; the storage medium is configured to store the bitstream; and the transmitter is configured to transmit the bitstream; The bitstream comprises coded data of one or more sub-pictures partitioned from a picture and a sequence parameter set (SPS), and when the picture has a picture width that is not an integer multiple of a coding tree unit (CTU) size, a CTU at a right boundary of the picture is incomplete, or when the picture has a picture height that is not an integer multiple of the CTU size, a CTU at a bottom boundary of the picture is incomplete; the SPS comprises a subpicture identifier (ID), a subpicture size, and a subpicture position for a current subpicture of the one or more subpictures; the bitstream further comprises a flag, wherein the flag equal to 1 specifies that in-loop filtering operations across sub-picture boundaries are enabled, and the flag equal to 0 specifies that in-loop filtering operations across sub-picture boundaries are disabled.

Citation Information

Patent Citations

  • Concept for picture / video data streams allowing efficient reducibility or efficient random access

    WO2017137444A1

  • Advanced video data stream extraction and multi-resolution video transmission

    WO2018172234A2