High-Level Syntax for Video Encoding and Decoding

By omitting specific syntax element parsing in the VVC bitstream decoding process, particularly in raster scan slice mode, the complexity and bitrate are reduced without degrading coding performance, addressing the challenges of high-level syntax changes in VVC.

JP7733777B2Active Publication Date: 2025-09-03CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024096894
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-20
Filing Date
2024-06-14
Publication Date
2025-09-03
Estimated Expiration
2041-03-17

AI Technical Summary

Technical Problem

The changes in high-level syntax structures due to the introduction of new tools in the Versatile Video Coding (VVC) standard affect the bitstream bitrate, particularly in low-delay and low-bitrate applications, requiring improvements to reduce complexity without degrading coding performance.

Method used

A method for decoding video data from a bitstream that omits parsing certain syntax elements, such as slice addresses and the number of tiles, when the picture header is signaled in the slice header, reducing parsing complexity and bitrate, especially in raster scan slice mode.

Benefits of technology

This approach reduces parsing complexity and bitrate, particularly beneficial for low-delay and low-bitrate applications, while maintaining coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007733777000023
    Figure 0007733777000023
  • Figure 0007733777000024
    Figure 0007733777000024
  • Figure 0007733777000025
    Figure 0007733777000025
Patent Text Reader

Abstract

To provide a method for decoding video data from a bitstream.SOLUTION: A bitstream has video data corresponding to one or more slices. The slices each have one or more tiles. The bitstream has a picture header having syntax elements to be used when decoding one or more slices, and a slice header having syntax elements to be used when decoding a slice. The method includes: parsing the syntax elements; when the slice (or picture) includes a plurality of tiles, if the syntax elements indicating that the picture header is signaled in the slice header are parsed, omitting parsing the syntax elements indicating the address of the slice; and decoding the bitstream by using the syntax elements.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to video encoding and decoding, and in particular to high-level syntax used in bitstreams. [Background technology]

[0002] Recently, the Joint Video Experts Team (JVET), a collaborative team formed by MPEG and ITU-T Study Group 16's VCEG, began research into a new video coding standard called VVC (Versatile Video Coding). The goal of VVC is to provide significant improvements in compression performance over the existing HEVC standard (i.e., typically double the previous standard) and to be completed in 2020. Primary target applications and services include, but are not limited to, 360-degree and high dynamic range (HDR) video. Overall, JVET evaluated responses from 32 organizations using formal subjective testing conducted by an independent testing laboratory. Several proposals demonstrated compression efficiency improvements of typically 40% or more compared to using HEVC. This was particularly evident for ultra-high-definition (UHD) video test material. Therefore, we expect compression efficiency improvements far exceeding the 50% target for the final standard. Summary of the Invention [Problem to be solved by the invention]

[0003] The JVET exploration model (JEM) uses all HEVC tools and introduces many new ones. These changes required changes to the bitstream structure, especially the high-level syntax, which can affect the overall bitstream bitrate. [Means for solving the problem]

[0004] The present invention relates to improvements in high-level syntax structures that reduce complexity without degrading coding performance.

[0005] According to a first aspect of the present invention, there is provided a method for decoding video data from a bitstream, wherein the bitstream comprises video data corresponding to one or more slices, each slice comprising one or more tiles; the bitstream comprising a picture header having syntax elements for use in decoding one or more slices, and a slice header having syntax elements for use in decoding a slice, the method comprising: parsing the syntax elements; if a slice (or picture) comprises multiple tiles, omitting parsing of syntax elements indicating slice addresses if a syntax element indicating that a picture header is signaled in a slice header is parsed; and decoding the bitstream using the syntax elements. According to another aspect, there is provided a method for decoding video data from a bitstream, wherein the bitstream has video data corresponding to one or more slices, each slice having one or more tiles; the bitstream having a picture header with syntax elements for use in decoding the one or more slices and a slice header with syntax elements for use in decoding the slices, the method comprising: parsing syntax elements; and, if a slice or picture contains multiple tiles, omitting parsing syntax elements indicating slice addresses if a syntax element indicating that a picture header is signaled in a slice header is parsed; and decoding the bitstream using the syntax elements.According to another aspect, there is provided a method for decoding video data from a bitstream, wherein the bitstream has video data corresponding to one or more slices, each slice having one or more tiles; the bitstream has a picture header with syntax elements for use in decoding one or more slices and a slice header with syntax elements for use in decoding a slice; the bitstream has a syntax element with a value indicating that a slice or picture has multiple tiles, and the bitstream is constrained to include a syntax element indicating that a stack element indicating an address of a slice shall not be parsed if the bitstream has a syntax element indicating that a picture header is signaled in a slice header, the method comprising: decoding the bitstream using the syntax elements.

[0006] Therefore, especially for low-delay and low-bitrate applications, when the picture header is within the slice header, the slice address is not parsed, which reduces the bit rate. Furthermore, when the picture is signaled in the slice header, the parsing complexity can be reduced. In one embodiment, the omission is performed (only) when raster scan slice mode is used to decode the slice, which reduces parsing complexity while still allowing some bitrate reduction. The omission may further include omission of parsing a syntax element that indicates the number of tiles in a slice. Thus, further bitrate reduction may be achieved.

[0007] In a second aspect, a method is provided for decoding video data from a bitstream, where the bitstream includes video data corresponding to one or more slices, each slice including one or more tiles. The bitstream includes a picture header with syntax elements for use in decoding the one or more slices, and a slice header with syntax elements for use in decoding the slice. The decoding includes parsing one or more syntax elements, and if the slice (or picture) includes multiple tiles, skipping parsing a syntax element indicating the number of tiles in the slice if a syntax element indicating that the picture header is signaled in the slice header is parsed, and decoding the bitstream using the syntax elements. According to another aspect, a method is provided for decoding video data from a bitstream, where the bitstream includes video data corresponding to one or more slices, each slice including one or more tiles. The bitstream includes a picture header with syntax elements for use in decoding the one or more slices, and a slice header with syntax elements for use in decoding the slice. and decoding includes parsing one or more syntax elements, and, if the slice or picture contains multiple tiles, omitting parsing a syntax element indicating the number of tiles in the slice if a syntax element indicating that a picture header is signaled in a slice header is parsed, and decoding the bitstream using the syntax elements. According to yet another aspect of the invention, there is provided a method for decoding video data from a bitstream, wherein the bitstream includes video data corresponding to one or more slices, each slice including one or more tiles. The bitstream includes a picture header with syntax elements for use in decoding the one or more slices, and a slice header with syntax elements for use in decoding the slices.and wherein the bitstream is constrained to have a syntax element indicating that a syntax element indicating the number of tiles in the slice is not parsed if the bitstream has a syntax element with a value indicating that a slice or picture has multiple tiles and the bitstream has a syntax element indicating that a picture header is signaled in a slice header, and the method comprises decoding the bitstream using the syntax elements.

[0008] Therefore, the bit rate can be reduced, which is particularly advantageous for low-delay and low-bit rate applications where the number of tiles does not need to be transmitted. The elision can be performed (only) when raster scan slice mode is used to decode the slice, which reduces the parsing complexity but still allows some bitrate reduction.

[0009] Moreover, the method may further include parsing a syntax element indicating the number of tiles in a picture, and determining the number of tiles in a slice based on the number of tiles in the picture indicated by the parsed syntax element. This is advantageous because it allows the number of tiles in a slice to be easily predicted when the picture header is signaled in the slice header without requiring further signaling. The omission may further include omitting parsing syntax elements that indicate slice addresses, thus further reducing the bit rate.

[0010] According to a third aspect of the present invention, there is provided a method for decoding video data from a bitstream, wherein the bitstream comprises video data corresponding to one or more slices, each slice comprising one or more tiles. The bitstream also comprises a picture header comprising syntax elements used when decoding the one or more slices, and a slice header comprising syntax elements used when decoding the slices. The decoding step then comprises parsing one or more syntax elements, and, if the slice (or picture) comprises multiple tiles, omitting parsing a syntax element indicating a slice address if the number of tiles in the slice equals the number of tiles in the picture, and decoding the bitstream using the syntax element. This exploits the insight that if the number of tiles in a slice equals the number of tiles in a picture, it is ensured that the current picture includes only one slice. Therefore, omitting the slice address can improve bitrate and reduce parsing and / or encoding complexity.

[0011] The omission may be performed (only) when raster scan slice mode is used to decode the slice, thus reducing complexity while still providing some bitrate reduction.

[0012] The decoding further includes parsing, in the slice, a syntax element indicating the number of tiles in the slice, and parsing, in the picture parameter set, a syntax element indicating the number of tiles in the picture, and omitting parsing of the syntax element indicating the slice address is based on the parsed syntax element. The decoding further includes parsing a syntax element in the slice that indicates the number of tiles in the slice before one or more syntax elements for signaling the slice address. The decoding may further include parsing a syntax element in the slice that indicates whether a picture header is signaled in the slice header, and determining (inferring) that the number of tiles in the slice is equal to the number of tiles in the picture if the parsed syntax element indicates that a picture header is signaled in the slice header.

[0013] In a fourth aspect, a method for decoding video data from a bitstream is provided, where the bitstream has video data corresponding to one or more slices, each slice having one or more tiles. The bitstream includes a picture header having syntax elements used when decoding the one or more slices and a slice header having syntax elements used when decoding the slice. The decoding includes parsing the one or more syntax elements and, if the syntax elements indicate that a raster scan decoding mode is enabled for the slice, decoding at least one of a slice address and a number of tiles in the slice from the one or more syntax elements. If the raster scan decoding mode is enabled, decoding at least one of the slice address and the number of tiles in the slice includes decoding the bitstream using the syntax elements, independent of the number of tiles in the picture. Therefore, the complexity of parsing a slice header can be reduced.

[0014] In a fifth aspect according to the present invention there is provided a method comprising the first and second aspects.

[0015] In a sixth aspect according to the present invention there is provided a method comprising the first, second and third aspects.

[0016] According to a seventh aspect of the present invention, there is provided a method for encoding video data into a bitstream, wherein the bitstream comprises video data corresponding to one or more slices, each slice comprising one or more tiles, and wherein the bitstream comprises a picture header comprising syntax elements used when decoding the one or more slices, and a slice header comprising syntax elements used when decoding the slice. The encoding step comprises determining one or more syntax elements for encoding the video data, omitting encoding of syntax elements indicating slice addresses if the picture header is a syntax element signaled in a slice header if the slice (or picture) comprises multiple tiles, and encoding the video data using the syntax elements. According to an additional aspect of the present invention, there is provided a method for encoding video data into a bitstream, wherein the bitstream comprises video data corresponding to one or more slices, each slice comprising one or more tiles, and wherein the bitstream comprises a picture header comprising syntax elements used when decoding the one or more slices, and a slice header comprising syntax elements used when decoding the slice. and encoding comprises determining one or more syntax elements for encoding the video data, omitting encoding of syntax elements indicating slice addresses if the picture header is a syntax element signaled in a slice header if the slice or picture includes multiple tiles, and encoding the video data using the syntax elements. According to a further supplemental aspect of the present invention, there is provided a method for encoding video data into a bitstream, wherein the bitstream comprises video data corresponding to one or more slices, each slice comprising one or more tiles, and the bitstream comprises a picture header comprising syntax elements used when decoding the one or more slices, and a slice header comprising syntax elements used when decoding the slices.and wherein the bitstream is constrained to have a syntax element indicating that syntax elements indicating slice addresses are not parsed if the bitstream has a syntax element with a value indicating that a slice or picture has multiple tiles and the bitstream has a syntax element indicating that a picture header is signaled in the slice header, and the method includes encoding the video data using the syntax elements.

[0017] In one or more embodiments, the skipping is performed (only) when a raster scan slice mode is used to encode the slice. The omitting may further include omitting encoding a syntax element that indicates the number of tiles in a slice.

[0018] According to an eighth aspect of the present invention, there is provided a method for encoding video data into a bitstream, wherein the bitstream has video data corresponding to one or more slices, each slice having one or more tiles, wherein the bitstream has a picture header having syntax elements used when decoding the one or more slices and a slice header having syntax elements used when decoding the slice. The encoding step comprises determining one or more syntax elements for encoding the video data, omitting encoding a syntax element indicating the number of tiles in the slice if the slice includes multiple tiles and the picture header has determined a syntax element for encoding indicating that the number of tiles in the slice is signaled in the slice header, and encoding the video data using the syntax elements. According to yet an additional aspect of the present invention, there is provided a method for encoding video data into a bitstream, wherein the bitstream has video data corresponding to one or more slices, each slice having one or more tiles, wherein the bitstream has a picture header having syntax elements used when decoding the one or more slices and a slice header having syntax elements used when decoding the slice. and encoding the video data comprises determining one or more syntax elements for encoding the video data, omitting encoding a syntax element indicating the number of tiles in a slice if a syntax element for encoding indicates that the slice or picture contains multiple tiles and a picture header is signaled in the slice header, and encoding the video data using the syntax elements. According to a further complementary aspect of the present invention, there is provided a method for encoding video data into a bitstream, wherein the bitstream comprises video data corresponding to one or more slices, each slice comprising one or more tiles. The bitstream comprises a picture header comprising syntax elements for use in decoding the one or more slices, and a slice header comprising syntax elements for use in decoding the slices.If the bitstream has a syntax element with a value indicating that a slice or picture has multiple tiles, and the bitstream has a determined syntax element for encoding that a picture header is signaled in the slice header, the bitstream is constrained to also have a syntax element indicating that a syntax element indicating the number of tiles in a slice is not parsed, and the method includes encoding the video data using the syntax element.

[0019] In one embodiment, the omission is performed (only) when a raster scan slice mode is used to encode the slice.

[0020] The encoding further includes encoding a syntax element indicating a number of tiles in the picture, and the number of tiles in the slice is based on the number of tiles in the picture indicated by the parsed syntax element. The omitting may further include omitting encoding a syntax element that indicates an address of the slice.

[0021] According to a ninth aspect of the present invention, there is provided a method for encoding video data into a bitstream, wherein the bitstream comprises video data corresponding to one or more slices, each slice comprising one or more tiles, and wherein the bitstream comprises a picture header having syntax elements for use in decoding the one or more slices, and a slice header having syntax elements for use in decoding the slices, and wherein the encoding comprises determining one or more syntax elements, and, if the slice (or picture) comprises multiple tiles, omitting encoding of syntax elements indicating slice addresses if the number of tiles in the slice is equal to the number of tiles in the picture, and encoding the video data using the syntax elements.

[0022] In one or more embodiments, the omission is performed (only) when a raster scan slice mode is used to encode the slice.

[0023] The encoding step further includes encoding, in a slice, a syntax element indicating the number of tiles in the slice, and encoding, in a picture parameter set, a syntax element indicating the number of tiles in the picture, and whether or not to omit encoding the syntax element indicating the slice address is based on the value of the encoded syntax element. The encoding may further include encoding a syntax element in the slice indicating the number of tiles in the slice before one or more syntax elements for signaling a slice address. The encoding further includes encoding a syntax element in the slice that indicates whether a picture header is signaled in the slice header, and determining that the number of tiles in the slice is equal to the number of tiles in the picture if the encoded syntax element indicates that a picture header is signaled in the slice header.

[0024] According to a tenth aspect of the present invention, there is provided a method for encoding video data into a bitstream, wherein the bitstream comprises video data corresponding to one or more slices, each slice comprising one or more tiles, and the bitstream comprises a picture header having syntax elements for use in decoding the one or more slices and a slice header having syntax elements for use in decoding the slice, and the encoding comprises determining one or more syntax elements for encoding the video data, and if the determined syntax elements for encoding indicate that a raster scan decoding mode is enabled for the slice, encoding a syntax element indicating at least one of a slice address and a number of tiles in the slice, wherein if the raster scan decoding mode is enabled for the slice, encoding at least one of the slice address and the number of tiles in the slice from the one or more syntax elements is independent of the number of tiles in the picture, and encoding the bitstream using the syntax elements.

[0025] In an eleventh aspect according to the present invention there is provided a method comprising the seventh and eighth aspects.

[0026] In a twelfth aspect according to the present invention there is provided a method comprising the seventh, eighth and ninth aspects.

[0027] According to a thirteenth aspect of the present invention, there is provided a decoder for decoding video data from a bitstream, the decoder being configured to perform the method of any of the first to sixth aspects.

[0028] According to a fourteenth aspect of the present invention, there is provided an encoder for encoding video data into a bitstream, the encoder being configured to perform a method according to any of the seventh to twelfth aspects.

[0029] According to a fifteenth aspect of the present invention, there is provided a computer program which, when executed, causes the computer to perform the method of any of the first to twelfth aspects. The program may be provided by itself or may be carried on, by, or in a carrier medium. The carrier medium may be non-transitory, for example, a storage medium, in particular a computer-readable storage medium. The carrier medium may also be transitory, for example, a signal or other transmission medium. The signal may be transmitted over any suitable network, including the Internet. Further features of the present invention are characterized by the independent and dependent claims.

[0030] Any feature of one aspect of the invention may be applied to other aspects of the invention in any appropriate combination, in particular method aspects may be applied to apparatus aspects and vice versa. Additionally, features implemented in hardware may be implemented in software and vice versa, and any references to software and hardware features herein should be construed accordingly.

[0031] Any apparatus features described herein may also be provided as method features, and vice versa. As used herein, means- and function-features may alternatively be expressed in terms of their corresponding structures, such as a suitably programmed processor and associated memory.

[0032] It should also be understood that specific combinations of the various features described and defined in any embodiment of the present invention may be implemented and / or provided and / or used independently. [Brief explanation of the drawings]

[0033] Reference will now be made, by way of example, to the accompanying drawings in which: [Figure 1] FIG. 1 is a diagram for use in explaining the coding structure used in HEVC and VVC. [Figure 2]1 is a block diagram that schematically illustrates a data communications system in which one or more embodiments of the present invention may be implemented. [Figure 3] 1 is a block diagram illustrating components of a processing device in which one or more embodiments of the present invention may be implemented. [Figure 4] 3 is a flowchart illustrating steps of an encoding method according to an embodiment of the present invention. [Figure 5] 4 is a flow chart illustrating steps of a decoding method according to an embodiment of the present invention. [Figure 6] FIG. 1 illustrates the structure of a bitstream in an exemplary coding system VVC. [Figure 7] FIG. 10 illustrates another structure of a bitstream in an exemplary coding system VVC. [Figure 8] FIG. 1 illustrates Luma Modeling Chroma Scaling (LMCS). [Figure 9] FIG. 1 is a diagram showing sub-tools of LMCS. [Figure 10] A diagram of the raster scan slice mode and rectangular slice mode in the current VVC draft standard. [Figure 11] 1 shows a system comprising an encoder or decoder and a communication network according to an embodiment of the invention; [Figure 12] 1 is a schematic block diagram of a computing device for implementation of one or more embodiments of the present invention. [Figure 13] FIG. 1 illustrates a network camera system. [Figure 14] FIG. 1 illustrates a smartphone. DETAILED DESCRIPTION OF THE INVENTION

[0034] Figure 1 relates to the coding structure used in the High Efficiency Video Coding (HEVC) video standard. A video sequence 1 consists of a sequence of digital images i. Each such digital image is represented by one or more matrices. The matrix coefficients represent pixels.

[0035] An image 2 of a sequence can be divided into slices 3. Slices may, in some instances, constitute the entire image. These slices are divided into non-overlapping coding tree units (CTUs). A code tree unit (CTU) is the basic processing unit of the High Efficiency Video Coding (HEVC) video standard and conceptually corresponds structurally to the macroblock unit used in some previous video standards. A CTU is sometimes referred to as a Largest Coding Unit (LCU). A CTU has a luma (luminance) component portion and a chroma (color difference) component portion, each of which is called a coding tree block (CTB). These different color components are not shown in FIG. 1.

[0036] A CTU is typically 64 pixels by 64 pixels in size. Each CTU may be iteratively divided into smaller, variable-sized coding units (CUs) 5 using quadtree decomposition.

[0037] A coding unit is a basic coding element and consists of two types of subunits called prediction units (PUs) and transform units (TUs). The maximum size of a PU or TU is equal to the size of a CU. A prediction unit corresponds to the partitioning of a CU for the prediction of pixel values. Various different partitions of a CU into PUs are possible, as shown by 606, including a partition into four square PUs and two different partitions into two rectangular PUs. A transform unit is a basic unit that performs spatial transformation using DCT. A CU can be partitioned into TUs based on a quadtree representation 607.

[0038] Each slice is embedded in one Network Abstraction Layer (NAL) unit. Furthermore, the coding parameters of a video sequence are stored in a dedicated NAL unit called a parameter set. HEVC and H.264 / AVC use two types of parameter set NAL units. The first is the sequence parameter set (SPS) NAL unit, which collects all parameters that do not change during the entire video sequence. Typically, it handles coding profiles, video frame sizes, and other parameters. The second, the picture parameter set (PPS) NAL unit, contains parameters that may change from one picture (or frame) of the sequence to another. HEVC also includes the video parameter set (VPS) NAL unit, which contains parameters that describe the overall structure of the bitstream. The VPS is a new type of parameter set defined in HEVC that applies to all layers of the bitstream. A layer can contain multiple temporal sublayers; all Version 1 bitstreams are limited to one layer. HEVC has specific layer extensions for scalability and multiview, which allow for multiple layers with a backward-compatible version 1 base layer.

[0039] In the current definition of Versatile Video Coding (VVC), there are three high-level possibilities for picture partitioning: subpictures, slices, and tiles. Each has its own characteristics and usefulness. Partitioning into subpictures is for spatial extraction and / or merging of regions of video. Partitioning into slices is based on a similar concept to previous standards and corresponds to packetization for video transmission, even though it can be used for other applications. Partitioning into tiles is conceptually a code parallelization tool that divides a picture into independent coding regions of (approximately) the same size as the picture, although this tool can also be used for other applications.

[0040] These three available high-level methods of picture partitioning can be used together, so there are several modes for their use. As defined in the current draft specification of VVC, two modes of slicing are defined. In the raster scan slice mode, a slice contains a series of complete tiles in the tile raster scan of the picture. This mode in the current VVC specification is illustrated in Figure 10(a). As shown in this figure, a picture is shown to contain an 18x12 luma CTU partitioned into 12 tiles and 3 raster scan slices.

[0041] In the second slice mode, rectangular slice mode, a slice collectively contains any number of complete tiles from a rectangular region of the picture. This mode in the current VVC specification is illustrated in Figure 10(b). This example shows a picture with 18x12 luma CTUs partitioned into 24 tiles and 9 rectangular slices.

[0042] 2 illustrates a data communication system in which one or more embodiments of the present invention may be implemented. The data communication system includes a transmitting device, in this case a server 201, operable to transmit data packets of a data stream to a receiving device, in this case a client terminal 202, via a data communication network 200. The data communication network 200 may be a wide area network (WAN) or a local area network (LAN). Such a network may be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a hybrid network made up of multiple different networks. In certain embodiments of the present invention, the data communication system may be a digital television broadcasting system in which a server 201 transmits the same data content to multiple clients.

[0043] The data stream 204 provided by the server 201 may be comprised of multimedia data representing video and audio data. The audio and video data streams may be captured by the server 201 using a microphone and a camera, respectively, in some embodiments of the present invention. In some embodiments, the data streams may be stored on the server 201, received by the server 201 from another data provider, or generated at the server 201. The server 201 is particularly provided with an encoder for encoding the video and audio streams to provide a compressed bitstream for transmission with a more compact representation of the data presented as input to the encoder.

[0044] In order to obtain a better ratio between the quality of the transmitted data and the amount of transmitted data, the compression of the video data is for example according to the HEVC or H.264 / AVC format.

[0045] Client 202 receives the transmitted bitstream, decodes the reconstructed bitstream, and plays the video images on a display device and the audio data through a speaker.

[0046] Although the example of Figure 2 considers a streaming scenario, it will be appreciated that in some embodiments of the present invention data communication between the encoder and decoder may be performed using a media storage device, such as an optical disc, for example.

[0047] In one or more embodiments of the present invention, a video image is transmitted along with data representing a compensation offset to be applied to the reconstructed pixels of the image to provide filtered pixels in the final image.

[0048] 3 shows a schematic diagram of a processing device 300 configured to implement at least one embodiment of the present invention. The processing device 300 may be a device such as a microcomputer, a workstation, or a lightweight handheld device. The device 300 comprises a communication bus 313 to which the following are connected: - a central processing unit 311 such as a microprocessor, denoted CPU; - a read-only memory 306, denoted ROM, for storing a computer program for implementing the invention; a random access memory 312, denoted RAM, for storing the executable code of the methods of the embodiments of the invention, as well as registers adapted to record variables and parameters necessary for implementing the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to an embodiment of the invention; a communication interface 302 connected to a communication network 303 over which the digital data to be processed is transmitted and received;

[0049] Optionally, the device 300 may also include the following components: - data storage means 304, such as a hard disk, for storing computer programs for implementing the methods of one or more embodiments of the present invention and data used or generated during the implementation of one or more embodiments of the present invention; - a disk drive 305 for a disk 306, the disk drive being configured to read data from or write data to the disk 306; a screen 309 for displaying data that serves as a graphical interface with the user by means of a keyboard 310 or any other pointing means.

[0050] The device 300 may be connected to a variety of peripheral devices, such as a digital camera 320 or a microphone 308, each connected to an input / output card (not shown) to provide multimedia data to the device 300.

[0051] The communication bus provides communication and interoperability between the various elements included in or connected to the device 300. The representation of the bus is not limiting, and in particular the central processing unit is operable to communicate instructions to any element of the device 300 directly or by means of another element of the device 300.

[0052] The disk 306 may be replaced by any information carrier, such as for example a compact disk (CD-ROM), a ZIP disk or a memory card, rewritable or not, and generally speaking by any information storage means readable by a microcomputer or microprocessor, integrated or not integrated into the device, possibly removable, and configured to store one or more programs whose execution enables the method of encoding a sequence of digital images and / or the method of decoding a bitstream according to the invention.

[0053] The executable code may be stored either in the read-only memory 306, on the hard disk 304 or on a removable digital medium as previously explained, such as for example the disk 306. According to a variant, the executable code of the program may be received by means of the communication network 303, via the interface 302, to be stored in one of the storage means of the device 300 before being executed, such as the hard disk 304.

[0054] The central processing unit 311 is adapted to control and direct the execution of the program or instructions of the program or parts of the software code according to the invention with instructions stored in one of the storage means mentioned above. On power-up, the program or programs stored in a non-volatile memory, for example on the hard disk 304 or on the read-only memory 306, are transferred to the random access memory 312, which contains registers for storing the program or the executable code of the program, as well as variables and parameters necessary to carry out the invention.

[0055] In this embodiment, the device is a programmable device that uses software to implement the invention, although the invention may also be implemented in hardware (e.g., in the form of an application specific integrated circuit or ASIC).

[0056] 4 shows a block diagram of an encoder according to at least one embodiment of the present invention, the encoder being represented by connected modules, each module adapted to be implemented, for example, in the form of program instructions executed by the CPU 311 of the device 300, and at least one corresponding step of a method for encoding images of a sequence of images according to one or more embodiments of the present invention.

[0057] An original sequence 401 of digital images i0 to in is received as input by an encoder 400. Each digital image is represented by a set of samples known as pixels.

[0058] A bitstream 410 is output by the encoder 400 after the encoding process has been performed. The bitstream 410 comprises multiple coding units or slices, each comprising a slice header for transmitting coded values ​​of coding parameters used to code the slice, and a slice body containing coded video data.

[0059] An input digital image i0~in 401 is divided into blocks of pixels by module 402. The blocks correspond to portions of the image and can be of variable size (for example, 4x4, 8x8, 16x16, 32x32, 64x64, 128x128 pixels, and several rectangular block sizes can also be considered). For each input block, a coding mode is selected. Two families of coding modes are provided: coding modes based on spatial predictive coding (intra prediction) and coding modes based on temporal prediction (inter coding, merge, SKIP). Possible coding modes are tested.

[0060] Module 403 implements an intra prediction process, whereby a given block to be coded is predicted by a predictor calculated from pixels neighboring said block to be coded. If intra prediction is selected, an indication of the selected intra predictor and the difference between the given block and its predictor is coded to provide a residual.

[0061] Temporal prediction is performed by a motion estimation module 404 and a motion compensation module 405. First, a reference image is selected from a set of reference images 416, and the portion of the reference image, also called a reference area or image portion, that is closest to the block to be coded is selected by the motion estimation module 404. Then, the motion compensation module 405 uses the selected area to predict the block to be coded. The difference between the selected reference area and a given block, also called a residual block, is calculated by the motion compensation module 405. The selected reference area is indicated by a motion vector.

[0062] Therefore, in both cases (spatial and temporal prediction), the residual is calculated by subtracting the prediction from the original block.

[0063] In the intra prediction performed by module 403, the prediction direction is coded. In the temporal prediction, at least one motion vector is coded. In the inter prediction performed by modules 404, 405, 416, 418, 417, at least one motion vector or data for identifying such a motion vector is coded for the temporal prediction.

[0064] If inter prediction is selected, information about the motion vector and the residual block are coded. To further reduce the bit rate, assuming uniform motion, the motion vector is coded by a differential to a motion vector predictor. The motion vector predictor of the set of motion information predictors is obtained from the motion vector field 418 by the motion vector prediction and coding module 417.

[0065] The encoder 400 further comprises a selection module 406 for selecting a coding mode by applying a coding cost criterion, such as a rate-distortion criterion. To further reduce redundancy, a transform (such as a DCT) is applied to the residual block by a transform module 407, and the resulting transformed data is quantized by a quantization module 408 and entropy coded by an entropy coding module 409. Finally, the coded residual block of the current block being coded is inserted into a bitstream 410.

[0066] The encoder 400 also decodes the coded images to generate reference images for motion estimation of subsequent images, allowing the encoder and the decoder receiving the bitstream to have the same reference frame. An inverse quantization module 411 performs inverse quantization of the quantized data, followed by inverse transformation by an inverse transform module 412. An inverse intra prediction module 413 uses the prediction information to decide which predictor to use for a given block, and an inverse motion compensation module 414 actually adds the residual obtained by module 412 to a reference area obtained from a set of reference images 416.

[0067] Post-filtering by module 415 is then applied to filter the frame of reconstructed pixels. In an embodiment of the present invention, an SAO loop filter is used to add a compensation offset to the pixel values ​​of the reconstructed pixels of the reconstructed image.

[0068] 5 shows a block diagram of a decoder 60 that may be used to receive data from an encoder, according to one embodiment of the present invention. The decoder is represented by connected modules, each configured to implement a corresponding step of a method implemented by the decoder 60, e.g., in the form of program instructions executed by the CPU 311 of the device 300.

[0069] Decoder 60 receives a bitstream 61 containing coding units. Each coding unit consists of a header containing information about coding parameters and a body containing the coded video data. The structure of a bitstream in VVC is explained in more detail below with reference to Figure 6. As explained with reference to Figure 4, the coded video data is entropy coded, and the index of the motion vector predictor is coded with a predetermined number of bits for a given block. The received coded video data is entropy decoded by module 62. The residual data is then inverse quantized by module 63, and then an inverse transform is applied by module 64 to obtain pixel values.

[0070] Mode data indicating the coding mode is also entropy decoded, and based on the mode, INTRA-type or INTER-type decoding is performed on the coded block of image data.

[0071] For INTRA mode, the INTRA predictor is determined by the intra inverse prediction module 65 based on the intra prediction mode specified in the bitstream.

[0072] If the mode is INTER, motion prediction information is extracted from the bitstream to find the reference region used by the encoder. The motion prediction information consists of a reference frame index and a motion vector residual. The motion vector predictor is added to the motion vector residual by the motion vector decoding module 70 to obtain the motion vector.

[0073] A motion vector decoding module 70 applies motion vector decoding for each current block coded by motion prediction. Once the motion vector predictor index is obtained, for the current block, the actual value of the motion vector associated with the current block can be decoded and used to apply inverse motion compensation by module 66. The reference image portion indicated by the decoded motion vector is extracted from reference image 68, and inverse motion compensation 66 is applied. Motion vector field data 71 is updated with the decoded motion vector for use in inverse prediction of subsequent decoded motion vectors.

[0074] Finally, the decoded blocks are obtained. Post-filtering is applied by a post-filtering module 67. A decoded video signal 69 is finally provided by the decoder 60.

[0075] FIG. 6 shows the structure of a bitstream in an exemplary coding system VVC as described in JVET-Q2001-vD.

[0076] A bitstream 61 from a VVC coding system consists of an ordered sequence of syntax elements and coded data. The syntax elements and coded data are arranged in Network Abstraction Layer (NAL) units 601-608. There are different types of NAL units. The Network Abstraction Layer provides the ability to encapsulate the bitstream in different protocols such as RTP / IP, Real Time Protocol / Internet Protocol, ISO Base Media File Format, etc. The Network Abstraction Layer also provides a framework for packet loss resilience.

[0077] NAL units are divided into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain the actual coded video data. Non-VCL NAL units contain additional information, either parameters required for decoding the coded video data or supplementary data that can improve the usability of the decoded video data. NAL units 606 correspond to slices and constitute the VCL NAL units of the bitstream.

[0078] Different NAL units 601-605 correspond to different parameter sets, and these NAL units are non-VCL NAL units. The Decoder Parameter Set (DPS) NAL unit 301 contains parameters that are constant for a given decoding process. The Video Parameter Set (VPS) NAL unit 602 contains parameters defined for the entire video, i.e., the entire bitstream. The DPS NAL unit can define parameters that are more static than those in the VPS. In other words, the parameters in the DPS change less frequently than the parameters in the VPS.

[0079] The sequence parameter set (SPS) NAL unit 603 contains parameters defined for a video sequence. In particular, the SPS NAL unit can define the subpicture layout and associated parameters of a video sequence. The parameters associated with each subpicture specify the coding constraints that apply to the subpicture. In particular, it comprises a flag that indicates that temporal prediction between subpictures is limited to data coming from the same subpicture. Another flag can enable or disable a loop filter that crosses subpicture boundaries.

[0080] The Picture Parameter Set (PPS) NAL unit 604 contains parameters defined for a picture or group of pictures. The Adaptation Parameter Set (APS) NAL unit 605 contains parameters for the loop filter, typically an Adaptation Loop Filter (ALF) or reshaper model (or Luma mapping with chroma scaling (LMCS) model) or scaling matrix used at the slice level.

[0081] The PPS syntax proposed in the current version of VVC provides syntax elements that specify the size of the picture in luma samples and the partitioning of each picture in tiles and slices.

[0082] The PPS contains syntax elements that allow determining the location of slices within a frame. Since subpictures form rectangular regions within a frame, it is possible to determine the set of slices, parts of tiles, or tiles that belong to a subpicture from the parameter set NAL unit. Like the APS, the PPS has an ID mechanism that limits the amount of transmission of the same PPS.

[0083] The main difference between the PPS and the Picture Header is that the PPS is transmitted, and typically is transmitted for a group of pictures, compared to the PH, which is transmitted systematically for each picture. Therefore, the PPS, as compared to the PH, contains parameters that may be constant for several pictures.

[0084] The bitstream may also contain Supplemental Enhancement Information (SEI) NAL units (not shown in FIG. 6). Within the bitstream, the periodicity with which these parameter sets occur is variable. A VPS defined for the entire bitstream may occur only once in the bitstream. In contrast, an APS defined for a slice may occur once for each slice in each picture. In fact, different slices may rely on the same APS; thus, in general, there are fewer APSs than slices in each picture. In particular, APSs are defined in the picture header. However, ALF APSs can be refined in the slice header.

[0085] An Access Unit Delimiter (AUD) NAL unit 607 separates two access units. An access unit is a set of NAL units that may comprise one or more coded pictures with the same decoding timestamp. This optional NAL unit contains only one syntax element, pic_type, in the current VVC specification. This syntax element, pic_type, indicates the slice_type values ​​of all slices of coded pictures in the AU. If pic_type is set equal to 0, the AU contains only intra slices. If it is set equal to 1, it contains P and I slices. If it is set equal to 2, it contains B, P, or intra slices. This NAL unit contains only one syntax element of type pic-type.

[0086] Table 1 Syntax AUD JPEG0007733777000001.jpg32126

[0087] In JVET-Q2001-vD, pic_type is defined as follows: "pic_type indicates that the slice_type values ​​of all slices of the coded picture within the AU containing the AU delimiter NAL unit are members of the set listed in Table 2 for the given value of pic_type. The value of pic_type shall be equal to 0, 1, or 2 in bitstreams conforming to this version of this specification. Other values ​​of pic_type are reserved for future use by ITUT-T |ISO / IEC. Decoders conforming to this version of this specification shall ignore reserved values ​​of pic_type." rbsp_trailing_bits() is a function that adds bits to align the end of a byte, so that after this function the amount of parsed bitstream is an integer number of bytes.

[0088] Table 2 Interpretation of pic_type JPEG0007733777000002.jpg31105

[0089] The PH NAL unit 608 is a picture header NAL unit that groups parameters common to a set of slices of one coded picture. A picture may reference one or more APSs to indicate the AFL parameters, reconstruction models, and scaling matrices used by the slices of the picture.

[0090] Each VCL NAL unit 606 contains a slice. A slice may correspond to an entire picture or subpicture, a single tile, multiple tiles, or a portion of a tile. For example, the slice in Figure 3 contains several tiles 620. A slice consists of a slice header 610 and a raw byte sequence payload (RBSP) 611 that contains coded pixel data encoded as coding blocks 640.

[0091] The PPS syntax proposed in the current version of VVC provides syntax elements that specify the size of the picture in luma samples and the partitioning of each picture in tiles and slices.

[0092] The PPS contains syntax elements that allow determining the location of slices within a frame. Since a subpicture forms a rectangular region within a frame, it is possible to determine from the parameter set NAL unit the set of slices, parts of tiles, or tiles that belong to a subpicture.

[0093] NAL Unit Slice The NAL unit slice layer includes a slice header and slice data as shown in Table 3.

[0094] Table 3 Slice layer syntax JPEG0007733777000003.jpg38125

[0095] APS The adaptation parameter set (APS) NAL unit 605 is defined in Table 4, which shows the syntax elements. As shown in Table 4, there are three types of APS that can be specified with the APS_params_type syntax element: ALF_AP: for ALF parameters LMCS_APS for LMCS parameters Scaling list Scaling_APS for relative parameters

[0096] Table 4. Syntax of adaptive parameter sets JPEG0007733777000004.jpg98125

[0097] These three types of APS parameters are discussed in turn below.

[0098] ALP AP The ALF parameters are described in the Adaptive Loop Filter Data syntax element (Table 5). First, four flags are dedicated to specify whether ALF filters are transmitted for luma and / or chroma, and whether CC-ALF (Cross Component Adaptive Loop Filtering) is enabled for the Cb and Cr components. If the luma filter flag is enabled, another flag is decoded to determine whether a clip value is signaled (alf_Luma_clip_flag). Next, the number of signaled filters is decoded using the alf_luma_num_filters_signalled_minus1 syntax element. If necessary, a syntax element representing the ALF coefficient delta "alf_luma_coeff_delta_idx" is decoded for each enabled filter. Then, the absolute value and sign of each coefficient of each filter are decoded.

[0099] If alf_luma_clip_flag is enabled, the clip index for each coefficient of each enabled filter is decoded.

[0100] Similarly, the ALF chroma coefficients are decoded as needed.

[0101] If CC-ALF is enabled for Cr or Cb, the number of filters is decoded (alf_cc_cb_filters_signalled_minus1 or alf_cc_cr_filters_signalled_minus1) and the associated coefficients are decoded (alf_cc_cb_mapped_coeff_abs and alf_cc_cb_coeff_sign or alf_cc_cr_mapped_coeff_abs and alf_cc_cr_coeff_sign respectively).

[0102] Table 5. Adaptive Loop Filter Data Syntax JPEG0007733777000005.jpg155126JPEG0007733777000006.jpg92124

[0103] LMCS syntax elements for both luma mapping and chroma scaling Table 6 below lists all LMCS syntax elements that are coded in the adaptation parameter set (APS) syntax structure when the APS_params_type parameter is set to 1 (LMCS_APS). Up to four LMCS APSs can be used in a coded video sequence, but only a single LMCS APS can be used for a given picture. These parameters are used to construct the forward and backward mapping functions for luma and the scaling function for chroma.

[0104] Table 6. Luma Mapping with Chroma Scaling Data Syntax JPEG0007733777000007.jpg71125

[0105] Scaling List APS The scaling list provides the possibility to update the quantization matrix used for quantization. In VVC, this scaling matrix is ​​signaled in the APS as described in the Scaling List Data syntax element (Table 7 Scaling List Data Syntax). The first syntax element specifies whether a scaling matrix is ​​used for the LFNST (Low Frequency Non-Separable Transform) tool based on the flag scaling_matrix_for_LFNST_disabled_flag. The second one specifies if a scaling list is used for the chroma components (scaling_list_Chroma_present_flag). Next, the syntax elements required to construct the scaling matrix are decoded (scaling_list_copy_mode_flag, scaling_list_pred_mode_flag, scaling_list_pred_id_delta, scaling_list_dc_coef, scaling_list_delta_coef).

[0106] Table 7. Scaling List Data Syntax JPEG0007733777000008.jpg140127

[0107] Picture Header The picture header is transmitted at the beginning of each picture before any other slice data. It is significantly larger than the previous header in previous drafts of the standard. A complete description of all these parameters can be found in JVET-Q2001-vD. Table 9 shows these parameters in the current picture header decoding syntax.

[0108] The relevant syntax elements that can be decoded are: Use of this picture, reference frame or not Picture type Output Frame Number of pictures Use subpictures where appropriate Reference picture list if needed Color planes as needed · Partition update when override flag is enabled Delta QP parameters if necessary Movement information parameters as needed ALF parameters as needed SAO parameters as needed Quantification parameters as needed LMCS parameters as needed Scaling list parameters as needed Picture header extension if needed Etc...

[0109] Picture "Type" The first flag is gdr_or_irap_pic_flag, which indicates whether the current picture is a resynchronisation picture (IRAP or GDR). If this flag is true, gdr_pic_flag is decoded to tell whether the current picture is an IRAP or GDR picture. Next, the ph_Inter_slice_allowed_flag is decoded to identify that inter-slices are allowed. If allowed, the flag ph_intra_slice_allowed_flag is decoded to tell whether intra slices are allowed for the current picture. Next, the non_reference_picture_flag, ph_pic_parameter_set_ID indicating the PPS ID, and the picture order count ph_pic_order_cnt_lsb are decoded. The picture order count gives the number of the current picture. If the picture is a GDR or IRAP picture, the flag no_output_of_prior_pics_flag is decoded. If the picture is GDR, recovery_poc_cnt is decoded, followed by ph_poc_msb_present_flag and poc_msb_val, if necessary.

[0110] ALF After these parameters describing important information about the current picture, a set of ALF APS id syntax elements is decoded if ALF is enabled at the SPS level and if ALF is enabled at the picture header level. ALF is enabled at the SPS level by the sps_alf_enabled_flag flag, and ALF signaling is enabled at the picture header level by alf_info_in_ph_flag being 1. Otherwise (alf_info_in_ph_flag is 0), ALF is signaled at the slice level.

[0111] alf_info_in_ph_flag is defined as follows: "alf_info_in_ph_flag equal to 1 specifies that ALF information is present in PH syntax structures and is not present in slice headers that reference PPSs that do not contain PH syntax structures. alf_info_in_ph_flag equal to 0 specifies that ALF information is not present in PH syntax structures and may be present in slice headers that reference PPSs that do not contain PH syntax statement structures."

[0112] First, decode ph_alf_enabled_present_flag and determine whether to decode ph_alf_enabled_flag. If ph_alf_enabled_flag is enabled, ALF is enabled for all slices of the current image.

[0113] If ALF is enabled, the amount of ALF APS ids for luma is decoded using the pic_num_alf_aps_ids_luma syntax element. For each APS id, the APS id value for luma is decoded to “ph_alf_APS_id_luma”.

[0114] For chrominance syntax elements, ph_alf_chroma_idc is decoded to determine if ALF is enabled for chroma, either for Cr only or for Cb only. If enabled, the APS ID value for chroma is decoded using the ph_alf_aps_ID_Chroma syntax element. In this way, the APS ID for the CC-ALF method is decoded as needed for the Cb and / or CR components.

[0115] LMCS Then, a set of LMCS APS ID syntax elements are decoded if LMCS is enabled at the SPS level. First, ph_lmcs_enabled_flag is decoded to determine if LMCS is enabled for the current picture. If LMCS is enabled, its ID value is decoded ph_lmcs_aps_ID. For chroma, only ph_Chroma_residual_scale_flag is decoded to enable or disable the chroma method.

[0116] Scaling List Then, if the scaling list is enabled at the SPS level, a set of scaling list APS IDs is decoded. ph_scaling_list_present_flag is decoded to determine whether the scaling matrix is ​​enabled for the current picture. Then, the APS ID value ph_scaling_list_aps_id is decoded.

[0117] Subpicture Subpicture parameters are enabled when enabled in the SPS and when subpicture id signaling is disabled. They also contain information about the virtual border. Eight syntax elements are defined for subpicture parameters: · ph_virtual_boundaries_present_flag · ph_num_ver_virtual_boundaries · ph_virtual_boundaries_pos_x[i] · ph_num_hor_virtual_boundaries · ph_virtual_boundaries_pos_y[i]

[0118] Output Flags These subpicture parameters are followed by the pic_output_flag, if present.

[0119] Reference Picture List If reference picture lists are signaled in the picture header (because rpl_info_in_ph_flag is 1), the reference picture list parameters are decoded in ref_pic_lists(), which contains the following syntax elements: rpl_sps_flag[] rpl_idx[] poc_lsb_lt[][] · delta_poc_msb_present_flag[][] delta_poc_msb_cycle_lt[][]

[0120] Partitioning A set of partitioning parameters is decoded as needed, which includes the following syntax elements: · partition_constraints_override_flag · ph_log2_diff_min_qt_min_cb_intra_slice_luma · ph_max_mtt_hierarchy_depth_intra_slice_luma · ph_log2_diff_max_bt_min_qt_intra_slice_luma · ph_log2_diff_max_tt_min_qt_intra_slice_luma · ph_log2_diff_min_qt_min_cb_intra_slice_chroma · ph_max_mtt_hierarchy_depth_intra_slice_chroma · ph_log2_diff_max_bt_min_qt_intra_slice_chroma · ph_log2_diff_max_tt_min_qt_intra_slice_chroma · ph_log2_diff_min_qt_min_cb_inter_slice · ph_max_mtt_hierarchy_depth_inter_slice · ph_log2_diff_max_bt_min_qt_inter_slice · ph_log2_diff_max_tt_min_qt_inter_slice

[0121] Weighted prediction The weighted prediction parameters pred_weight_table() are decoded if the weighted prediction method is enabled at the PPS level and the weighted prediction parameters are signaled in the picture header (wp_info_in_ph_flag is 1). pred_weight_table() contains the weighted prediction parameters for list L0 and list L1 if bi-prediction weighted prediction is enabled. If weighted prediction parameters are sent in the picture header, the weight number for each list is explicitly sent, as shown in pred_weight_table() syntax table 8.

[0122] Table 8. Weighted Prediction Parameter Syntax JPEG0007733777000009.jpg173124

[0123] Delta QP If the picture is intra, ph_cu_qp_delta_subdiv_intra_slice and ph_cu_chroma_qp_offset_subdiv_intra_slice are decoded as needed. Also, if inter slicing is allowed, ph_cu_qp_delta_subdiv_inter_slice and ph_cu_chroma_qp_offset_subdiv_inter_slice are decoded as needed. Finally, picture header extension syntax elements are decoded as needed.

[0124] All parameters alf_info_in_ph_flag, rpl_info_in_ph_flag, qp_delta_info_in_ph_flag, sao_info_in_ph_flag, dbf_info_in_ph_flag, wp_info_in_ph_flag are signaled in the PPS.

[0125] Table 9 Picture Header Structure JPEG0007733777000010.jpg186125JPEG0007733777000011.jpg186124JPEG0007733777000012.jpg187124JPEG0007733777000013.jpg100124

[0126] Slice Header The slice header is transmitted at the beginning of each slice. The slice header contains approximately 65 syntax elements, which is significantly larger than previous slice headers in previous video coding standards. A complete description of all slice header parameters can be found in JVET-Q2001-vD. Table 10 shows these parameters in the current slice header decoding syntax statement.

[0127] Table 10 Partial Slice Header JPEG0007733777000014.jpg186125JPEG0007733777000015.jpg184124JPEG0007733777000016.jpg71123

[0128] First, picture_header_in_slice_header_flag is decoded to know if picture_header_structure() is present in the slice header. Then, slice_subpic_id is decoded, if necessary, to determine the subpicture id of the current slice. Next, slice_address is decoded to determine the address of the current slice. The slice address is decoded if the current slice mode is rectangular slice mode (rect_slice_flag is equal to 1) and the number of slices in the current subpicture is greater than 1. The slice address can also be decoded if the current slice mode is raster scan mode (rect_slice_flag is 0) and the number of tiles in the current picture is greater than 1, calculated based on variables defined in the PPS.

[0129] Next, num_tiles_in_slice_minus1 is decoded if the number of tiles in the current picture is greater than 1 and the current slice mode is not a rectangular slice mode. In the current VVC draft specification, num_tiles_In_slice_minus1 is defined as follows: "num_tiles_in_slice_minus1 + 1 (if present) specifies the number of tiles in the slice. The value of num_tiles_in_slice_minus1 ranges from 0 to NumTilesInPic-1."

[0130] Then the slice_type is decoded. If ALF is enabled at the SPS level (sps_alf_enabled_flag) and ALF is signaled in the slice header (alf_info_in_ph_flag is 0), the ALF information is decoded. This includes a flag (slice_alf_enabled_flag) indicating that ALF is enabled for the current slice. If enabled, the number of ALF IDs for luma (slice_num_alf_aps_ids_luma) is decoded, and the APS ID is decoded (slice_alf_aps_id_luma[i]). Next, slice_ALF_chroma_idc is decoded to know if ALF is enabled for a chroma component and for which chroma component. Then, the APS ID for chroma is decoded to slice_alf_aps_id_chroma, if necessary. Similarly, slice_cc_alf_cb_enabled_flag is decoded as needed to know whether the CC ALF method is enabled. If CC ALF is enabled, the associated APS IDs for CR and / or CB are decoded if CC ALF is enabled for CR and / or CB.

[0131] If the color planes are sent separately (separate_colour_plane_flag equals 1), colour_plane_id is decoded.

[0132] If the reference picture list is not transmitted in the picture header (rpl_info_in_ph_flag is equal to 0) and the Nal unit is not IDR, or if the reference picture list is transmitted for an IDR picture (sps_idr_rpl_present_flag is equal to 1), the reference picture list parameters are decoded. They are the same as the parameters in the picture header.

[0133] The override flag num_ref_idx_active_override_flag is decoded if reference picture lists are transmitted in the picture header (rpl_info_in_ph_flag is 1) and the Nal unit is not IDR, or if reference picture lists are transmitted for an IDR picture (sps_idr_rpl_present_flag is equal to 1) and the number of references in at least one list exceeds 1. If this flag is enabled, the reference index of each list is decoded.

[0134] cabac_init_flag is decoded when the slice type is not intra and as needed. slice_collocated_from_l0_flag and slice_collocated_ref_idx are decoded if a reference picture list is transmitted in the slice header and other conditions apply. These data are related to CABAC coding and located motion vectors.

[0135] Similarly, if the slice type is not intra, the weighted prediction parameters pred_weight_table() are decoded.

[0136] If delta QP information is sent in the slice header (QP_delta_info_in_ph_flag equals 0), slice_qp_delta is decoded. If necessary, the syntax elements slice_cb_qp_offset, slice_cr_qp_offset, slice_joint_cbcr_qp_offset, and cu_chroma_qp_offset_enabled_flag are decoded.

[0137] If SAO information is sent in the slice header (sao_info_in_ph_flag equals 0) and enabled at the SPS level (sps_sao_enabled_flag), the SAO enabled flags are decoded for both luma and chroma: slice_sao_luma_flag, slice_sao_chroma_flag. Then, the deblocking filter parameters are decoded if signaled in the slice header (dbf_info_in_ph_flag equals 0).

[0138] The flag slice_ts_residual_coding_disabled_flag is systematically decoded to know whether the Transform Skip residual coding method is enabled for the current slice. If LMCS was enabled in the picture header (ph_lmcs_enabled_flag equals 1), the flag slice_lmcs_enabled_flag is decoded. Similarly, if scaling lists were enabled in the picture header (phpic_scaling_list_presentenabled_flag equals 1), the flag slice_scaling_list_present_flag is decoded. Other parameters are then decoded as needed.

[0139] Picture Header in Slice Header In a specific signaling method, a picture header (708) can be signaled within a slice header (710), as shown in Figure 7. In that case, there is no NAL unit that contains only a picture header (608). NAL units 701-707 correspond to NAL units 601-607, respectively, in Figure 6. Similarly, a coding tile 720 and a coding block 740 correspond to blocks 620 and 640 in Figure 6. Therefore, the description of these units and blocks will not be repeated here. This can be enabled in the slice header by the flag picture_header_in_slice_header_flag. Furthermore, when a picture header is signaled within a slice header, a picture contains only one slice. Therefore, there is always only one picture header per picture. Furthermore, the flag picture_header_in_slice_header_flag has the same value for all pictures in a CLVS (Coded Layer Video Sequence). This means that all pictures between two IRAPs, including the first IRAP, have only one slice per picture.

[0140] The flag picture_header_in_slice_header_flag is defined as follows: If picture_header_in_slice_header_flag is 1, it specifies that the PH syntax structure is present in the slice header. If picture_header_in_slice_header_flag is 0, it specifies that the PH syntax structure is not present in the slice header. It is a bitstream conformance requirement that the value of picture_header_in_slice_header_flag be the same for all coded slices in CLVS. If picture_header_in_slice_header_flag is equal to 1 for a coded slice, it is a bitstream conformance requirement that no VCL NAL units with nal_unit_type equal to PH_NUT exist in the CLVS. When picture_header_in_slice_header_flag is equal to 0, all of the coded slices in the current picture have picture_header_in_slice_header_flag equal to 0, and the current PU has a PH NAL unit. The picture_header_structure() contains the syntax elements of picture_rbsp() except for the stuffing bits rbsp_trailing_bits().

[0141] Streaming Applications Some streaming applications extract only specific portions of the bitstream. These extractions can be spatial (as subpictures) or temporal (subportions of a video sequence). These extracted portions can then be merged with other bitstreams. Some others reduce the frame rate by extracting only a few frames. In general, the main goal of these streaming applications is to maximize the use of the available bandwidth in order to present the maximum quality to the end user.

[0142] In VVC, APS ID numbering is limited to reduce the frame rate so that a new APS id number of a frame cannot be used for frames at higher levels in the temporal hierarchy. However, for streaming applications that extract portions of the bitstream, it is necessary to keep track of the APS IDs to determine which APSs should be retained for a sub-portion of the bitstream, since frames (IRAPs) do not reset the APS ID numbering.

[0143] LMCS (Luma Mapping with Chroma Scaling) The Luma Mapping with Chroma scaling (LMCS) technique is a sample value transformation method applied to blocks before applying a loop filter in a video decoder such as VVC.

[0144] LMCS can be divided into two sub-tools: the first one is applied on the luma block, while the second one is applied on the chroma block, as follows: 1) The first sub-tool is in-loop mapping for the luma component based on adaptive piecewise linear models. In-loop mapping for the luma component adjusts the dynamic range of the input signal by redistributing codewords across the dynamic range to improve compression efficiency. Luma mapping utilizes a forward mapping function to the "mapped domain" and a corresponding backward mapping function back to the "input domain." 2) The second sub-tool relates to the chroma components to which luma-dependent chroma residual scaling is applied. The chroma residual scaling is designed to correct the interaction between the luma signal and the corresponding chroma signal. The chroma residual scaling depends on the average value of the top and / or left reconstructed neighboring luma samples of the current block.

[0145] Like most other tools in video coders such as VVC, LMCS can be enabled / disabled at the sequence level using the SPS flag. Whether chroma residual scaling is enabled is signaled at the slice level. When luma mapping is enabled, an additional flag is signaled to indicate whether luma-dependent chroma residual scaling is enabled. When luma mapping is not used, luma-dependent chroma residual scaling is completely disabled. Additionally, luma-dependent chroma residual scaling is always disabled for chroma blocks whose size is 4 or less.

[0146] FIG. 8 illustrates the principles of LMCS, as described above for the luma mapping subtool. The shaded blocks in FIG. 8 are new LMCS functional blocks that include forward and backward mapping of the luma signal. When using LMCS, it is important to note that several decoding operations are applied in the "mapped domain." These operations are represented by dashed blocks in this FIG. 8 and typically correspond to inverse quantization, inverse transform, luma intra prediction, and a reconstruction step that adds the luma prediction to the luma residual. Conversely, the solid blocks in FIG. 8 indicate where decoding processes are applied in the original (i.e., unmapped) domain, which includes deblocking, loop filtering such as ALF and SAO, motion compensation prediction, and storage of the decoded picture as a reference picture (DPB).

[0147] Figure 9 shows a diagram similar to Figure 8, but this time for the chroma scaling sub-tool of the LMCS tool. The shaded blocks in Figure 9 are new LMCS functional blocks that contain the luma-dependent chroma scaling process. However, there are some important differences for chroma compared to the luma case. Here, only the inverse quantization and inverse transform, represented by the dashed blocks, are performed in the "mapped domain" for chroma samples. The other steps of intra-chroma prediction, motion compensation, and loop filter are performed in the original domain. As shown in Figure 9, there is only the scaling process; there is no forward and backward processing as in the luma mapping case.

[0148] Luma mapping using piecewise linear models The luma mapping sub-tool uses a piecewise linear model, which separates the input signal dynamic range into 16 equal sub-ranges, where for each sub-range, its linear mapping parameters are expressed in terms of the number of codewords assigned to that range.

[0149] Luma mapping semantics The syntax element lmcs_min_bin_idx specifies the minimum bin index used in luma mapping with chroma scaling. The value of lmcs_min_bin_idx is in the range of 0 to 15.

[0150] The syntax element lmcs_delta_max_bin_idx specifies the delta value from 15 to the maximum bin index LmcsMaxBinIdx used in luma mapping with chroma scaling. The value of lmcs_delta_max_bin_idx is in the range of 0 to 15. The value of LmcsMaxBinIdx is set to 15 - lmcs_delta_max_bin_idx. The value of LmcsMaxBinIdx is greater than or equal to lmcs_min_bin_idx.

[0151] The syntax element lmcs_delta_cw_prec_minus1 +1 specifies the number of bits used to represent the syntax element lmcs_delta_abs_cw[i]. The syntax element lmcs_delta_abs_cw[i] specifies the absolute delta codeword value for the ith bin. The syntax element lmcs_delta_sign_cw_flag[i] specifies the sign of the variable lmcsDeltaCW[i]. If lmcs_delta_sign_cw_flag[i] is not present, it is inferred to be equal to 0.

[0152] LMCS intermediate variable operations for luma mapping To apply the forward and inverse luma mapping process, several intermediate variables and data arrays are required.

[0153] First, the variable OrgCW is derived as follows: OrgCW = (1 << BitDepth) / 16

[0154] Next, the variable lmcsDeltaCW[i] is calculated as follows, for i = lmcs_min_bin_idx .. LmcsMaxBinIdx: lmcsDeltaCW[i]=(1 - 2 * lmcs_delta_sign_cw_flag[i]) * lmcs_delta_abs_cw[i]

[0155] The new variable lmcsCW[i] is derived as follows: - For i = 0.. lmcs_min_bin_idx-1, set lmcsCW[i] equal to 0. - For i = lmcs_min_bin_idx.. LmcsMaxBinIdx, the following applies: lmcsCW[i]= OrgCW + lmcsDeltaCW[i] The value of lmcsCW[i] is in the range of (OrgCW>>3) to (OrgCW<<3-1). - For i = LmcsMaxBinIdx+1 .. 15, set lmcsCW[i] equal to 0.

[0156] Calculate the variable InputPivot[i] for i = 0..16 as follows: InputPivot[i] = i * OrgCW

[0157] The variable LmcsPivot[i] is calculated for i=0..16, and the variables ScaleCoeff[i] and InvScaleCoeff[i] are calculated for i=0..15 as follows: LmcsPivot[0]=0; for(i=0; i <=15; i++){ LmcsPivot[i+1] =LmcsPivot[i] + lmcsCW[i] ScaleCoeff[i] = (lmcsCW[i] * (1<<11) + (1 << (Log2(OrgCW)-1))) >> (Log2(OrgCW)) if (lmcsCW[i] ==0) invScaleCoeff[i] = 0 else invScaleCoeff[i] = OrgCW * (1 << 11) / lmcsCW[i]

[0158] Forward Luma Mapping As shown in FIG. 8, when LMCS is applied to luma, luma-remapped samples, called predMapSamples[i][j], are obtained from the predicted samples predSamples[i][j]. predMapSamples[i][j] is calculated as follows: First, an index idxY is calculated from the predicted sample predSamples[i][j] at position (i,j). idxY = preSamples[i][j] >> log2(OrgCW) Next, predMapSamples[i][j] is derived as follows using the intermediate variables idxY, LmcsPivot[idxY], and InputPivot[idxY] in section 0. predMapSamples[i][j] = LmcsPivot[idxY] + (ScaleCoeff[idxY] * (predSamples[i][j] - InputPivot[idxY])+(1 << 10)) >> 11

[0159] Luma reconstruction samples The reconstruction process is obtained from the predicted luma samples predMapSample[i][j] and the residual luma samples resiSamples[i][j]. The reconstructed luma picture samples recSamples[i][j] are obtained simply by adding predMapSample[i][j] to resiSamples[i][j] as follows. recSamples[i][j] = Clip1(predMapSamples[i][j] + resiSamples[i][j]]) In the above relationship, the Clip1 function is a clipping function that ensures that the reconstructed samples are properly within 0 and 1<<BitDepth-1.

[0160] Inverse luma mapping When applying inverse luma mapping according to FIG. 8, the following operation is applied to each sample recSample[i][j] of the current block to be processed: First, the index idxY is calculated from the reconstruction sample recSamples[i][j] at position (i,j). idxY = recSamples[i][j] >> Log2(OrgCW) The inverse-mapped luma sample invLumaSample[i][j] is derived as follows: invLumaSample[i][j] = InputPivot[idxYInv] + (InvScaleCoeff[idxYInv] * (recSample[i][j] - LmcsPivot[idxYInv]) + (1 << 10)) >> 11 Then a clipping operation is performed to obtain the final sample: finalSample[i][j] = Clip1(invLumaSample[i][j])

[0161] Chroma Scaling LMCS Semantics for Chroma Scaling The syntax element lmcs_delta_abs_crs in Table 6 specifies the absolute codeword value of the variable lmcsDeltaCrs. The value of lmcs_delta_abs_crs shall be in the range 0 to 7, inclusive. If not present, lmcs_delta_abs_crs is inferred to be equal to 0. The syntax element lmcs_delta_sign_crs_flag specifies the sign of the variable lmcsDeltaCrs. If not present, lmcs_delta_sign_crs_flag is inferred to be equal to 0.

[0162] LMCS intermediate variable calculation for chroma scaling To apply the chroma scaling process, several intermediate variables are required. The variable lmcsDeltaCrs is derived as follows: lmcsDeltaCrs = (1 - 2 * lmcs_delta_sign_crs_flag) * lmcs_delta_abs_crs The variable ChromaScaleCoeff[i], for i = 0..15, is derived as follows: if(lmcsCW[i] == 0) ChromaScaleCoeff[i] = (1 << 11) else ChromaScaleCoeff[i] = OrgCW * (1 << 11) / (lmcsCW[i] + lmcsDeltaCrs)

[0163] Chroma Scaling Process In the first step, a variable invAvgLuma is derived to calculate the average luma value of the reconstructed luma samples around the current corresponding chroma block. The average luma is calculated from the luma blocks to the left and above surrounding the corresponding chroma block. If there are no samples, the variable invAvgLuma is set as follows: invAvgLuma = 1 << (BitDepth - 1)

[0164] Then, based on the intermediate array LmcsPivot[ ] of section 0, the variable idxYInv is derived as follows: For (idxYInv = lmcs_min_bin_idx; idxYInv <= LmcsMaxBinIdx; idxYInv++) { if(invAvgLuma < LmcsPivot [idxYInv + 1]) break } IdxYInv = Min(idxYInv, 15)

[0165] The variable varScale is derived as follows: varScale = ChromaScaleCoeff[idxYInv]

[0166] When the transform is applied to the current chroma block, the reconstructed chroma picture sample array recSamples is derived as follows: recSamples[i][j] = Clip1(predSamples[i][j] + Sign(resiSamples[i][j]) * ((Abs(resiSamples[i][j]) * varScale + (1 << 10)) >> 11)) If no transformations apply to the current block, the following applies: recSamples[i][j]= Clip1(predSamples[i][j])

[0167] Encoder Considerations The basic principle of an LMCS encoder is to first assign more codewords to those dynamic range segments whose codewords have a lower average variance. Another formulation of this is to assign fewer codewords to dynamic range segments whose main goal is LMCS whose codewords have a higher average variance. In this way, smooth areas of the picture are coded with more codewords than average, and vice versa.

[0168] All parameters of the LMCS tool (see Table 6) stored in the APS are determined on the encoder side. The LMCS encoding encoder algorithm is based on the evaluation of the local luma variance and optimizes the determination of the LMCS parameters according to the basic principles described above. Optimization is then performed to obtain the best PSNR metric for the final reconstructed samples of a given block.

[0169] Embodiment Avoid slice address syntax elements when not needed In one embodiment, when a picture header is signaled within a slice header, the slice address syntax element (slice_address) is inferred to be equal to the value 0, even if the number of tiles is greater than 1. Table 11 illustrates this embodiment.

[0170] The advantage of this embodiment is that slice addresses are not parsed when the picture header is in the slice header which reduces the bit rate and reduces the parsing complexity of some implementations when pictures are signaled in slice headers, especially for low-delay and low-bit-rate applications.

[0171] In one embodiment, this only applies to raster scan slice mode (rect_slice_flag equals 0), which reduces parsing complexity in some implementations.

[0172] Table 11 Partial slice header showing changes JPEG0007733777000017.jpg49125

[0173] Avoid sending tile counts in slices if not necessary In one embodiment, when the picture header is transmitted in the slice header, the number of tiles in the slice is not transmitted. Table 12 shows this embodiment, where the num_tiles_in_slice_minus1 syntax element is not transmitted when the flag picture_header_in_slice_header_flag is set to 1. The advantage of this embodiment is bitrate reduction, especially for low-delay and low-bitrate applications, since the number of tiles does not need to be transmitted. In one embodiment, this only applies to raster scan slice mode (rect_slice_flag equals 0), which reduces parsing complexity in some implementations.

[0174] Table 12 Partial slice header showing modifications JPEG0007733777000018.jpg41124

[0175] PPS value predicted by NumTilesInPic (Semantics) In a further embodiment, when a picture header is transmitted in a slice header, the number of tiles in the current slice is inferred to be equal to the number of tiles in the picture. This can be configured by adding the following statement to the semantics of the syntax element num_tiles_in_slice_minus1: "If the variable num_tiles_in_slice_minus1 does not exist, its value is set to NumTilesInPic-1." Here, the variable NumTilesInPic gives the maximum number of tiles in a picture. This variable is calculated based on the syntax elements sent in the PPS.

[0176] Set the number of tiles before the slice address to avoid unnecessary sending of slice_address. In one embodiment, a syntax element dedicated to the number of tiles in a slice is transmitted before the slice address, and its value is used to know whether it is necessary to decode the slice address. More precisely, the number of tiles in a slice is compared with the number of tiles in a picture to know whether it is necessary to decode the slice address. Indeed, if the number of tiles in a slice is equal to the number of tiles in a picture, it is guaranteed that the current picture contains only one slice. In one embodiment, this only applies to raster scan slice mode (rect_slice_flag equals 0), which reduces parsing complexity in some implementations.

[0177] Table 13 illustrates this embodiment, where the syntax element slice_address is not decoded if the value of the syntax element num_tiles_in_slice_minus1 is equal to the variable NumTilesInPic-1. If num_tiles_in_slice_minus1 is equal to the variable NumTilesInPic-1, slice_address is inferred to be equal to 0.

[0178] Table 13 Partial slice header showing modifications JPEG0007733777000019.jpg57123

[0179] The advantage of this embodiment is bit rate reduction and parsing complexity reduction when the condition is set equal to true, since slice addresses are not transmitted. In one embodiment, when a picture header is transmitted in a slice header, the syntax element indicating the number of tiles in the current slice is not decoded and the number of tiles in the slice is inferred to be equal to 1. And, the slice address is inferred to be equal to 0 and the associated syntax element is not decoded when the number of tiles in the slice is equal to the number of tiles in the picture. Table 14 illustrates this embodiment. This increases the bitrate reduction obtained by combining these two embodiments.

[0180] Table 14 Partial slice header showing modifications JPEG0007733777000020.jpg55124

[0181] Remove the unnecessary condition numTileInPic > 1 In one embodiment, because the syntax element slice_address and / or the number of tiles in the current slice are decoded, the condition that the number of tiles in the current picture must be greater than 1 does not need to be tested when the raster scan slice mode is enabled. Specifically, if the number of tiles in the current picture is equal to 1, the rect_slice_flag value is inferred to be equal to 1. Therefore, in this case, the raster scan slice mode cannot be enabled. Table 15 illustrates this embodiment. This embodiment reduces the complexity of parsing the slice header.

[0182] Table 15 Partial slice header showing modifications JPEG0007733777000021.jpg49124

[0183] In one embodiment, when a picture header is sent in a slice header and raster scan slice mode is enabled, the syntax element indicating the number of tiles in the current slice is not decoded and the number of tiles in the slice is inferred to be equal to 1. And, the slice address is inferred to be equal to 0 and the associated syntax element slice_address is not decoded when the number of tiles in the slice is equal to the number of tiles in the picture and raster scan slice mode is enabled. Table 16 illustrates this embodiment. The advantages are reduced bit rate and reduced analysis complexity.

[0184] Table 16 Partial Slice Header Showing Modifications JPEG0007733777000022.jpg53123

[0185] implementation 11 illustrates a system 191, 195 comprising at least one of the encoder 150 or the decoder 100 and a communication network 199 according to an embodiment of the present invention. According to one embodiment, the system 195 is for processing and providing content (e.g., video and audio content for displaying / outputting or streaming the video / audio content) to a user who has access to the decoder 100, for example, via a user interface of a user terminal including the decoder 100 or a user terminal capable of communicating with the decoder 100. Such a user terminal may be a computer, a mobile phone, a tablet, or any other type of device capable of providing / displaying (provided / streamed) content to a user. The system 195 obtains / receives a bitstream 101 (e.g., in the form of a continuous stream or signal while previous video / audio is being displayed / output) via the communication network 199. According to one embodiment, the system 191 is for processing content and storing the processed content, for example, the processed video and audio content for later display / output / streaming. The system 191 obtains / receives content including an original image sequence 151 that is received and processed by an encoder 150 (including filtering by a deblocking filter according to the present invention), which generates a bitstream 101 that is communicated to the decoder 100 via a communications network 199. The bitstream 101 is then communicated to the decoder 100 in several ways, for example being pre-generated by the encoder 150 and stored as data in a storage device (e.g., a server or cloud storage) within the communications network 199 until a user requests the content (i.e., bitstream data) from the storage device, at which point the data is communicated / streamed from the storage device to the decoder 100.The system 191 may also provide / stream content information (e.g., content titles and other meta / storage location data for identifying, selecting, and requesting content) for content stored in the storage device to the user (e.g., by communicating data for a user interface displayed on the user terminal), and may include a content providing device for receiving and processing user requests for content so that the requested content can be delivered / streamed from the storage device to the user terminal. Alternatively, the encoder 150 generates the bitstream 101 and communicates / streams it directly to the decoder 100 when a user requests content. The decoder 100 then receives the bitstream 101 (or signal) and performs filtering using a deblocking filter in accordance with the present invention to obtain / generate a video signal 109 and / or an audio signal, which is then used by the user terminal to provide the requested content to the user.

[0186] Any step of a method / process according to the present invention or function described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the step / function may be stored as one or more instructions, codes, or programs, or as one or more instructions, codes, or programs on a computer-readable medium, and executed by one or more hardware-based processing units, such as a programmable computing machine ("personal computer"), a DSP ("digital signal processor"), a circuit, a processor and memory, a general-purpose microprocessor or central processing unit, a microcontroller, an ASIC ("application-specific integrated circuit"), a field programmable logic array (FPGA), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein.

[0187] Embodiments of the present invention may also be implemented by a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of JCs (e.g., chipsets). Various components, modules, or units are described herein to illustrate functional aspects of devices / apparatuses configured to perform the embodiments, but do not necessarily require implementation by different hardware units. Rather, the various modules / units may be combined in a codec hardware unit or may be provided by a collection of interoperating hardware units including one or more processors along with appropriate software / firmware.

[0188] Embodiments of the present invention can be realized by a computer of a system or device that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium to execute modules / units / functions of one or more of the above-described embodiments and / or one or more central processing units or circuits for executing the functions of one or more of the above-described embodiments, and by a method executed by the computer of the system or device, e.g., by reading and executing the computer-executable instructions from the storage medium to execute the functions of one or more of the above-described embodiments and / or by controlling the one or more central processing units or circuits to execute the functions of one or more of the above-described embodiments. The computer may include separate computers or a network of separate processing units to read and execute the computer-executable instructions. The computer-executable instructions may be provided to the computer from a computer-readable medium, such as a communication medium, for example, via a network or a tangible storage medium. The communication medium may be a signal / bitstream / carrier wave. The tangible storage medium may be, for example, a hard disk, a random access memory, a read-only memory, a storage device in a distributed computing system, an optical disc (compact disc (CD), digital versatile disc (DVD), or Blu-ray disc (BD)). TMThe "non-transitory computer-readable storage medium" may include one or more of a memory card, flash memory device, memory card, etc. At least some of the steps / functions may also be implemented in hardware by machines or dedicated components such as FPGAs ("field programmable gate arrays") or ASICs ("application-specific integrated circuits").

[0189] FIG. 12 is a schematic block diagram of a computing device 2000 for implementing one or more embodiments of the present invention. The computing device 2000 may be a device such as a microcomputer, a workstation, or a light portable device. The computing device 2000 includes a communication bus connected to: a central processing unit 2001, such as a microprocessor; a random access memory 2002 for storing executable code of embodiments of the present invention; and registers adapted to record variables and parameters necessary for implementing a method for encoding or decoding at least a portion of an image according to embodiments of the present invention, the memory capacity of which may be expanded, for example, by an optional RAM connected to an expansion port; a read-only memory 2003 for storing a computer program for implementing embodiments of the present invention; and a network interface 2004, typically connected to a communications network over which digital data to be processed is transmitted or received. The network interface (NET) 2004 may be a single network interface or may consist of a set of different network interfaces (e.g., wired and wireless interfaces, or different types of wired or wireless interfaces). Data packets are written to the network interface for transmission or read from the network interface for reception under the control of a software application running on the CPU 2001. - A user interface (UI) 2005 may be used to receive input from the user and display information to the user. - A hard disk (HD) 2006 may be provided as mass storage. - An input / output module (IO) 2007 may be used to send and receive data to external devices such as video sources and displays. The executable code may be stored either in the ROM 2003, the HD 2006 or on a removable digital medium such as a disk.According to a variant, the executable code of the program can be received by means of a communications network, via NET 2004, to be stored in one of the storage means of the communications device 2000, such as HD 2006, before being executed. CPU 2001 is adapted to control and direct the execution of instructions of a program or part of a software code according to an embodiment of the invention, which instructions are stored in one of the aforementioned storage means. After power-on, CPU 2001 can execute instructions for a software application, for example from main RAM memory 2002, after these instructions have been loaded from program ROM 2003 or HD 2006. Such a software application, when executed by CPU 2001, causes the steps of the method according to the invention to be carried out.

[0190] It will also be appreciated that, according to another embodiment of the present invention, the decoder according to the above-described embodiments is provided in a user terminal such as a computer, a mobile phone (cell phone), a table, or any other type of device (e.g., a display device) that can provide / display content to a user. According to yet another embodiment, the encoder according to the above-described embodiments is provided in an image capture device that also comprises a camera, a digital video camera, or a network camera (e.g., a closed-circuit television or video surveillance camera) that captures and provides content for the encoder to encode. Two such examples are provided below with reference to Figures 13 and 14.

[0191] Network Camera FIG. 13 is a diagram illustrating a network camera system 2100 including a network camera 2102 and a client device 2104 . The network camera 2102 includes an imaging unit 2106 , an encoding unit 2108 , a communication unit 2110 , and a control unit 2112 . A network camera 2102 and a client device 2104 are connected via a network 200 so that they can communicate with each other.

[0192] The imaging unit 2106 includes a lens and an imaging element (e.g., a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS)) to capture an image of a subject and generate image data based on the image, which can be a still image or a video image.

[0193] The encoding unit 2108 encodes the image data using the encoding method described above. The communication unit 2110 of the network camera 2102 transmits the encoded image data encoded by the encoding unit 2108 to the client device 2104 . The communication unit 2110 also receives commands from the client device 2104. The commands include commands to set parameters for encoding in the encoding unit 2108. The control unit 2112 controls the other units in the network camera 2102 according to the commands received by the communication unit 2110 .

[0194] The client device 2104 includes a communication unit 2114 , a decoding unit 2116 , and a control unit 2118 . The communication unit 2114 of the client device 2104 sends a command to the network camera 2102 . Additionally, the communication unit 2114 of the client device 2104 receives the encoded image data from the network camera 2102 . The decoding unit 2116 decodes the coded image data using the above-mentioned decoding method.

[0195] The control unit 2118 of the client device 2104 controls other units within the client device 2104 in response to user operations and commands received by the communication unit 2114 . The control unit 2118 of the client device 2104 controls the display device 2120 to display the image decoded by the decoding unit 2116 . In addition, the control unit 2118 of the client device 2104 controls the display device 2120 to display a GUI (Graphical User Interface) and specifies parameter values ​​for the network camera 2102, including parameters for encoding by the encoding unit 2108. Furthermore, the control unit 2118 of the client unit 2104 controls other units within the client unit 2104 in response to user operation inputs to the GUI displayed by the display unit 2120 . The control unit 2119 of the client device 2104 controls the communication unit 2114 of the client device 2104 to send a command to the network camera 2102 specifying the parameter values ​​of the network camera 2102 in response to a user operation input to the GUI displayed by the display device 2120.

[0196] Smartphone FIG. 14 is a diagram illustrating a smartphone 2200. The smartphone 2200 comprises a communication unit 2202 , a decoding unit 2204 , a control unit 2206 , a display unit 2208 , an image recording device 2210 and a sensor 2212 . The communication unit 2202 receives the encoded image data via the network 200 . The decoding unit 2204 decodes the coded image data received by the communication unit 2202 . The decoding unit 2204 decodes the coded image data using the above-mentioned decoding method. The control unit 2206 controls other units within the smartphone 2200 according to user operations or commands received by the communication unit 2202 . For example, the control unit 2206 controls the display unit 2208 to display the image decoded by the decoding unit 2204 .

[0197] Although the present invention has been described above using exemplary embodiments, the present invention is not limited to these embodiments. It will be understood by those skilled in the art that various changes and modifications can be made without departing from the scope of the present invention, as defined by the appended claims. All features disclosed in the specification (including any accompanying claims, abstract, and drawings) and / or all steps of any method or process so disclosed may be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. Each feature disclosed in the specification (including any accompanying claims, abstract, and drawings) may be replaced by an alternative feature serving the same, equivalent, or similar purpose, unless expressly stated otherwise. Thus, unless otherwise specified, each disclosed feature is merely an example of a generic series of equivalent or similar functions.

[0198] It will also be understood that any result of the above comparison, determination, evaluation, selection, execution, performance, or consideration, e.g., a selection made during an encoding or filtering process, may be indicated in or determinable / inferable from data in the bitstream, e.g., a flag or data indicating the result, so that the indicated or determined / inferred result may be used in processing, e.g., during a decoding process, in place of actually performing the comparison, determination, evaluation, selection, execution, performance, or consideration.

[0199] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.

[0200] Reference signs appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims.

Claims

1. 1. A method for decoding video data from a bitstream, the bitstream having coded video data corresponding to one or more slices, each slice having one or more tiles, the bitstream having a picture header including a plurality of syntax elements used when decoding the one or more slices, and a slice header including a plurality of syntax elements used when decoding the slices; The method comprises: Parsing a plurality of syntax elements; decoding the weighted prediction parameters from the picture header according to a value of a first flag in a picture parameter set included in the bitstream, the first flag being a flag related to the presence of weighted prediction parameters in the picture header; decoding the video data from the bitstream using the parsed syntax elements; If the second syntax element parsed from the slice header indicates that the picture header is in the slice header, (a) constraining parsing of a first syntax element indicating an address of a slice included in a picture to be omitted; (b) parsing a third syntax element, which is parsed in response to two conditions being satisfied, and which represents the result of subtracting 1 from the number of tiles included in the slice, is omitted when the second syntax element indicates that the picture header is in the slice header because one of the two conditions is not satisfied, and the value of the third syntax element is inferred to be equal to 0 regardless of the number of tiles in the picture. A method characterized by:

2. 2. The method of claim 1, wherein if the second syntax element indicates that the picture header is in the slice header, the value of the first syntax element is inferred to be 0.

3. 2. The method of claim 1, wherein each of the two conditions does not include the logical operator &&.

4. The method of claim 3, wherein && means and.

5. 1. A method for encoding video data into a bitstream, the bitstream comprising encoded video data corresponding to one or more slices, each slice comprising one or more tiles, the bitstream comprising a picture header comprising a plurality of syntax elements used when decoding the one or more slices, and a slice header comprising a plurality of syntax elements used when decoding the slices; The method comprises: encoding a plurality of syntax elements; encoding the video data using the plurality of syntax elements; encoding the weighted prediction parameters in the picture header according to a value of a first flag in a picture parameter set included in the bitstream, the first flag being a flag related to the presence of weighted prediction parameters in the picture header; When a second syntax element coded in the slice header indicates that the picture header is in the slice header, (a) a value of a first syntax element indicating an address of a slice included in a picture is constrained to be 0; (b) a third syntax element, which is coded in response to two conditions being satisfied and represents the result of subtracting 1 from the number of tiles included in the slice, is not coded when the second syntax element indicates that the picture header is in the slice header because one of the two conditions is not satisfied, and the value of the third syntax element is inferred to be equal to 0 regardless of the number of tiles in the picture. A method characterized by:

6. 6. The method of claim 5, wherein if the second syntax element indicates that the picture header is in the slice header, the value of the second syntax element is inferred to be 0.

7. 6. The method of claim 5, wherein each of the two conditions does not include the logical operator &&.

8. The method of claim 7, wherein && means and.

9. 1. A decoder for decoding video data from a bitstream, the bitstream having coded video data corresponding to one or more slices, each slice having one or more tiles, the bitstream having a picture header including a plurality of syntax elements used when decoding the one or more slices, and a slice header including a plurality of syntax elements used when decoding the slices; The decoder means for parsing a plurality of syntax elements; means for decoding the weighted prediction parameters from the picture header according to a value of a first flag in a picture parameter set included in the bitstream, the first flag being a flag related to the presence of weighted prediction parameters in the picture header; means for decoding the video data from the bitstream using the parsed syntax elements; If the second syntax element parsed from the slice header indicates that the picture header is in the slice header, (a) constraining parsing of a first syntax element indicating an address of a slice included in a picture to be omitted; (b) parsing a third syntax element, which is parsed in response to two conditions being satisfied, and which represents the result of subtracting 1 from the number of tiles included in the slice, is omitted when the second syntax element indicates that the picture header is in the slice header because one of the two conditions is not satisfied, and the value of the third syntax element is inferred to be equal to 0 regardless of the number of tiles in the picture. A decoder characterized by:

10. 1. An encoder for encoding video data into a bitstream, the bitstream having encoded video data corresponding to one or more slices, each slice having one or more tiles, the bitstream having a picture header including a plurality of syntax elements used when decoding the one or more slices, and a slice header including a plurality of syntax elements used when decoding the slices; The encoder comprises: means for encoding a plurality of syntax elements; means for encoding the video data using the plurality of syntax elements; means for encoding the weighted prediction parameters in the picture header according to a value of a first flag in a picture parameter set included in the bitstream, the first flag being a flag related to the presence of weighted prediction parameters in the picture header; When a second syntax element coded in the slice header indicates that the picture header is in the slice header, (a) a value of a first syntax element indicating an address of a slice included in a picture is constrained to be 0; (b) a third syntax element, which is coded in response to two conditions being satisfied and represents the result of subtracting 1 from the number of tiles included in the slice, is not coded when the second syntax element indicates that the picture header is in the slice header, because one of the two conditions is not satisfied, and the value of the third syntax element is inferred to be equal to 0 regardless of the number of tiles in the picture. An encoder characterized by:

11. A computer program product for causing a computer to carry out the method of claim 1.

12. A computer program for causing a computer to carry out the method of claim 5.

Citation Information

Patent Citations

  • Encoder, decoder and corresponding method for simplifying signaling slice header syntax elements

    JP2023515175A

  • An encoder, a decoder and corresponding methods simplifying signalling slice header syntax elements

    WO2021170132A1